<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2022.872107</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning Based Automatic Grape Downy Mildew Detection</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Zhao</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1366337/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Qiao</surname> <given-names>Yongliang</given-names></name>
<xref ref-type="aff" rid="aff5"><sup>5</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1401473/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Guo</surname> <given-names>Yangyang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>He</surname> <given-names>Dongjian</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1367254/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>College of Mechanical and Electronic Engineering, Northwest A&#x00026;F University</institution>, <addr-line>Xianyang</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>College of Electronic and Electrical Engineering, Baoji University of Arts and Sciences</institution>, <addr-line>Baoji</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>Key Laboratory of Agricultural Internet of Things, Ministry of Agriculture and Rural Affairs, Northwest A&#x00026;F University</institution>, <addr-line>Xianyang</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>Shaanxi Key Laboratory of Agricultural Information Perception and Intelligent Service, Northwest A&#x00026;F University</institution>, <addr-line>Xianyang</addr-line>, <country>China</country></aff>
<aff id="aff5"><sup>5</sup><institution>Faculty of Engineering, Australian Centre for Field Robotics (ACFR), The University of Sydney</institution>, <addr-line>Sydney, NSW</addr-line>, <country>Australia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yiannis Ampatzidis, University of Florida, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Brun Francois, Association de Coordination Technique Agricole, France; Harald Scherm, University of Georgia, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Yongliang Qiao <email>y.qiao&#x00040;acfr.usyd.edu.au</email></corresp>
<corresp id="c002">Dongjian He <email>hdj168&#x00040;nwsuaf.edu.cn</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Sustainable and Intelligent Phytoprotection, a section of the journal Frontiers in Plant Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>872107</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Zhang, Qiao, Guo and He.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhang, Qiao, Guo and He</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Grape downy mildew (GDM) disease is a common plant leaf disease, and it causes serious damage to grape production, reducing yield and fruit quality. Traditional manual disease detection relies on farm experts and is often time-consuming. Computer vision technologies and artificial intelligence could provide automatic disease detection for real-time controlling the spread of disease on the grapevine in precision viticulture. To achieve the best trade-off between GDM detection accuracy and speed under natural environments, a deep learning based approach named YOLOv5-CA is proposed in this study. Here coordinate attention (CA) mechanism is integrated into YOLOv5, which highlights the downy mildew disease-related visual features to enhance the detection performance. A challenging GDM dataset was acquired in a vineyard under a nature scene (consisting of different illuminations, shadows, and backgrounds) to test the proposed approach. Experimental results show that the proposed YOLOv5-CA achieved a detection precision of 85.59%, a recall of 83.70%, and a mAP&#x00040;0.5 of 89.55%, which is superior to the popular methods, including Faster R-CNN, YOLOv3, and YOLOv5. Furthermore, our proposed approach with inference occurring at 58.82 frames per second, could be deployed for the real-time disease control requirement. In addition, the proposed YOLOv5-CA based approach could effectively capture leaf disease related visual features resulting in higher GDE detection accuracy. Overall, this study provides a favorable deep learning based approach for the rapid and accurate diagnosis of grape leaf diseases in the field of automatic disease detection.</p></abstract>
<kwd-group>
<kwd>grape downy mildew</kwd>
<kwd>disease detection</kwd>
<kwd>deep learning</kwd>
<kwd>attention mechanism</kwd>
<kwd>data augmentation</kwd>
<kwd>digital agriculture</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Key Research and Development Program of China<named-content content-type="fundref-id">10.13039/501100012166</named-content></contract-sponsor>
<contract-sponsor id="cn002">Key Research and Development Program of Ningxia<named-content content-type="fundref-id">10.13039/100016692</named-content></contract-sponsor>
<counts>
<fig-count count="6"/>
<table-count count="3"/>
<equation-count count="13"/>
<ref-count count="59"/>
<page-count count="12"/>
<word-count count="7828"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>Grape as an important fruit crop makes a large economic income contribution in many countries (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>; Zhou et al., <xref ref-type="bibr" rid="B57">2021</xref>). As the grape grows in a natural condition, diseases will often appear on the leaves due to the complex weather condition and changing surrounding environments. Grape downy mildew (GDM) is one of the serious diseases caused by the oomycete pathogen Plamopara viticola, which seriously affects the growth of the grapes, causes a decrease in quality and yield, and results in huge economic losses in the grape industry (Chen et al., <xref ref-type="bibr" rid="B7">2020</xref>; Ji et al., <xref ref-type="bibr" rid="B19">2020</xref>). Downy mildew often happened in wet and rainy areas in spring and summer, it is initiated at the stomata on the underside of the leaf, and then on the whole leaf (Chen et al., <xref ref-type="bibr" rid="B7">2020</xref>). Monitoring grape leaf health and detecting pathogen are essential to reduce disease spread and facilitate effective management practices. Grape leaf diseases are currently controlled by repetitive fungicide treatments throughout the season. Reducing the treatment costs is a major challenge from both environmental and economic views. Timely detection and treatment at the initial stage of downy mildew infection (Adeel et al., <xref ref-type="bibr" rid="B3">2019</xref>) is a good solution to control and cut down the spread of downy mildew in a large area. Therefore, if an automatic detection method can be achieved when the spots appear, the leaf disease control plan can be made to control the diseases, guarantee the grape plant health, and improve the quality and yield of the grapes. Vision based detection approaches have been developed to detect plant diseases, which is performed by extracting visual features (e.g., texture, shape, and color of leaf lesions) of leaf images and using models (e.g., support vector machine, linear regression) to recognize and detect the diseases (Tang et al., <xref ref-type="bibr" rid="B42">2020</xref>; Hern&#x000E1;ndez et al., <xref ref-type="bibr" rid="B15">2021</xref>). Zhu et al. (<xref ref-type="bibr" rid="B58">2020</xref>) identified grape diseases using image analysis and BP neural networks. Chen et al. (<xref ref-type="bibr" rid="B7">2020</xref>) developed and compared several generalized linear models to predict the probability of high incidence and severity in the Bordeaux vineyard region. Abdelghafour et al. (<xref ref-type="bibr" rid="B2">2020</xref>) detected downy mildew symptoms using proximal color imaging and achieved 83% pixel-wise precision. However, the traditional image processing technology needs to manually extract the leaf disease characteristics, which is often time-consuming, and easy to miss the best disease prevention time. In addition, under the nature scene (e.g., different illumination, symptoms, camera viewpoints), classical algorithms or models lack robustness and cannot achieve stable detect performance.</p>
<p>Many scholars have proposed approaches for earlier plant disease detection and monitoring of the disease symptoms (Mutka and Bart, <xref ref-type="bibr" rid="B31">2015</xref>). At the earlier stage, the human-crafted features such as texture, color, or shape characteristics are extracted from RGB or hyper-spectral plant leaf images for identifying the plant diseases (Mahlein, <xref ref-type="bibr" rid="B28">2016</xref>). For example, Atanassova et al. (<xref ref-type="bibr" rid="B5">2019</xref>) proposed spectral data based classification models to predict the infection in plants, which achieved over 78% accuracy. Waghmare et al. (<xref ref-type="bibr" rid="B46">2016</xref>) proposed an automatic grape diseases detection system using the extracted color Local Binary Pattern (LBP) features. Mohammadpoor et al. (<xref ref-type="bibr" rid="B30">2020</xref>) proposed a support vector machine for grape fanleaf virus detection and achieved 98.6% average accuracy. However, this kind of method mostly depends on selected features and their extraction is easy to be influenced by the camera viewpoints, shadows, and lighting.</p>
<p>In recent years, deep learning methods such as convolutional neural networks (CNN) have been widely implemented in leaf disease detection, scene perception, and smart agriculture. Variates CNN based detection methods have been proposed for leaf disease recognition and monitoring (Liu et al., <xref ref-type="bibr" rid="B22">2020</xref>). Ferentinos (<xref ref-type="bibr" rid="B12">2018</xref>) proposed convolutional neural network architectures to identify healthy or diseased plants. Arsenovic et al. (<xref ref-type="bibr" rid="B4">2019</xref>) developed a two-stage architecture of neural networks to classify plant disease and achieved an accuracy of 93.67%. Zhang et al. (<xref ref-type="bibr" rid="B53">2019a</xref>) proposed an AlexNet based cucumber disease identification approach, achieving 94.65% recognition accuracy. Ji et al. (<xref ref-type="bibr" rid="B19">2020</xref>) proposed CNN based approach to classify common grape leaf diseases and obtained average classification accuracy of 98.57%. Liu et al. (<xref ref-type="bibr" rid="B22">2020</xref>) proposed Inception convolutional neural network (DICNN) for identifying grape leaf diseases and realized an overall accuracy of 97.22% on single-leaf datasets. Thet et al. (<xref ref-type="bibr" rid="B43">2020</xref>) used an improved VGG16 model that achieved 98.4% classification accuracy for five different leaf diseases. Tang et al. (<xref ref-type="bibr" rid="B42">2020</xref>) classified grape disease types using lightweight convolution neural networks and channel-wise attention, which achieved 99.14% accuracy. Liu and Wang (<xref ref-type="bibr" rid="B24">2020</xref>) improved the YOLOv3 model to directly generate the bounding box coordinates for tomato diseases and pests detection, which achieved a detection accuracy of 92.39%. According to these studies, CNNs can learn advanced robust features of leaf diseases directly from original images, outperforming the traditional feature extraction approaches. Yu and Son (<xref ref-type="bibr" rid="B50">2020</xref>) proposed a leaf spot attention mechanism to increase apple leaf disease discriminative power and enhance the identification performance. Hern&#x000E1;ndez and Lopez (<xref ref-type="bibr" rid="B16">2020</xref>) developed Bayesian deep learning techniques and an uncertainty probabilistic programming approach for plant disease detection.</p>
<p>With the continuous development of smart sensors, big data, and cloud computing, many automatic approaches have been proposed to identify and detect plant leaf diseases (Vishnoi et al., <xref ref-type="bibr" rid="B45">2021</xref>). The rapid development of artificial intelligence and the Internet of Things (IoT) has significantly facilitated automatic disease detection (Zhang et al., <xref ref-type="bibr" rid="B52">2020</xref>). Using deep learning models and noninvasive sensors to identify plant diseases has drawn more attention in the field of precision agriculture and plant phenotyping (Nagaraju and Chawla, <xref ref-type="bibr" rid="B32">2020</xref>; Singh et al., <xref ref-type="bibr" rid="B40">2020</xref>). Hern&#x000E1;ndez et al. (<xref ref-type="bibr" rid="B15">2021</xref>) investigated hyperspectral sensing technologies and artificial intelligence applications for assessing downy mildew in grapevine under laboratory conditions. Guti&#x000E9;rrez et al. (<xref ref-type="bibr" rid="B13">2021</xref>) differentiated downy mildew and spider mite in grapevine under field conditions using the CNN model. Liu et al. (<xref ref-type="bibr" rid="B23">2021</xref>) proposed Hierarchical Multi-Scale Attention Semantic Segmentation (HMASS) to identify GDM infected regions, and the calculated infection severity percentage was highly correlated (<italic>R</italic> = 0.96) with the human field assessment.</p>
<p>Choi and Hsiao (<xref ref-type="bibr" rid="B8">2021</xref>) classified Cassava leaf diseases using the Residual Network. Zhang et al. (<xref ref-type="bibr" rid="B51">2021</xref>) developed a multi-feature fusion Faster R-CNN model and achieved 83.34% detection accuracy for soybean leaf disease. Dinata et al. (<xref ref-type="bibr" rid="B11">2021</xref>) proposed CNN based approach for 6 types of strawberry disease classification and achieved 63.7% accuracy. Abbas et al. (<xref ref-type="bibr" rid="B1">2021</xref>) detected tomato plant disease using transfer learning and C-GAN synthetic images, which achieved 99.51% accuracy. Cristin et al. (<xref ref-type="bibr" rid="B10">2020</xref>) proposed a deep neural network based Rider-Cuckoo Search Algorithm and achieved 87.7% plant disease detection accuracy. Roy and Bhaduri (<xref ref-type="bibr" rid="B39">2021</xref>) proposed deep learning-based multi-class plant disease and achieved 91.2% mean average precision. However, most of these methods are only tested in experimental situations, which need to be verified on the complex background situation.</p>
<p>Despite deep learning based approaches demonstrating its facilitate in GDM detection, the detection accuracy and speed restricted its application in autonomous viticulture management. Plant leaf disease detection in the real vineyard is facing many challenges, such as the small difference between the lesion area and the background, different scales of the spots, variation of symptoms, and camera viewpoints (Liu and Wang, <xref ref-type="bibr" rid="B25">2021</xref>). Also, light changing in a real complex natural environment further increased the difficulty to achieve high detection accuracy. Therefore, real-time and accurate detection of grape downy mildew is of great significance for the scientific management and control of grape diseases in precision vineyard farming.</p>
<p>Recently, the attention mechanisms such as Squeeze-and-Excitation Networks (SE) (Hu et al., <xref ref-type="bibr" rid="B18">2018</xref>), Convolutional block attention module (CBAM) (Woo et al., <xref ref-type="bibr" rid="B49">2018</xref>), and CA (Hou et al., <xref ref-type="bibr" rid="B17">2021</xref>) have been widely used to enhance the deep learning model performances. SE simply squeezes each 2D feature map to efficiently build interdependencies among channels (Hu et al., <xref ref-type="bibr" rid="B18">2018</xref>). CBAM introduces spatial information encoding via convolutions with large-size kernels and gathers channel-wise and spatial-wise attention sequentially. The recently proposed CA adopts different spatial attention mechanisms and designs advanced attention blocks. Zhang et al. (<xref ref-type="bibr" rid="B54">2019b</xref>) applied an attention mechanism to object detection networks, enhancing the impact of significant features and weakening background interference. Experimental results show that the proposed approach achieved an object detection accuracy of 75.9% on PASCAL VOC 2007, which is 6% higher than Faster R-CNN. Liu et al. (<xref ref-type="bibr" rid="B26">2019</xref>) presented a deep neural network architecture based on information transmission and attention mechanisms. Zhao et al. (<xref ref-type="bibr" rid="B56">2021</xref>) diagnosed tomato leaf disease using an attention module improved network, which achieved 96.81% average identification accuracy on the tomato leaf diseases dataset. Ravi et al. (<xref ref-type="bibr" rid="B36">2021</xref>) integrated the attention module into the EfficientNet model to locate and identify the tiny infected regions in the Cassava leaf. The proposed method achieved better performance than non-attention-based CNN pre-trained models. Wang et al. (<xref ref-type="bibr" rid="B48">2021b</xref>) proposed a Fine-Grained grape leaf disease recognition method using a lightweight attention network, which can efficiently diagnose orchard grape leaf diseases with low computing cost. The above studies have demonstrated that attention mechanisms could enhance feature extraction ability for leaf disease detection and identification.</p>
<p>In this study, to improve GDM detection accuracy in the natural grape farm environment, we proposed YOLOv5-CA based GDM detection approach by combing YOLOv5 and CA mechanism. Different scales of image features were extracted through CNN layers of YOLOv5, and these features were weighted by CA for GDM detection. By using CA, the features&#x00027; effectiveness for GDM detection is highlighted and those less effective features are suppressed. The proposed YOLOv5-CA based GDM detection is tested on our acquired grape leaf image dataset.</p>
<p>The remaining part of the article is organized as follows. Section 2 illustrates the used datasets, the proposed approaches, and evaluation indicators. Experimental results are presented in Section 3. Discussions of the performance are presented in Section 4. Finally, conclusions and future areas for research are given in Section 5.</p></sec>
<sec id="s2">
<title>2. Material and Methodology</title>
<sec>
<title>2.1. Plant Material and Image Acquisition</title>
<p>Grape leaf image data were acquired in a commercial vineyard located in the college of Enology, Northwest A&#x00026;F University, north of China (Yangling, Shaanxi Province). The vineyard manifested downy mildew (Plasmopara viticola) in many plants. Images were taken manually for several days (each day is from 8:00 a.m. to 16:00 p.m.) in early August on a partly cloudy day (<xref ref-type="fig" rid="F1">Figure 1</xref>). The used camera is a Canon EDS 1200D (a field of view of approximately 504 mm horizontally and 360 mm vertically), and the external conditions for shooting are automatic mode. There is approximately 30 cm between the camera lens and the grape leaves.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Commercial vineyard and acquired images under natural light conditions.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0001.tif"/>
</fig>
<p>A total of 820 leaf samples were collected from different lights, leaf overlapping, and disease severity. The dataset is challenging considering the complex background, occlusions, different disease spot-areas, and shadow influence. <xref ref-type="fig" rid="F1">Figure 1</xref> shows images of diseased leaves in a typical complex environment in the dataset. Downy mildew first appears as brown patches. These patches gradually spread and a leaf that is severely affected may have a reduced yield with a shorter lifetime and fruits with a small size.</p>
<p>To validate the proposed YOLOv5-CA based GDM detection approach, the randomly selected 500 leaf images were used as training datasets, while the remaining 320 images were used as testing data. For experiment testing, the LabelImg annotation tool (Tzutalin, <xref ref-type="bibr" rid="B44">2015</xref>) was used to manually label the leaf disease areas.</p></sec>
<sec>
<title>2.2. YOLOv5-CA Based GDM Detection</title>
<p>In order to make YOLOv5 more suitable for GDM detection in complex natural scenarios such as complex background, occlusions, different disease spot-areas, and shadow influence, YOLOv5-CA based GDM detection approach is proposed to improve the GDM detection performance for real farming applications. Grape leaves&#x00027; RGB images were acquired under field conditions from a commercial vineyard. These collected images contain healthy and downy mildew infected leaves. Then detection model YOLOv5-CA was trained to identify the GDM infected regions. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the proposed YOLOv5-CA approach extracted features using YOLOv5 and learned key features through CA, enhancing the feature extraction ability and improving the leaf disease detection performance. As YOLOv5 could adjust the width and depth of the backbone network according to application requirements, for GDM detection, moderate model parameters (i.e., width and depth parameters are 0.75 and 0.67, respectively) were used to achieve reasonable detection speed.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The architecture of the proposed YOLOv5-CA based GDM detection.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0002.tif"/>
</fig>
<p>YOLOv5-CA network is mainly composed of backbone part, neck network, and head part: 1) The backbone of YOLOv5 is responsible for extracting image features, which includes several different layers types such as Focus, Conv, C3, CA, and Spatial Pyramid Pooling (SPP) layer. 2) The neck module generates a feature pyramid based on the PANet (Path Aggregation Network) (Liu et al., <xref ref-type="bibr" rid="B27">2018</xref>). It is a series of feature aggregation layers of mixed and combined image features, enhancing the ability to detect objects with different scales by fusing low-level spatial features and high-level semantic features. 3) The head module generates detection boxes, indicating the category, coordinates, and confidence by applying anchor boxes to multi-scale feature maps from the neck module. The proposed YOLOv5-CA boosts the detection ability of different GDM infection regions through an attention mechanism, which provides a feasible GDM detection and monitoring solution for automatic disease control.</p>
<sec>
<title>2.2.1. Backbone of YOLOv5-CA</title>
<p>The backbone of the YOLOV5-CA object detector mainly contains Focus, Conv, C3, CA, and Spatial Pyramid Pooling (SPP) layer. The features from deeper layers are more abstract and semantic, while the low-layer features contain spatial information and fine-grained features. For an input image, the Focus module rearranged it through stridden slice operations in both width and height dimensions, which reduces model calculation time. C3 module contains three convolutions and is used to extract the deep features of the image. The following SPP is used to improve the receptive field of the network by converting any size of the feature map into a fixed-size feature vector. SPP (He et al., <xref ref-type="bibr" rid="B14">2015</xref>) concatenates layer outputs with different kernel sizes (e.g., 13 &#x000D7; 13, 9 &#x000D7; 9, 5 &#x000D7; 5) to boost multi-scale image feature representation ability. All the convolutions utilize Swish activation:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>S</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3; denotes the sigmoid function.</p>
<p>In our study, we integrated the CA layer into the YOLOv5 backbone, CA layer factorizes channel attention into two 1D feature encoding processes and preserves the precise positional information, which augments the representations of the leaf disease regions.</p></sec>
<sec>
<title>2.2.2. Neck of YOLOv5-CA</title>
<p>The neck structure used in YOLOv5-CA is a PANet (Liu et al., <xref ref-type="bibr" rid="B27">2018</xref>), which fuses the information of all layers to aggregate features by combing bottom-up pyramid and element-wise max operations. PANet combines convolution features of different layers for images, thus the useful information in each feature layer can be directly propagated to the following subnetwork. By this, PANet can not only realize the abstract description of large objects but also retains the feature details of small objects. In addition, C3 modules are also added at this stage to enhance the feature fusion capability. Through the neck part, the features of infected areas can be extracted to maintain the detection performance.</p></sec>
<sec>
<title>2.2.3. CA Layer</title>
<p>In terms of GDM detection, because the GDM is randomly distributed in the grape leaf, there is inevitably a mix of overlapping occlusion, and the GDM infection regions account for a relatively small percentage of the images, resulting in missed and mis-detected. In our study, a plug-and-play CA layer was introduced to assist YOLOv5 focused on key disease-related features, and improve the detection accuracy.</p>
<p>The CA layer embeds the location-aware information into the channel attention simultaneously, which increases the spatial range of attention and avoids a lot of computational overhead (Hou et al., <xref ref-type="bibr" rid="B17">2021</xref>). CA layer can be regarded as a computational unit that enhances the representation ability of the learned features. For any intermediate feature <inline-formula><mml:math id="M2"><mml:mi>X</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x022EF;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, CA could outputs a transformed feature with augmented representations <italic>Y</italic> &#x0003D; [<italic>y</italic>1, <italic>y</italic>2, &#x022EF;&#x02009;, <italic>y</italic><sub><italic>c</italic></sub>] of the same size to <italic>X</italic>.</p>
<p>As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the CA mechanism can be divided into two parts: the coordinate information embedding part (encodes the information of the channels in the horizontal and vertical coordinates) and the coordinate attention generation part (captures the positional information and generates the weight values).</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Schematic of coordinate attention module.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0003.tif"/>
</fig></sec>
<sec>
<title>2.2.4. Coordinate Information Embedding</title>
<p>Attention mechanisms have been demonstrated helpful to enhance the overall performance of deep learning models (Chorowski et al., <xref ref-type="bibr" rid="B9">2015</xref>). The attention mechanism can be regarded as a feature weighting scheme, which helps the deep learning model to pay more attention to the task-related information, and suppress or ignore the less-contribution features (Li et al., <xref ref-type="bibr" rid="B21">2020</xref>; Mi et al., <xref ref-type="bibr" rid="B29">2020</xref>). Through this, the attention mechanism strengthens the deep learning model&#x00027;s learning ability and boosts performance (Niu et al., <xref ref-type="bibr" rid="B33">2021</xref>). In recent years, attention mechanisms based on deep learning networks have been applied to a wide variety of computer vision tasks such as image classification, object detection, and image segmentation (Qiao et al., <xref ref-type="bibr" rid="B35">2019</xref>, <xref ref-type="bibr" rid="B34">2021</xref>). Wang et al. (<xref ref-type="bibr" rid="B47">2021a</xref>) developed a deep attention module for vegetable and fruit leaf plant disease detection. Kerkech et al. (<xref ref-type="bibr" rid="B20">2020</xref>) used a fully convolutional neural network approach to classify Unmanned Aerial Vehicle (UAV) image pixels for detecting mildew disease.</p>
<p>It is known that channel attention could increase the value of the important channel while punishing the non-significant channels, however, channel attention is difficult to preserve positional information (Zhang et al., <xref ref-type="bibr" rid="B55">2018</xref>). To capture precise positional information, the global average pooling was factorized into the average pooling from two directions of each channel. Specifically, given the input <italic>X</italic>, two spatial extents of pooling kernels (H, 1) and (W, 1) were used to encode each channel along the horizontal coordinate and the vertical dimensions, respectively. The output of the <italic>c</italic>-<italic>th</italic> channel along height <italic>h</italic> and width <italic>w</italic> dimensions can be formulated as:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>H</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mn>0</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>j</mml:mi><mml:mo>&#x0003C;</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>z</italic><sup><italic>h</italic></sup> and <italic>z</italic><sup><italic>w</italic></sup> are the outputs of the transform at <italic>h</italic> direction width <italic>w</italic>, respectively; <italic>x</italic><sub><italic>c</italic></sub> is the feature map at <italic>c</italic>-<italic>th</italic> channel; <italic>W</italic> and <italic>H</italic> are the width and height dimensions of the feature map separately.</p>
<p>The Equation (2) encodes each channel along with the horizontal and vertical coordinates, preserving the positional information of each channel of feature maps, which facilitates the network to locate the GDM-related visual features precisely.</p></sec>
<sec>
<title>2.2.5. Coordinate Attention Generation</title>
<p>To further exploit resulting expressive representations, a simple and effective coordinate attention generation was used as the second transformation. Here, the obtained feature maps from the coordinate information embedding stage were concatenated and then sent to a shared 1 &#x000D7; 1 convolution layer. The relevant process is defined as:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where [, ] indicates concatenate operation, <italic>F</italic> is 1 &#x000D7; 1 convolution operation; <inline-formula><mml:math id="M5"><mml:mi>f</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>W</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>H</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is the output feature map of the ReLU layer, <italic>r</italic> is reduction rate.</p>
<p>Next, the feature map <italic>f</italic> was decomposed into two separate tensors: <inline-formula><mml:math id="M6"><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mi>H</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M7"><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Then the following two 1 &#x000D7; 1 convolution layers for <italic>f</italic><sup><italic>h</italic></sup> and <italic>f</italic><sup><italic>w</italic></sup>, respectively, are recovered to the same shape as <italic>z</italic><sup><italic>h</italic></sup> and <italic>z</italic><sup><italic>w</italic></sup>. The operation is formulated as:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M8"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3; is the sigmoid activation function, and <italic>F</italic><sub><italic>h</italic></sub> and <italic>F</italic><sub><italic>w</italic></sub> are the convolution manipulation for <italic>f</italic><sup><italic>h</italic></sup> and <italic>f</italic><sup>&#x003C9;</sup> separately.</p>
<p>The obtained feature maps <italic>g</italic><sup><italic>h</italic></sup> and <italic>g</italic><sup><italic>w</italic></sup> are then expanded and used as attention weights for the horizontal and vertical coordinates, respectively. This operation can enhance the effective leaf disease related features and reduce the influence of less important information. The reweighing process of the original input feature map can be defined as:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>h</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:msubsup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>y</italic><sub><italic>c</italic></sub> is the <italic>c</italic>-<italic>th</italic> channel in the generated feature map <italic>y</italic> of the attention block.</p></sec></sec>
<sec>
<title>2.3. YOLOv5-CA Model Training for GDM Detection</title>
<sec>
<title>2.3.1. Network Training Parameters</title>
<p>In our study, the experimental platform is based on a computer equipped with an NVIDIA RTX 1080Ti GPU, Ryzen 7 3600 CPU&#x00040;3.6 GHz. The proposed GDM detection approach was implemented using Pytorch.</p>
<p>In addition, to verify the effectiveness of the YOLOv5-CA based GDM detection approach, Faster R-CNN (Ren et al., <xref ref-type="bibr" rid="B38">2015</xref>), YOLOv4 (Bochkovskiy et al., <xref ref-type="bibr" rid="B6">2020</xref>), and YOLOv5 (Tzutalin, <xref ref-type="bibr" rid="B44">2015</xref>) were also used for comparison. Faster R-CNN generates regions of interest (RoIs) candidates and then classifies them into objects (and background) and refines the boundaries of those regions. YOLOv4 and YOLOv5 are the two widely used detection methods from the YOLO series (Redmon et al., <xref ref-type="bibr" rid="B37">2016</xref>).</p>
<p>For network training, the network&#x00027;s input size was set to 416 &#x000D7; 416 &#x000D7; 3, the training epoch was set to 1000, batch size was set to 16, and the learning rate was 0.0013. The momentum factor (momentum) was set to 0.937, the initial learning rate was 1 &#x000D7; 10<sup>&#x02212;5</sup> and the decay rate of weight was set to 0.001. The other parameters of each network are their default settings. In the training process, the network predicts the bounding box based on the initial anchor box. The gap between the prediction and ground truth was calculated to update the network in reverse and adjusts the network parameters. After training, the weight file of the detection model obtained was saved.</p></sec>
<sec>
<title>2.3.2. Network Loss Function</title>
<p>YOLOV5-CA automatically updates the best bounding box for GDM detection during the training process. The default optimization method of the model is the gradient descent method. The loss function <italic>L</italic><sub><italic>loss</italic></sub> used in YOLOv5-CA includes bounding box location loss <italic>L</italic><sub><italic>Ciou</italic></sub>, confidence loss <italic>L</italic><sub><italic>conf</italic></sub> and classification loss <italic>L</italic><sub><italic>cls</italic></sub>:</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Classification loss <italic>L</italic><sub><italic>cls</italic></sub> computes the loss of class probability using Cross Entropy:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M11"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:munderover><mml:mrow><mml:msubsup><mml:mi>&#x02113;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mstyle><mml:mstyle displaystyle='true'><mml:munder><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mover><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mo stretchy='false'>(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>c</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mstyle></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>&#x02113;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is used to judge whether there is an object center. <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability of class <italic>c</italic>; <italic>p</italic><sub><italic>i</italic></sub>(<italic>c</italic>) is the probability of predicted box that belongs to class <italic>c</italic>.</p>
<p>Confidence loss <italic>L</italic><sub><italic>conf</italic></sub> penalizes object confidence error if that predictor is responsible for the ground truth box, which is computed using mean squared error:</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M14"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>&#x02113;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mover><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>+</mml:mo><mml:msub><mml:mtext>&#x003BB;</mml:mtext><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:munderover><mml:mrow><mml:mstyle displaystyle='true'><mml:munderover><mml:mo>&#x02211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>&#x02113;</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mstyle></mml:mrow></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mover><mml:mrow><mml:msub><mml:mi>C</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>&#x02227;</mml:mo></mml:mover><mml:mo stretchy='false'>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BB;<sub><italic>noobj</italic></sub> represents the weight of the classification error, <italic>S</italic> is the number of grids, and <italic>B</italic> is the number of prior boxes in each grid; <italic>C</italic><sub><italic>i</italic></sub> is the confidence of the predicted box; &#x00108;<sub><italic>i</italic></sub> is the confidence of the ground-truth (&#x00108;<sub><italic>i</italic></sub> is always 1).</p>
<p>The <italic>L</italic><sub><italic>Ciou</italic></sub> computes the loss related to the predicted bounding box and ground truth, it can be defined as follows:</p>
<disp-formula id="E9"><label>(9)</label><mml:math id="M"><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x003C1;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msup><mml:mi>e</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>&#x003BD;</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo stretchy='false'>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x02212;</mml:mo><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo stretchy='false'>)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x003BD;</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x02229;</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x0222A;</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>v</italic> represents the coincidence degree of the two frame aspect ratios, <italic>b</italic> and <italic>b</italic><sup><italic>gt</italic></sup> are the center coordinates of the prediction box and the real box respectively; &#x003C1; is the Euclidean distance between the two center points, and <italic>e</italic> represents the diagonal distance of the smallest closed area containing both the prediction and real boxes. <italic>IoU</italic> means the ratio of the intersection and union of the prediction bounding box and the actual bounding box.</p></sec>
<sec>
<title>2.3.3. Performance Evaluation</title>
<p>The used performance evaluation indicators for GDM detection include precision, recall, <italic>F</italic><sub>1</sub>-score, mAP (mean average precision), and FPS (frame per second). Precision shows the ability of the model to accurately identify targets; recall reflects the ability of the model to detect targets; the <italic>F</italic><sub>1</sub>-score is a harmonic mean of the precision and recall; FPS is the average inference speed. The <italic>F</italic><sub>1</sub>-score is the reconciled mean of precision and recall, taking into account both the precision and recall of the classification model. Based on tp (the number of hlcorrectly detected downy mildew areas), fp (the number of incorrectly detected downy mildew areas), and fn (the number of disease regions that are incorrectly identified as background), the relevant calculation equations are as follows:</p>
<disp-formula id="E10"><label>(10)</label><mml:math id="M16"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>f</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mn>100</mml:mn><mml:mi>%</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E11"><label>(11)</label><mml:math id="M17"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>t</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>p</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>f</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mn>100</mml:mn><mml:mi>%</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E12"><label>(12)</label><mml:math id="M18"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x000D7;</mml:mo><mml:mn>100</mml:mn><mml:mi>%</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>From the values of precision and recall, a precision-recall curve can be plotted to observe their distribution. The value of AP is the area under the precision-recall curve, and a larger value means better model performance. mAP&#x00040;0.5 is the average value of precision under different recall values when the intersection over union (IoU) is 0.5. The calculation of mAP is as follows:</p>
<disp-formula id="E13"><label>(13)</label><mml:math id="M19"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:mi>A</mml:mi><mml:msub><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>N</italic> denotes the number of disease types (<italic>N</italic> is 1 in our study).</p></sec></sec></sec>
<sec id="s3">
<title>3. Experimental Results</title>
<sec>
<title>3.1. Comparison of Different Object Detection Algorithms</title>
<p>There are varieties of deep learning based detection methods, in order to verify the effectiveness of the proposed method for GDM detection, three popular detection algorithms&#x02014;Faster R-CNN, YOLOv4, and YOLOv5 were compared. The GDM detection results were presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Comparison of different GDM methods.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Precision (%)</bold></th>
<th valign="top" align="center"><bold>Recall  (%)</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub>  (%)</bold></th>
<th valign="top" align="center"><bold>mAP&#x00040;0.5 (%)</bold></th>
<th valign="top" align="center"><bold>FPS (Frame/s)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Faster R-CNN</td>
<td valign="top" align="center">79.97</td>
<td valign="top" align="center">87.80</td>
<td valign="top" align="center">83.70</td>
<td valign="top" align="center">80.65</td>
<td valign="top" align="center">35.90</td>
</tr>
<tr>
<td valign="top" align="left">YOLOv4</td>
<td valign="top" align="center">82.69</td>
<td valign="top" align="center">83.63</td>
<td valign="top" align="center">83.15</td>
<td valign="top" align="center">82.65</td>
<td valign="top" align="center">75.20</td>
</tr>
<tr>
<td valign="top" align="left">YOLOv5</td>
<td valign="top" align="center">85.35</td>
<td valign="top" align="center">81.45</td>
<td valign="top" align="center">83.36</td>
<td valign="top" align="center">87.41</td>
<td valign="top" align="center">84.74</td>
</tr>
<tr>
<td valign="top" align="left">YOLOv5-CA</td>
<td valign="top" align="center">85.59</td>
<td valign="top" align="center">83.70</td>
<td valign="top" align="center">84.63</td>
<td valign="top" align="center">89.55</td>
<td valign="top" align="center">58.82</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In <xref ref-type="table" rid="T1">Table 1</xref>, the proposed YOLOv5-CA based approach achieved 85.59% precision, 83.70% recall and 84.63% <italic>F</italic><sub>1</sub>, and 89.55% mAP, respectively. Compared with the other methods, the proposed YOLOv5-CA GDM detection method is better than that of Faster R-CNN (80.65% mAP), YOLOv4 (82.65 % mAP), and YOLOv5 (87.41% mAP). From these results, it is clear that the CA mechanism of YOLOv5-CA improves the feature representation ability, enhancing the final detection accuracy for identifying the leaf disease areas. Meanwhile, the proposed approach could detect the GDM with a speed of 58.82 frames per second. These results illustrated that the proposed method could achieve high precision with a fast speed to meet real-time requirements, which is favorable for the deployment of the GDM detection model in spraying robots for the plant diseases control in smart vineyard farming.</p></sec>
<sec>
<title>3.2. Qualitative GDM Detection Comparison</title>
<p><xref ref-type="fig" rid="F4">Figure 4</xref> demonstrates the comparison of different methods&#x00027; qualitative results on our acquired grape leaf dataset. It can be seen that the proposed YOLOv5-CA could detect GDM at different leaf parts (e.g., leaf edge, the leaf central parts). Especially, the proposed YOLOv5-CA method could detect less obvious GDM lesions on the leaves, which outperformed the other methods such as Faster R-CNN, YOLOv4, and YOLOv5. The main reason could be that the CA mechanism strengthens the feature representation ability, which enhances the GDM detection performance.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Examples of different GDM detection attention methods. <bold>(A)</bold> Faster R-CNN, <bold>(B)</bold> YOLOv4, <bold>(C)</bold> YOLOv5, and <bold>(D)</bold> YOLOV5-CA.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0004.tif"/>
</fig>
<p>Additionally, more examples of YOLOv5-CA based GDM detection are presented in <xref ref-type="fig" rid="F5">Figure 5</xref>. It can be seen that the GDM infected regions are well detected (blue bounding box) under complex background, especially, YOLOv5-CA could well detect the GDM regions nearby the leaf edge and petioles. It also can be noted that the YOLOv5-CA could detect both large and small GDM regions. The main reason is that the YOLOv5-CA makes the network pay more attention to the GDM-related visual features, reducing the false or mis-detection cases. The good detection performance of YOLOv5-CA provides valuable information for automatic disease control.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Examples of YOLOv5-CA based GDM detection results.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0005.tif"/>
</fig></sec>
<sec>
<title>3.3. Influence of Different Network Input-Sizes on GDM Detection</title>
<p>The network input size is one factor that would influence the GDM detection performance. Here, we also investigate different input-sizes&#x00027; influence on YOLOv5-CA based GDM detection. In <xref ref-type="table" rid="T2">Table 2</xref>, five typical network input sizes, namely, 112 &#x000D7;112, 224 &#x000D7;224, 320 &#x000D7;320, 416 &#x000D7;416, and 512 &#x000D7;512 were compared in terms of GDM detection performance.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Grape downy mildew Detection performance with different network input sizes.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Network input size</bold></th>
<th valign="top" align="center"><bold>Precision  (%)</bold></th>
<th valign="top" align="center"><bold>Recall  (%)</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub> (%)</bold></th>
<th valign="top" align="center"><bold>mAP&#x00040;0.5</bold></th>
<th valign="top" align="center"><bold>FPS (Frame/s)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">112 &#x000D7;112</td>
<td valign="top" align="center">80.32</td>
<td valign="top" align="center">72.76</td>
<td valign="top" align="center">76.35</td>
<td valign="top" align="center">76.71</td>
<td valign="top" align="center">102.04</td>
</tr>
<tr>
<td valign="top" align="left">224 &#x000D7;224</td>
<td valign="top" align="center">83.73</td>
<td valign="top" align="center">79.32</td>
<td valign="top" align="center">81.47</td>
<td valign="top" align="center">82.63</td>
<td valign="top" align="center">92.63</td>
</tr>
<tr>
<td valign="top" align="left">320 &#x000D7;320</td>
<td valign="top" align="center">84.75</td>
<td valign="top" align="center">84.32</td>
<td valign="top" align="center">84.53</td>
<td valign="top" align="center">85.25</td>
<td valign="top" align="center">76.92</td>
</tr>
<tr>
<td valign="top" align="left">416 &#x000D7;416</td>
<td valign="top" align="center">85.59</td>
<td valign="top" align="center">83.70</td>
<td valign="top" align="center">84.63</td>
<td valign="top" align="center">89.55</td>
<td valign="top" align="center">58.82</td>
</tr>
<tr>
<td valign="top" align="left">512 &#x000D7;512</td>
<td valign="top" align="center">86.71</td>
<td valign="top" align="center">82.80</td>
<td valign="top" align="center">84.71</td>
<td valign="top" align="center">87.89</td>
<td valign="top" align="center">45.45</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to <xref ref-type="table" rid="T2">Table 2</xref>, the network input with 416 &#x000D7;416 size achieved 85.59% precision, 83.70% recall, 84.63% <italic>F</italic><sub>1</sub>-score, and 89.55% mAP&#x00040;0.5, which outperformed the performance of input size with 112 &#x000D7;112, 224 &#x000D7;224, and 512 &#x000D7;512. This means the proposed YOLOv5-CA could extract and learn the more useful information from the large input size. However, when the network input-size increases to 512 &#x000D7;512, there is not much performance improvement but significantly increased the processing time and calculating memory size, which is not favorable for fast detection and real applications. By balancing the speed and accuracy, the input size of 416 &#x000D7;416 was selected in our work for real-time GDM detection.</p></sec>
<sec>
<title>3.4. Data Augmentation for YOLOv5-CA Detection</title>
<p>Offline data augmentation could increase the dataset diversity, explore the network hyperparameters, and finally enhance the accuracy and robustness of the trained model (Zoph et al., <xref ref-type="bibr" rid="B59">2020</xref>; Su et al., <xref ref-type="bibr" rid="B41">2021</xref>). To further improve the GDM detection performance, in our study, bounding box based data augmentation was used. The augmentation technique was only applied to disease areas within the manually labeled bounding boxes. The transformations for data augmentation implemented include: flipping horizontally and vertically, randomly cropping between 0 and 20% of the bounding box, random rotation, random shear of between &#x02212;15&#x000B0; to &#x0002B;15&#x000B0; horizontally and vertically, random brightness adjustment (between 0 and &#x0002B;10%), and Gaussian blur (between 0 and 5 pixels). The original 500 training images were expanded to 2000 images, and then they were used to train the YOLOv5-CA network, which forces neural nets to optimize hyperparameters and generate a high-robust model. Some augmented bounding boxes on grape leaves can be seen in <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Examples of Bounding box based data augmentation. <bold>(A)</bold> Manual label, <bold>(B)</bold> Flip (vertical), <bold>(C)</bold> Flip (horizontal), <bold>(D)</bold> Crop, <bold>(E)</bold> Rotation, <bold>(F)</bold> Shear, <bold>(G)</bold> Brightness, and <bold>(H)</bold> Gaussian blur.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-872107-g0006.tif"/>
</fig>
<p>As illustrated in <xref ref-type="table" rid="T3">Table 3</xref>, the data augmentation based YOLOv5-CA detection achieved a precision of 88.82%, a recall of 83.63%, and an <italic>F</italic><sub>1</sub>-score of 86.15%, which is slightly higher than those without data augmentation. The data augmentation positively influences the model&#x00027;s performance by increasing the size of the dataset and mitigating the over-fitting. The overall improvements demonstrated that the data augmentation module is helpful in the GDM detection, enlarging model learning ability and significantly improving detection performance.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Comparison of different GDM methods.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Precision (%)</bold></th>
<th valign="top" align="center"><bold>Recall (%)</bold></th>
<th valign="top" align="center"><bold><italic>F</italic><sub>1</sub> (%)</bold></th>
<th valign="top" align="center"><bold>mAP&#x00040;0.5 (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">YOLOv5-CA</td>
<td valign="top" align="center">85.59</td>
<td valign="top" align="center">83.70</td>
<td valign="top" align="center">84.63</td>
<td valign="top" align="center">89.55</td>
</tr>
<tr>
<td valign="top" align="left">YOLOv5-CA  (with data augmentation)</td>
<td valign="top" align="center">88.82</td>
<td valign="top" align="center">83.63</td>
<td valign="top" align="center">86.15</td>
<td valign="top" align="center">90.02</td>
</tr>
</tbody>
</table>
</table-wrap></sec></sec>
<sec id="s4">
<title>4. Discussions</title>
<p>This study presents a deep learning-based pipeline for automatic GDM detection in the vineyard. The grape leaf images acquired directly from the plants under field conditions were used to verify our proposed YOLOv5-CA approach. According to our experimental results presented in <xref ref-type="table" rid="T1">Table 1</xref>, a precision of 85.59%, a recall of 83.70%, an <italic>F</italic><sub>1</sub>-score of 84.63%, and a mAP&#x00040;0.5 of 89.55% with the inference speed of 58.82 frames per second (FPS) was obtained for GDM detection. The detection accuracy of the proposed YOLOv5-CA is superior to that of state-of-the-art methods such as Faster R-CNN, YOLOv4, and YOLOv5. This high accuracy demonstrates the effectiveness of YOLOv5-CA for GDM detection of grapevine leaf images taken under field conditions. There yield results show that it is feasible to model visual symptoms for automatic GDM detection using a combination of the YOLOv5 and the CA mechanism. The proposed YOLOv5-CA automatically finds complex features capable of differentiating leaves with downy mildew symptoms and without any, which provides a precise and effective method for automatic disease detection.</p>
<p>On the other hand, the results presented in <xref ref-type="table" rid="T2">Table 2</xref> reveal the appropriate network input size in our experiments is 416 &#x000D7;416. Additionally, <xref ref-type="table" rid="T3">Table 3</xref> compared the GDM detection performance with and without data augmentation, it shows that data augmentation enhances the GDM detection performance. The possible reason is that data augmentation increases the size of the dataset and brings more diversity to leverage the model training.</p>
<p>Although this study mainly focuses on GDM detection, it is suitable for multi-diseases detection (e.g., black spot, powdery mildew) after the model was re-trained with the dataset containing these diseases. As our approach uses RGB images, it would be a restriction for detecting GDM in the very earlier stage (i.e., non-visible symptoms) detection. However, if the multi-spectral images were acquired and used, our proposed YOLOv5-CA could be a potential tool to distinguish downy mildew from other leaf diseases/damage.</p></sec>
<sec id="s5">
<title>5. Conclusions and Future Study</title>
<p>To achieve an accurate and real-time intelligent detection of GDM under natural environments, an automatic YOLOv5-CA based detection method was proposed in this study. By combing YOLOv5 and coordinate attention, the GDM related visual features are well focused on and extracted, which boosts the GDM detection performance. Our proposed YOLOv5-CA achieved 85.59% detection precision, 89.55% mAP&#x00040;0.5 with 58.82 FPS, which outperformed Faster R-CNN, YOLOv4, and YOLOv5. Moreover, the test results showed that the different disease levels of GDM and the illumination influence would not have a great impact on the GDM detection results, indicating the proposed method is feasible for the rapid and accurate detection of GDM. Ablation studies show that a network input size of 416 &#x000D7;416 is favorable for fast GDM detection, and bounding box-based data augmentation boosts the GDM detection precision by 3.23%. The results exposed in this work indicate that downy mildew in grapevine can be automatically evaluated using artificial intelligence technology.</p>
<p>Overall, our approach achieved a good trade-off between speed and accuracy for GDM, and can be adapted to applications with autonomous-based smart farming. For future study, the multi-spectral information and edge computing will be exploited to further improve detection performance and computational efficiency.</p></sec>
<sec sec-type="data-availability" id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding authors.</p></sec>
<sec id="s7">
<title>Author Contributions</title>
<p>ZZ: investigation, methodology, writing-review, and editing. YQ: data curation, methodology, formal analysis, and writing-original draft. YG: writing-review and editing. DH: resources and article revising. All authors contributed to the article and approved the submitted version.</p></sec>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This research was funded by the National Key Research and Development Program of China (2019YFD1002500), Ningxia Hui Autonomous Region Key Research and Development Program (2021BEF02015), and Ningxia Hui Autonomous Region Flexible Introduction Team Project (2020RXTDLX08).</p></sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec> </body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Abbas</surname> <given-names>A.</given-names></name> <name><surname>Jain</surname> <given-names>S.</given-names></name> <name><surname>Gour</surname> <given-names>M.</given-names></name> <name><surname>Vankudothu</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Tomato plant disease detection using transfer learning with c-gan synthetic images</article-title>. <source>Comput. Electron. Agric</source>. <volume>187</volume>, <fpage>1106279</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.106279</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abdelghafour</surname> <given-names>F.</given-names></name> <name><surname>Keresztes</surname> <given-names>B.</given-names></name> <name><surname>Germain</surname> <given-names>C.</given-names></name> <name><surname>Da Costa</surname> <given-names>J.-P.</given-names></name></person-group> (<year>2020</year>). <article-title>In field detection of downy mildew symptoms with proximal colour imaging</article-title>. <source>Sensors</source> <volume>20</volume>, <fpage>4380</fpage>. <pub-id pub-id-type="doi">10.3390/s20164380</pub-id><pub-id pub-id-type="pmid">32764472</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Adeel</surname> <given-names>A.</given-names></name> <name><surname>Khan</surname> <given-names>M. A.</given-names></name> <name><surname>Sharif</surname> <given-names>M.</given-names></name> <name><surname>Azam</surname> <given-names>F.</given-names></name> <name><surname>Shah</surname> <given-names>J. H.</given-names></name> <name><surname>Umer</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Diagnosis and recognition of grape leaf diseases: An automated system based on a novel saliency approach and canonical correlation analysis based multiple features fusion</article-title>. <source>Sustainable Comput</source>. <volume>24</volume>, <fpage>1100349</fpage>. <pub-id pub-id-type="doi">10.1016/j.suscom.2019.08.002</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Arsenovic</surname> <given-names>M.</given-names></name> <name><surname>Karanovic</surname> <given-names>M.</given-names></name> <name><surname>Sladojevic</surname> <given-names>S.</given-names></name> <name><surname>Anderla</surname> <given-names>A.</given-names></name> <name><surname>Stefanovic</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Solving current limitations of deep learning based approaches for plant disease detection</article-title>. <source>Symmetry</source> <volume>11</volume>, <fpage>939</fpage>. <pub-id pub-id-type="doi">10.3390/sym11070939</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Atanassova</surname> <given-names>S.</given-names></name> <name><surname>Nikolov</surname> <given-names>P.</given-names></name> <name><surname>Valchev</surname> <given-names>N.</given-names></name> <name><surname>Masheva</surname> <given-names>S.</given-names></name> <name><surname>Yorgov</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Early detection of powdery mildew (podosphaera xanthii) on cucumber leaves based on visible and near-infrared spectroscopy</article-title>. <source>AIP Conf. Proc</source>. <volume>2075</volume>, <fpage>160014</fpage>. <pub-id pub-id-type="doi">10.1063/1.5091341</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bochkovskiy</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>C.-Y.</given-names></name> <name><surname>Liao</surname> <given-names>H.-Y. M.</given-names></name></person-group> (<year>2020</year>). <article-title>Yolov4: optimal speed and accuracy of object detection</article-title>. <source>arXiv preprint</source> arXiv:2004.10934. <pub-id pub-id-type="doi">10.48550/arXiv.2004.10934</pub-id><pub-id pub-id-type="pmid">34300543</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Brun</surname> <given-names>F.</given-names></name> <name><surname>Raynal</surname> <given-names>M.</given-names></name> <name><surname>Makowski</surname> <given-names>D.</given-names></name></person-group> (<year>2020</year>). <article-title>Forecasting severe grape downy mildew attacks using machine learning</article-title>. <source>PLoS ONE</source> <volume>15</volume>, <fpage>e0230254</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0230254</pub-id><pub-id pub-id-type="pmid">32163490</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Choi</surname> <given-names>H. C.</given-names></name> <name><surname>Hsiao</surname> <given-names>T.-C.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Image classification of cassava leaf disease based on residual network,&#x0201D;</article-title> in <source>2021 IEEE 3rd Eurasia Conference on Biomedical Engineering, Healthcare and Sustainability (ECBIOS)</source> (<publisher-loc>Tainan</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>185</fpage>&#x02013;<lpage>186</lpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chorowski</surname> <given-names>J.</given-names></name> <name><surname>Bahdanau</surname> <given-names>D.</given-names></name> <name><surname>Serdyuk</surname> <given-names>D.</given-names></name> <name><surname>Cho</surname> <given-names>K.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Attention-based models for speech recognition</article-title>. <source>arXiv preprint</source> arXiv:1506.07503. <pub-id pub-id-type="doi">10.48550/arXiv.1506.07503</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cristin</surname> <given-names>R.</given-names></name> <name><surname>Kumar</surname> <given-names>B. S.</given-names></name> <name><surname>Priya</surname> <given-names>C.</given-names></name> <name><surname>Karthick</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Deep neural network based rider-cuckoo search algorithm for plant disease detection</article-title>. <source>Artif. Intell. Rev</source>. <volume>53</volume>, <fpage>4993</fpage>&#x02013;<lpage>5018</lpage>. <pub-id pub-id-type="doi">10.1007/s10462-020-09813-w</pub-id><pub-id pub-id-type="pmid">35140732</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dinata</surname> <given-names>M. I.</given-names></name> <name><surname>Nugroho</surname> <given-names>S. M. S.</given-names></name> <name><surname>Rachmadi</surname> <given-names>R. F.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Classification of strawberry plant diseases with leaf image using CNN,&#x0201D;</article-title> in <source>2021 International Conference on Artificial Intelligence and Computer Science Technology (ICAICST)</source> (<publisher-loc>Yogyakarta</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>68</fpage>&#x02013;<lpage>72</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferentinos</surname> <given-names>K. P..</given-names></name></person-group> (<year>2018</year>). <article-title>Deep learning models for plant disease detection and diagnosis</article-title>. <source>Comput. Electron. Agric</source>. <volume>145</volume>, <fpage>3111</fpage>&#x02013;<lpage>3318</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2018.01.009</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Guti&#x000E9;rrez</surname> <given-names>S.</given-names></name> <name><surname>Hern&#x000E1;ndez</surname> <given-names>I.</given-names></name> <name><surname>Ceballos</surname> <given-names>S.</given-names></name> <name><surname>Barrio</surname> <given-names>I.</given-names></name> <name><surname>D&#x000ED;ez-Navajas</surname> <given-names>A. M.</given-names></name> <name><surname>Tardaguila</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Deep learning for the differentiation of downy mildew and spider mite in grapevine under field conditions</article-title>. <source>Comput. Electron. Agric</source>. <volume>182</volume>, <fpage>1105991</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.105991</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Spatial pyramid pooling in deep convolutional networks for visual recognition</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>. <volume>37</volume>, <fpage>1904</fpage>&#x02013;<lpage>1916</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2389824</pub-id><pub-id pub-id-type="pmid">26353135</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez</surname> <given-names>I.</given-names></name> <name><surname>Guti&#x000E9;rrez</surname> <given-names>S.</given-names></name> <name><surname>Ceballos</surname> <given-names>S.</given-names></name> <name><surname>I&#x000F1;&#x000ED;guez</surname> <given-names>R.</given-names></name> <name><surname>Barrio</surname> <given-names>I.</given-names></name> <name><surname>Tardaguila</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Artificial intelligence and novel sensing technologies for assessing downy mildew in grapevine</article-title>. <source>Horticulturae</source> <volume>7</volume>, <fpage>103</fpage>. <pub-id pub-id-type="doi">10.3390/horticulturae7050103</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hern&#x000E1;ndez</surname> <given-names>S.</given-names></name> <name><surname>Lopez</surname> <given-names>J. L.</given-names></name></person-group> (<year>2020</year>). <article-title>Uncertainty quantification for plant disease detection using bayesian deep learning</article-title>. <source>Appl. Soft. Comput</source>. <volume>96</volume>, <fpage>1106597</fpage>. <pub-id pub-id-type="doi">10.1016/j.asoc.2020.106597</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hou</surname> <given-names>Q.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Feng</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Coordinate attention for efficient mobile network design,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Nashville, TN</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>13713</fpage>&#x02013;<lpage>13722</lpage>.</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Shen</surname> <given-names>L.</given-names></name> <name><surname>Sun</surname> <given-names>G.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Squeeze-and-excitation networks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>7132</fpage>&#x02013;<lpage>7141</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ji</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name></person-group> (<year>2020</year>). <article-title>Automatic grape leaf diseases identification via unitedmodel based on multiple convolutional neural networks</article-title>. <source>Inf. Process. Agric</source>. <volume>7</volume>, <fpage>418</fpage>&#x02013;<lpage>426</lpage>. <pub-id pub-id-type="doi">10.1016/j.inpa.2019.10.003</pub-id></citation>
</ref>
<ref id="B20">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kerkech</surname> <given-names>M.</given-names></name> <name><surname>Hafiane</surname> <given-names>A.</given-names></name> <name><surname>Canals</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>Vine disease detection in uav multispectral images using optimized image registration and deep learning segmentation approach</article-title>. <source>Comput. Electron. Agric</source>. <volume>174</volume>, <fpage>1105446</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2020.105446</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Cheng</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>Object detection based on an adaptive attention mechanism</article-title>. <source>Sci Rep</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-67529-x</pub-id><pub-id pub-id-type="pmid">32647299</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Ding</surname> <given-names>Z.</given-names></name> <name><surname>Tian</surname> <given-names>L.</given-names></name> <name><surname>He</surname> <given-names>D.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Grape leaf disease identification using improved deep convolutional neural networks</article-title>. <source>Front. Plant. Sci</source>. <volume>11</volume>, <fpage>11082</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.01082</pub-id><pub-id pub-id-type="pmid">32760419</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>E.</given-names></name> <name><surname>Gold</surname> <given-names>K. M.</given-names></name> <name><surname>Combs</surname> <given-names>D.</given-names></name> <name><surname>Cadle-Davidson</surname> <given-names>L.</given-names></name> <name><surname>Jiang</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Deep learning-based autonomous downy mildew detection and severity estimation in vineyards,&#x0201D;</article-title> in <source>2021 ASABE Annual International Virtual Meeting</source> (<publisher-loc>American Society of Agricultural and Biological Engineers</publisher-loc>).</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Tomato diseases and pests detection based on improved yolo v3 convolutional neural network</article-title>. <source>Front. Plant Sci</source>. <volume>11</volume>, <fpage>8198</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.00898</pub-id><pub-id pub-id-type="pmid">32612632</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Plant diseases and pests detection based on deep learning: a review</article-title>. <source>Plant Methods</source> <volume>17</volume>, <fpage>1</fpage>&#x02013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1186/s13007-021-00722-9</pub-id><pub-id pub-id-type="pmid">33627131</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>R.</given-names></name> <name><surname>Cheng</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Remote sensing image change detection based on information transmission and attention mechanism</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>1156349</fpage>&#x02013;<lpage>1156359</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2947286</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Qi</surname> <given-names>L.</given-names></name> <name><surname>Qin</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Path aggregation network for instance segmentation,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8759</fpage>&#x02013;<lpage>8768</lpage>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mahlein</surname> <given-names>A.-K..</given-names></name></person-group> (<year>2016</year>). <article-title>Plant disease detection by imaging sensors-parallels and specific demands for precision agriculture and plant phenotyping</article-title>. <source>Plant Dis</source>. <volume>100</volume>, <fpage>241</fpage>&#x02013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.1094/PDIS-03-15-0340-FE</pub-id><pub-id pub-id-type="pmid">30694129</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mi</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Su</surname> <given-names>J.</given-names></name> <name><surname>Han</surname> <given-names>D.</given-names></name> <name><surname>Su</surname> <given-names>B.</given-names></name></person-group> (<year>2020</year>). <article-title>Wheat stripe rust grading by deep learning with attention mechanism and images from mobile devices</article-title>. <source>Front. Plant Sci</source>. <volume>11</volume>, <fpage>558126</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.558126</pub-id><pub-id pub-id-type="pmid">33013976</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohammadpoor</surname> <given-names>M.</given-names></name> <name><surname>Nooghabi</surname> <given-names>M. G.</given-names></name> <name><surname>Ahmedi</surname> <given-names>Z.</given-names></name></person-group> (<year>2020</year>). <article-title>An intelligent technique for grape fanleaf virus detection</article-title>. <source>Int. J. Interact. Multim. Artif. Intell</source>. <volume>6</volume>, <fpage>62</fpage>&#x02013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.9781/ijimai.2020.02.001</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Mutka</surname> <given-names>A. M.</given-names></name> <name><surname>Bart</surname> <given-names>R. S.</given-names></name></person-group> (<year>2015</year>). <article-title>Image-based phenotyping of plant disease symptoms</article-title>. <source>Front. Plant Sci</source>. <volume>5</volume>, <fpage>7134</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2014.00734</pub-id><pub-id pub-id-type="pmid">25601871</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nagaraju</surname> <given-names>M.</given-names></name> <name><surname>Chawla</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Systematic review of deep learning techniques in plant disease detection</article-title>. <source>Int. J. Syst Assurance Eng. Manag</source>. <volume>11</volume>, <fpage>547</fpage>&#x02013;<lpage>560</lpage>. <pub-id pub-id-type="doi">10.1007/s13198-020-00972-1</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Niu</surname> <given-names>Z.</given-names></name> <name><surname>Zhong</surname> <given-names>G.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>A review on the attention mechanism of deep learning</article-title>. <source>Neurocomputing</source> <volume>452</volume>, <fpage>418</fpage>&#x02013;<lpage>462</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.03.091</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>Y.</given-names></name> <name><surname>Kong</surname> <given-names>H.</given-names></name> <name><surname>Clark</surname> <given-names>C.</given-names></name> <name><surname>Lomax</surname> <given-names>S.</given-names></name> <name><surname>Su</surname> <given-names>D.</given-names></name> <name><surname>Eiffert</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Intelligent perception-based cattle lameness detection and behaviour recognition: a review</article-title>. <source>Animals</source> <volume>11</volume>, <fpage>3033</fpage>. <pub-id pub-id-type="doi">10.3390/ani11113033</pub-id><pub-id pub-id-type="pmid">34827766</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Qiao</surname> <given-names>Y.</given-names></name> <name><surname>Truman</surname> <given-names>M.</given-names></name> <name><surname>Sukkarieh</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Cattle segmentation and contour extraction based on mask r-cnn for precision livestock farming</article-title>. <source>Comput. Electron. Agric</source>. <volume>165</volume>, <fpage>1104958</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.104958</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ravi</surname> <given-names>V.</given-names></name> <name><surname>Acharya</surname> <given-names>V.</given-names></name> <name><surname>Pham</surname> <given-names>T. D.</given-names></name></person-group> (<year>2021</year>). <article-title>Attention deep learning-based large-scale learning classifier for cassava leaf disease classification</article-title>. <source>Expert. Syst</source>. <volume>39</volume>, <fpage>e12862</fpage>. <pub-id pub-id-type="doi">10.1111/exsy.12862</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Redmon</surname> <given-names>J.</given-names></name> <name><surname>Divvala</surname> <given-names>S.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;You only look once: Unified, real-time object detection,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>779</fpage>&#x02013;<lpage>788</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Faster R-CNN: towards real-time object detection with region proposal networks</article-title>. <source>Adv. Neural Inf. Process. Syst</source>. <volume>28</volume>, <fpage>911</fpage>&#x02013;<lpage>999</lpage>. <pub-id pub-id-type="doi">10.48550/arXiv.1506.01497</pub-id><pub-id pub-id-type="pmid">27295650</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roy</surname> <given-names>A. M.</given-names></name> <name><surname>Bhaduri</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>A deep learning enabled multi-class plant disease detection model based on computer vision</article-title>. <source>AI</source> <volume>2</volume>, <fpage>413</fpage>&#x02013;<lpage>428</lpage>. <pub-id pub-id-type="doi">10.3390/ai2030026</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>V.</given-names></name> <name><surname>Sharma</surname> <given-names>N.</given-names></name> <name><surname>Singh</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>A review of imaging techniques for plant disease detection</article-title>. <source>Artif Intell Agric</source>. <volume>4</volume>, <fpage>229</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1016/j.aiia.2020.10.002</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Su</surname> <given-names>D.</given-names></name> <name><surname>Kong</surname> <given-names>H.</given-names></name> <name><surname>Qiao</surname> <given-names>Y.</given-names></name> <name><surname>Sukkarieh</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Data augmentation for deep learning based semantic segmentation and crop-weed classification in agricultural robotics</article-title>. <source>Comput. Electron. Agric</source>. <volume>190</volume>, <fpage>1106418</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.106418</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Qi</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>Grape disease image classification based on lightweight convolution neural networks and channelwise attention</article-title>. <source>Comput. Electron. Agric</source>. <volume>178</volume>, <fpage>1105735</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2020.105735</pub-id></citation>
</ref>
<ref id="B43">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Thet</surname> <given-names>K. Z.</given-names></name> <name><surname>Htwe</surname> <given-names>K. K.</given-names></name> <name><surname>Thein</surname> <given-names>M. M.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Grape leaf diseases classification using convolutional neural network,&#x0201D;</article-title> in <source>2020 International Conference on Advanced Information Technologies (ICAIT)</source> (<publisher-loc>Yangon</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>147</fpage>&#x02013;<lpage>152</lpage>.<pub-id pub-id-type="pmid">35062534</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="web"><person-group person-group-type="author"><collab>Tzutalin</collab></person-group> (<year>2015</year>). <source>Labelimg. git code (2015)</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://github.com/tzutalin/labelImg">https://github.com/tzutalin/labelImg</ext-link></citation>
</ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vishnoi</surname> <given-names>V. K.</given-names></name> <name><surname>Kumar</surname> <given-names>K.</given-names></name> <name><surname>Kumar</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Plant disease detection using computational intelligence and image processing</article-title>. <source>J. Plant Dis. Protect</source>. <volume>128</volume>, <fpage>19</fpage>&#x02013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1007/s41348-020-00368-0</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Waghmare</surname> <given-names>H.</given-names></name> <name><surname>Kokare</surname> <given-names>R.</given-names></name> <name><surname>Dandawate</surname> <given-names>Y.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Detection and classification of diseases of grape plant using opposite colour local binary pattern feature and machine learning for automated decision support system,&#x0201D;</article-title> in <source>2016 3rd International Conference on Signal Processing and Integrated Networks (SPIN)</source> (<publisher-loc>Noida</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>513</fpage>&#x02013;<lpage>518</lpage>.</citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>L.</given-names></name> <name><surname>Dong</surname> <given-names>H.</given-names></name> <name><surname>Yun</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2021a</year>). <article-title>Dba_ssd: a novel end-to-end object detection using deep attention module for helping smart device with vegetable and fruit leaf plant disease detection</article-title>. <source>Information</source> <volume>12</volume>, <fpage>474</fpage>. <pub-id pub-id-type="doi">10.21203/rs.3.rs-166579/v1</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Niu</surname> <given-names>T.</given-names></name> <name><surname>Mao</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2021b</year>). <article-title>Fine-grained grape leaf diseases recognition method based on improved lightweight attention network</article-title>. <source>Front. Plant Sci</source>. <volume>12</volume>, <fpage>738042</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2021.738042</pub-id><pub-id pub-id-type="pmid">34745172</pub-id></citation></ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Woo</surname> <given-names>S.</given-names></name> <name><surname>Park</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J.-Y.</given-names></name> <name><surname>Kweon</surname> <given-names>I. S.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Cbam: convolutional block attention module,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>3</fpage>&#x02013;<lpage>19</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>H.-J.</given-names></name> <name><surname>Son</surname> <given-names>C.-H.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Leaf spot attention network for apple leaf disease identification,&#x0201D;</article-title> in <source>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</source> (<publisher-loc>Seattle, WA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>52</fpage>&#x02013;<lpage>53</lpage>.</citation>
</ref>
<ref id="B51">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name></person-group> (<year>2021</year>). <article-title>Detecting soybean leaf disease from synthetic image using multi-feature fusion faster r-cnn</article-title>. <source>Comput. Electron. Agric</source>. <volume>183</volume>, <fpage>1106064</fpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2021.106064</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>N.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Pan</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Zhao</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>A review of advanced technologies and development for hyperspectral-based plant disease detection in the past three decades</article-title>. <source>Remote Sens</source>. <volume>12</volume>, <fpage>3188</fpage>. <pub-id pub-id-type="doi">10.3390/rs12193188</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Shi</surname> <given-names>Y.</given-names></name></person-group> (<year>2019a</year>). <article-title>Cucumber leaf disease identification with global pooling dilated convolutional neural network</article-title>. <source>Comput. Electron. Agric</source>. <volume>162</volume>, <fpage>4122</fpage>&#x02013;<lpage>4430</lpage>. <pub-id pub-id-type="doi">10.1016/j.compag.2019.03.012</pub-id></citation>
</ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Huang</surname> <given-names>C.</given-names></name> <name><surname>Gao</surname> <given-names>M.</given-names></name></person-group> (<year>2019b</year>). <article-title>Object detection network based on feature fusion and attention mechanism</article-title>. <source>Future Internet</source> <volume>11</volume>, <fpage>9</fpage>. <pub-id pub-id-type="doi">10.3390/fi11010009</pub-id></citation>
</ref>
<ref id="B55">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zhong</surname> <given-names>B.</given-names></name> <name><surname>Fu</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Image super-resolution using very deep residual channel attention networks,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision (ECCV)</source> (<publisher-loc>Munich</publisher-loc>), <fpage>286</fpage>&#x02013;<lpage>301</lpage>.<pub-id pub-id-type="pmid">34059829</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>S.</given-names></name> <name><surname>Peng</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>Tomato leaf disease diagnosis based on improved convolution neural network by attention module</article-title>. <source>Agriculture</source> <volume>11</volume>, <fpage>651</fpage>. <pub-id pub-id-type="doi">10.3390/agriculture11070651</pub-id></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>S.</given-names></name> <name><surname>Xing</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>Q.</given-names></name> <name><surname>Song</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Grape leaf spot identification under limited samples by fine grained-gan</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>1100480</fpage>&#x02013;<lpage>1100489</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3097050</pub-id></citation>
</ref>
<ref id="B58">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Identification of grape diseases using image analysis and bp neural networks</article-title>. <source>Multimed Tools Appl</source>. <volume>79</volume>, <fpage>14539</fpage>&#x02013;<lpage>14551</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-018-7092-0</pub-id><pub-id pub-id-type="pmid">32829161</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zoph</surname> <given-names>B.</given-names></name> <name><surname>Cubuk</surname> <given-names>E. D.</given-names></name> <name><surname>Ghiasi</surname> <given-names>G.</given-names></name> <name><surname>Lin</surname> <given-names>T.-Y.</given-names></name> <name><surname>Shlens</surname> <given-names>J.</given-names></name> <name><surname>Le</surname> <given-names>Q. V.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Learning data augmentation strategies for object detection,&#x0201D;</article-title> in <source>European Conference on Computer Vision</source> (<publisher-loc>Springer</publisher-loc>), <fpage>566</fpage>&#x02013;<lpage>583</lpage>.<pub-id pub-id-type="pmid">34659392</pub-id></citation></ref>
</ref-list> 
</back>
</article>