<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2023.1204569</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>ALAD-YOLO:an lightweight and accurate detector for apple leaf diseases</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Xu</surname>
<given-names>Weishi</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<xref ref-type="author-notes" rid="fn003">
<sup>&#x2020;</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Runjie</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn003">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2259422"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Intelligent Science and Technology, East China University of Science and Technology</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>School of Geological Engineering, Tongji University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Chaojun Hou, Zhongkai University of Agriculture and Engineering, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Prof. Dr Noor Zaman Jhanjhi, Taylor&#x2019;s University, Malaysia; Pushkar Gole, University of Delhi, India</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Weishi Xu, <email xlink:href="mailto:20002139@mail.ecust.edu.cn">20002139@mail.ecust.edu.cn</email>
</p>
</fn>
<fn fn-type="equal" id="fn003">
<p>&#x2020;These authors have contributed equally to this work and share first authorship</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1204569</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Xu and Wang</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Xu and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Suffering from various apple leaf diseases, timely preventive measures are necessary to take. Currently, manual disease discrimination has high workloads, while automated disease detection algorithms face the trade-off between detection accuracy and speed. Therefore, an accurate and lightweight model for apple leaf disease detection based on YOLO-V5s (ALAD-YOLO) is proposed in this paper. An apple leaf disease detection dataset is collected, containing 2,748 images of diseased apple leaves under a complex environment, such as from different shooting angles, during different spans of the day, and under different weather conditions. Moreover, various data augmentation algorithms are applied to improve the model generalization. The model size is compressed by introducing the Mobilenet-V3s basic block, which integrates the coordinate attention (CA) mechanism in the backbone network and replacing the ordinary convolution with group convolution in the Spatial Pyramid Pooling Cross Stage Partial Conv (SPPCSPC) module, depth-wise convolution, and Ghost module in the C3 module in the neck network, while maintaining a high detection accuracy. Experimental results show that ALAD-YOLO balances detection speed and accuracy well, achieving an accuracy of 90.2% (an improvement of 7.9% compared with yolov5s) on the test set and reducing the floating point of operations (FLOPs) to 6.1 G (a decrease of 9.7 G compared with yolov5s). In summary, this paper provides an accurate and efficient detection method for apple leaf disease detection and other related fields.</p>
</abstract>
<kwd-group>
<kwd>apple leaf disease</kwd>
<kwd>lightweight object detection</kwd>
<kwd>ALAD-YOLO</kwd>
<kwd>coordinate attention</kwd>
<kwd>group convolution</kwd>
</kwd-group>
<counts>
<fig-count count="12"/>
<table-count count="2"/>
<equation-count count="11"/>
<ref-count count="29"/>
<page-count count="15"/>
<word-count count="7653"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Technical Advances in Plant Science</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Apple is one of the most important crops with rich nutritional and medicinal values and is widely grown in the world. China, one of the largest apple producers, produced more than 41 million tons of apples in 2019, accounting for 54.07% of the global total (<xref ref-type="bibr" rid="B9">Hu et&#xa0;al., 2022</xref>). However, apples are often threatened by various foliar diseases caused by bacterium, such as brown spot disease and mosaic disease, which will lead to a drastic decrease in apple yield and quality without timely detection and prevention, causing significant economic losses to farmers. Therefore, an efficient and accurate diagnosis method for apple foliar diseases is essential to promoting the development of the apple industry, which is of great practical value.</p>
<p>Traditionally, apple leaf diseases are usually detected by manual inspection, which has many limitations and drawbacks. First, the method relies on professional inspectors for detection (<xref ref-type="bibr" rid="B15">Liu et&#xa0;al., 2017</xref>), thus making the limitation of human resources a serious problem. Secondly, the accuracy of the method is difficult to guarantee due to factors such as the vision and fatigue of human eyes (<xref ref-type="bibr" rid="B4">Dutot et&#xa0;al., 2013</xref>). Especially for large-scale planting areas such as apple plantations, the use of manual detection methods will lead to a large workload, which will not only be labor-intensive but also cause missed or false detection. Therefore, automated apple leaf disease detection has become a hotspot for research (<xref ref-type="bibr" rid="B28">Zhang and Tao, 2021</xref>) to improve the efficiency and accuracy of detection.</p>
<p>With the development of computer science and technology, machine learning algorithms have been applied to the agricultural field. For example (<xref ref-type="bibr" rid="B16">Pallathadka et&#xa0;al., 2022</xref>), preprocessed the images with histogram equalization, then applied the principal component analysis algorithm for feature extraction, and finally used the support vector machine and na&#xef;ve Bayes to classify rice leaf diseases. However, machine learning algorithms are usually constrained by the ultra-high computational effort during the stages of data preprocessing and feature extracting, making the utility of these methods generally poor (<xref ref-type="bibr" rid="B20">Sujatha et&#xa0;al., 2021</xref>). compared performances of machine learning and deep learning algorithms for plant leaf disease detection, and the experimental results show that the latter has better performance on this kind of task.</p>
<p>With the rise of convolutional neural networks and the creation of residual structures, deep learning techniques have achieved a technological leap in a very short period of time, which also improves the performances of target detection algorithms with their excellent feature extraction and model migration ability. Target detection algorithms have evolved from two-stage detection algorithms to one-stage detection algorithms in this phase. Among them, two-stage detection algorithms such as Faster R-CNN (<xref ref-type="bibr" rid="B18">Ren et&#xa0;al., 2015</xref>) and Mask-RCNN have been applied to detect plant leaf diseases by many scholars, such as (<xref ref-type="bibr" rid="B13">Li et&#xa0;al., 2023</xref>) combined Mask-RCNN with the geometric model to improve the ground-penetrating radar (GPR), and the pixel-level segmentation for localization can effectively improve the detection accuracy, reaching an average depth prediction error of 2.78 cm (<xref ref-type="bibr" rid="B3">Du et&#xa0;al., 2022</xref>). proposed a corn pest detection algorithm based on faster R-CNN, called pest R-CNN, which classified pest invasion severity into four categories, juvenile, mild, moderate, and severe based on feeding severity, and determined the severity of infestation and specific forage location. One-stage detection algorithms include YOLO (<xref ref-type="bibr" rid="B22">Tian et&#xa0;al., 2019</xref>) and SSD (<xref ref-type="bibr" rid="B10">Jiang et&#xa0;al., 2019</xref>), an end-to-end detection method that enables high-speed and real-time detection (<xref ref-type="bibr" rid="B12">LI et&#xa0;al., 2022</xref>). constructed the YOLO-JD network by introducing DSCFEM and SPPM modules, increasing the mean average precision (mAP) of jute disease detection to 96.63% (<xref ref-type="bibr" rid="B22">Tian et&#xa0;al., 2019</xref>). applied the V-space-based SSD algorithm for multiscale feature fusion and introduced an attention mechanism, increasing the mAP of apple leaf disease detection to 83.19% and the detection speed to 27.53 frames per second (FPS). However, the above models still have the problem of too large model size with a large number of parameters and high computational cost, so lightweight models have become the focus of scholars&#x2019; research in recent years (<xref ref-type="bibr" rid="B14">Lin et&#xa0;al., 2023</xref>). applied Tiny-YOLOv4 to realize strawberry real-time counting by comparing three frameworks (Darknet, TensorRT, and TensorFlow Lite) and images with different resolutions, allowing the model to achieve an accuracy of 14.6% with 91.95 FPS (<xref ref-type="bibr" rid="B19">Sozzi et&#xa0;al., 2022</xref>). proposed an improved Tiny-YOLOX model, called YOLO-Tobacco, for detecting tobacco brown spot disease in an open-air scenario. They incorporated hierarchical mixed-scale (HMU) units and convolutional block attention modules (CBAM), making the detection speed reach 69 FPS but with less accuracy compared with the non-lightweight model.</p>
<p>In summary, scholars have provided a lot of excellent ideas and methods in the field of target detection, achieving good results for plant leaf disease detection. In order to better apply to the actual situation of agricultural production, this paper proposes accurate and lightweight apple detection based on YOLOv5 (ALAD-YOLO), a lightweight network model, taking both detection accuracy and detection speed into account, which can be more easily deployed on mobiles regardless of computational resources. The main contributions are as follows:</p>
<list list-type="simple">
<list-item>
<p>(1) We replace the backbone network of YOLOv5 with a more lightweight MobilenetV3s network to reduce the parameter quantity and increase the operational efficiency.</p>
</list-item>
<list-item>
<p>(2) We use the C3 module refined by depth-wise convolution and GhostNet module (DWC3_ghost) to replace all C3 modules in the neck network, improving feature fusion efficiency, reducing computational cost, and maintaining feature expressiveness while minimizing the impact on detection accuracy.</p>
</list-item>
<list-item>
<p>(3) We develop a new module called SPPCSPC_GC, using Cross Stage Partial (CSP) structure, spatial pyramid pooling (SPP) module, and group convolution (GC) to replace the original SPP module at the interface between the backbone network and neck network, making this section better adapted to images of different resolutions, effectively avoiding overfitting, and making the model more lightweight.</p>
</list-item>
<list-item>
<p>(4) We apply the CA mechanism to improve the model&#x2019;s accuracy. Compared with the CBAM mechanism and the Squeeze-and-Excitation Networks (SE) attention mechanism, the CA mechanism can better capture object spatial and channel information without compromising model lightweighting, thereby improving detection accuracy.</p>
</list-item>
</list>
<p>The remaining parts of this article are organized as follows: Section 2 provides a detailed introduction of the dataset and network modules we used. Section 3 presents our experimental results and provides a visual analysis. Next, Section 4 compares and discusses our proposed model with current mainstream networks. Finally, we conclude by summarizing our model and discussing its potential applications.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Materials and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Data collection and preprocessing</title>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Data collection</title>
<p>To improve the generalization ability of the model, the dataset in this paper includes apple leaves with different shooting angles, backgrounds, times, disease ranges, and densities. The selected large amount of image data ensures the model&#x2019;s ability to detect small disease ranges.</p>
<p>Due to the scarcity of public datasets for apple leaf diseases, this paper collected the Kashmiri Apple Plant Disease Dataset (<xref ref-type="bibr" rid="B11">Kaur et&#xa0;al., 2022</xref>) and the public dataset for Plant Pathology 2020-FGVC7 (<xref ref-type="bibr" rid="B17">Raman et&#xa0;al., 2022</xref>).</p>
<p>However, the above datasets all have problems that do not match the actual detection environment, such as overly clear image backgrounds and mostly displaying single leaves. To enhance the generalization effect of the model and improve its ability to detect small target disease ranges, the dataset also includes 961 small target apple leaf cluster data that we collected ourselves (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1C</bold>
</xref>). The final experimental data consists of 2,748 apple disease images.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Samples of datasets, where <bold>(A)</bold> is a diseased apple leaf with a normal background, <bold>(B)</bold> is a diseased apple leaf in a real environment, and <bold>(C)</bold> is in an intensive situation where samples of apple leaves have multiple diseases collected in this paper.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g001.tif"/>
</fig>
<p>Most of the images have varying resolutions and large or small detection targets, and they were captured at different angles, providing sufficient overall diversity of the data, as shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. The details of the apple leaves are preserved while also being more in line with the actual detection environment; the background is influenced by real outdoor lighting and shadow occlusion. This reflects the actual apple disease leaf detection scene and enhances the robustness and generalization ability of the model training.</p>
<p>LabelMe software was used to generate XML files, and the images were marked as mosaic disease, spot wilt disease, and leaf blight. In the experiment, the dataset was divided into training, validation, and testing sets in an 8:1:1 ratio.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Data preprocessing</title>
<p>To improve the generalization ability of object detection models, data augmentation techniques are widely used. We used the mosaic data augmentation technique to preprocess the apple leaf disease detection dataset (<xref ref-type="bibr" rid="B21">Thapa et&#xa0;al., 2020</xref>). Mosaic data augmentation is a technique that combines multiple data augmentation operations. It combines four randomly selected images into one and then applies random transformations to the entire image, such as random scaling, flipping, translation, and color change, as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. The probability of scaling and flipping the image is 50%, while the probability of adjusting the hue, saturation, and brightness in color change is 1.5%, 70%, and 40%, respectively (<xref ref-type="bibr" rid="B25">Wang et&#xa0;al., 2021</xref>). The probability of translation is 10%.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Mosaic data augmentation. Four images are randomly cropped and stitched onto one image as training data.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g002.tif"/>
</fig>
<p>In the apple leaf disease detection dataset, the distribution of small targets is uneven, which may lead to insufficient model training. By using the mosaic data augmentation technique, we can increase the number of small targets and make their distribution more uniform, thereby improving the model&#x2019;s detection ability. In addition, mosaic data augmentation can also reduce overfitting and improve the model&#x2019;s generalization ability.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Design for ALAD-YOLO</title>
<p>In order to achieve model lightweight and ensure accurate detection of different categories of apple leaf diseases, this paper proposes an efficient and accurate detection network ALAD-YOLO based on YOLOv5s. <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref> shows the detailed structure of the ALAD-YOLO model proposed in this paper.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>The architecture of the proposed ALAD-YOLO model. The entire network is divided into four parts: input network <bold>(A)</bold>, backbone network <bold>(B)</bold>, neck network <bold>(C)</bold>, and head network <bold>(D)</bold>.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g003.tif"/>
</fig>
<p>YOLOv5 and the proposed ALAD-YOLO consist of four main components: the input layer, the backbone network, the neck network, and the prediction head. The basic unit CBS is composed of regular convolution, batch normalization (BN), and activation function SiLU. The backbone network of YOLOv5 is stacked with a large number of CBS modules and C3 modules. The Spatial Pyramid Pooling (SPP) module increases the receptive field of the feature map through three different sizes of pooling kernels, solves the problem of multiscale detection, and connects to the neck network. The prediction head achieves predictions for three different scales of objects and outputs the detection results for large, medium, and small objects.</p>
<p>Due to the large number of CBS and C3 modules in the network, the original YOLOv5 has poor portability (<xref ref-type="bibr" rid="B23">Wadhawan et&#xa0;al., 2012</xref>), making it difficult to embed in mobile devices for use in smart agriculture. Therefore, this work proposes an efficient and accurate ALAD-YOLO for detecting apple leaf diseases, as shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>. The main improvements are as follows: 1) MobileNetV3s basic blocks containing lightweight depth-wise separable convolutions, SE modules, and inverted residual structures are used instead of stacked CBS and C3 modules to improve feature extraction efficiency and compress model size. 2) The DWC3-ghost module is proposed to replace the original C3 module in the neck network to reduce parameter count and FLOPs. 3) The SPPCSPC_GC structure is proposed to replace the original SPP module with group convolution to further compress model size and improve efficiency in the feature fusion stage. 4) A lightweight coordinate attention (CA) module is embedded in the neck network to refine the key information for detecting apple leaf diseases and improve the detection accuracy for different types of diseases.</p>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Lightweight backbone network establishment</title>
<p>The backbone network of YOLOv5 mainly consists of CBS modules and C3 modules, which include convolution operations and residual structures with high parameter count and FLOPs. In order to be applied to embedded mobile devices, while ensuring the detection accuracy of the model, this paper compresses the model parameters as much as possible to improve its portability. A lightweight backbone network for ALAD-YOLO was designed using efficient MobileNetV3 (<xref ref-type="bibr" rid="B7">Howard et&#xa0;al., 2019</xref>) building blocks.</p>
<p>In this paper, the MobileNetV3s basic block is used, which can reduce the number of parameters and computations in the feature extraction process and achieve a good balance between speed and accuracy. As shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>, the MobileNetV3s basic block is mainly composed of four modules: SE module, and DW convolution module (<xref ref-type="bibr" rid="B27">Zhang et&#xa0;al., 2020</xref>), combined with an inverted residual structure and a linear bottleneck structure.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>The architecture of the MobileNetV3s Basic Block. The basic block consists of four parts: depthwise convolution, linear bottleneck and inverse residual structure, and SE attention mechanism.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g004.tif"/>
</fig>
<sec id="s2_2_1_1">
<label>2.2.1.1</label>
<title>Depth-wise separable convolutions</title>
<p>Depth-wise separable convolution decomposes a standard convolution into a concatenation of two layers. First, it uses a lightweight depth-wise convolution layer (<inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>) to apply a single-channel convolution filter to each channel of the input feature map. Then, it connects a pointwise convolution layer (<inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>), which generates a new feature map by linearly combining the features of each channel through a pointwise convolution operation.</p>
<p>The ratio of the total number of parameters between depth-wise separable convolution and normal convolution can be calculated as follows:</p>
<disp-formula>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>Y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>*</mml:mo>
<mml:mi>f</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The formula for traditional convolutional operation is shown in equation n, where * represents convolution operation. The output feature map dimension is <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, where N is the number of output channels, and the convolutional kernel dimension is <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, where M is the number of input feature map channels. Therefore, the calculation of FLOPs for normal convolution is <inline-formula>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Depth-wise separable convolution can be divided into two stages: depth-wise convolution and point-wise convolution. In the first stage, the input feature dimension of the depth-wise convolution is <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, the convolution kernel parameter is <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and each channel corresponds to only one convolution kernel during convolution. The output dimension is <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, so the FLOPs of the depth-wise convolution is <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. In the second stage, the point-wise convolution has a convolution kernel parameter of <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and performs a <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> standard convolution on each feature, with an output dimension of <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. The FLOP calculation is <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. Therefore, the parameter and FLOP ratio of the normal convolution and the depth-wise separable convolution can be obtained as</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>r</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>+</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msubsup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>r</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>O</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msubsup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where M is the number of input channels, <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the size of the input feature map, <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the size of the convolutional kernel, and N is the number of output channels.</p>
<p>The reduction of the parameter and <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>O</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> in depth-wise separable convolution can be approximately attributed to the parameter <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:msubsup>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>. When choosing <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, compared with the traditional convolution, depth-wise separable convolution reduces the parameter and <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>O</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> to <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">/</mml:mo>
<mml:mn>8</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">/</mml:mo>
<mml:mn>9</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> of the regular convolution while ensuring a slight drop in accuracy.</p>
</sec>
<sec id="s2_2_1_2">
<label>2.2.1.2</label>
<title>Linear bottleneck and reverse residual structure</title>
<p>The reverse residual structure is as follows: The 1 &#xd7; 1 convolution, also known as expansion convolution, expands the number of channels in the input feature map by a factor of &#x201c;factor&#x201d;, mapping the low-dimensional space to a high-dimensional space. Then, the 3 &#xd7; 3 depth-wise separable convolution is used to greatly reduce the network&#x2019;s parameter and computational complexity while ensuring the feature extraction capability and prediction accuracy of regular convolutions. Finally, the 1 &#xd7; 1 convolution, combined with a linear activation function (Linear Bottleneck structure), is used to restore the number of output feature map channels to 1/factor. The use of a linear activation function can effectively reduce information loss during the transformation process from the high-dimensional to low-dimensional feature space. The input and output are only connected through the residual structure when they have the same number of channels. This structure reflects the compactness of the input and output, implements an internal nonlinear transformation to expand the features to a higher dimension, retains all the necessary information in the bottleneck, increases the feature expression ability, and reduces the network&#x2019;s parameter count.</p>
</sec>
<sec id="s2_2_1_3">
<label>2.2.1.3</label>
<title>SE module</title>
<p>Firstly, the channels of the input feature matrix are pooled to obtain an <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mtext>channel</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> -dimensional vector. Then, two fully connected layers are connected, with neuron numbers of channel and channel*1/4, respectively, and ReLu and h-swish activation functions are applied respectively. The SE module increases the weight of important parts of the results and reduces the weight of ineffective or less effective parts, thus training a better model.</p>
<p>In summary, compared with the CBS and C3 modules, the MobileNetV3s basic block has three advantages: 1) introducing depth-wise separable convolutions with lower computational cost to replace ordinary convolutions, thus reducing the number of parameters while maintaining network performance; 2) using linear bottlenecks and inverted residual structures to make more efficient layer structures by utilizing the low-rank properties of the problem; 3) using SE attention mechanism modules to increase the network&#x2019;s sensitivity to effective information.</p>
</sec>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>DWC3-ghost</title>
<p>To further reduce model parameters and computational complexity for apple leaf detection on embedded devices, this paper proposes a new DWC3-ghost module using DWConv and GhostNet (<xref ref-type="bibr" rid="B5">Han et al., 2020</xref>) basic blocks, which is embedded in the neck network to replace the original C3 module, improving the efficiency of feature fusion in the model.</p>
<p>As shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>, in the ghost basic block, the input is first processed by a normal convolution (convolution + BN + activation function) to generate an intrinsic feature map with fewer channels, reducing the number of parameters. Then, the identity and inexpensive linear operation <inline-formula>
<mml:math display="inline" id="im22">
<mml:mi>&#x3c6;</mml:mi>
</mml:math>
</inline-formula> (only convolution) are used to enhance the features, generating the complete feature map.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Ghost basic block computation: <bold>(A)</bold> general convolution with reduced number of channels. <bold>(B)</bold> Lightweight cheap linear transformation. <bold>(C)</bold> Stacking of the results of <bold>(B)</bold> computation.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g005.tif"/>
</fig>
<p>In addition, lightweight operations such as low-channel convolution and inexpensive linear operations are used in the ghost module instead of traditional convolutional layers to generate redundant features, greatly reducing computational costs and achieving more efficient feature mapping and computation.</p>
<p>The FLOPs in the ghost module consist of two parts: the traditional convolutional layer that outputs a small number of feature maps and the lightweight and inexpensive linear transformation layer. The calculation process of the ghost module is shown in formula n:</p>
<disp-formula>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mo>`</mml:mo>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>X</mml:mi>
</mml:mstyle>
<mml:mo>*</mml:mo>
<mml:msup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>f</mml:mi>
</mml:mstyle>
<mml:mo>`</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>&#x3c6;</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>j</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>i</mml:mi>
</mml:mstyle>
<mml:mo>`</mml:mo>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>j</mml:mi>
</mml:mstyle>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mo>`</mml:mo>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>Y</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The first stage outputs a small number of feature maps. The input feature map <inline-formula>
<mml:math display="inline" id="im23">
<mml:mi>X</mml:mi>
</mml:math>
</inline-formula> has dimensions of <inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and the convolutional kernel has dimensions of <inline-formula>
<mml:math display="inline" id="im25">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. The&#xa0;output feature map <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:msup>
<mml:mi>Y</mml:mi>
<mml:mo>`</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> has dimensions of <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. To simplify the calculation, the bias in the convolution calculation is omitted and the <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>O</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> are calculated as <inline-formula>
<mml:math display="inline" id="im29">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>. The second stage represents a lightweight and inexpensive linear transformation layer. <inline-formula>
<mml:math display="inline" id="im30">
<mml:mrow>
<mml:msubsup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>`</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> represents the <inline-formula>
<mml:math display="inline" id="im31">
<mml:mi>i</mml:mi>
</mml:math>
</inline-formula> th feature map output from the first stage, <inline-formula>
<mml:math display="inline" id="im32">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c6;</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the <inline-formula>
<mml:math display="inline" id="im33">
<mml:mi>j</mml:mi>
</mml:math>
</inline-formula> th ghost feature map generated by the <inline-formula>
<mml:math display="inline" id="im34">
<mml:mi>j</mml:mi>
</mml:math>
</inline-formula> th linear operation, and <inline-formula>
<mml:math display="inline" id="im35">
<mml:mrow>
<mml:msubsup>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>`</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> generates <inline-formula>
<mml:math display="inline" id="im36">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> ghost feature maps. Finally, the ghost feature map <inline-formula>
<mml:math display="inline" id="im37">
<mml:mrow>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> has dimensions of <inline-formula>
<mml:math display="inline" id="im38">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>F</mml:mi>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The FLOP ratio of traditional convolutional layers and Ghost modules with the same number of channels and the same feature map size is shown in formula n. The Ghost module reduces the model FLOPs by around s times.</p>
<disp-formula>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>r</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mfrac>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>N</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>K</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>M</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>D</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
</mml:mstyle>
</mml:msub>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>d</mml:mi>
</mml:mstyle>
<mml:mo>&#xd7;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>d</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2248;</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>s</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where s represents the total mappings produced by each channel (1 intrinsic feature map and s-1 ghost feature maps) and d represents the average size of the convolutional kernel for linear operations.</p>
<p>Due to its simplicity and efficiency, the proposed GhostNet block was used to replace the Bottleneck in the original C3 convolution for lightweight implementation, as shown in <xref ref-type="fig" rid="f6">
<bold>Figures&#xa0;6A, B</bold>
</xref>. The CBS module was replaced by DWConv to design an efficient DWC3-ghost module, as shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6C</bold>
</xref>. The proposed DWC3-ghost module was embedded in the neck network to improve feature fusion efficiency. This can reduce the computational cost of the neck network while maintaining the expressive power of the features, thereby reducing the impact on detection accuracy.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>
<bold>(A)</bold> Basic Ghost module. <bold>(B)</bold> Ghost-Bottleneck. <bold>(C)</bold> DWC3-Ghost.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g006.tif"/>
</fig>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>SPPCSCP_GC module</title>
<p>To make the model more lightweight and improve its real-time performance, this paper proposes using an improved SPPCSCP module (<xref ref-type="bibr" rid="B24">Wang et&#xa0;al., 2022</xref>). The SPPCSCP_GC replaces the original SPP module to further compress the model size, reduce the number of model parameters, and improve the efficiency of the model in the feature fusion stage.</p>
<p>The SPP module uses Maxpool on the feature map input into four branches with different scales, giving the model the ability to adapt to images of different resolutions. This effectively avoids image distortion caused by cropping or scaling operations on image regions and improves the scale-invariance of the image while effectively avoiding overfitting.</p>
<p>The CSP module halves the number of channels in the feature map and splits it into two branches, with one branch going through convolution processing and the other going through Bottleneck * N operations. The two branches are then concatenated, ensuring both accuracy and reduced computational cost.</p>
<p>The proposed SPPCSPC module combines the SPP and CSP modules, as shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>, with CSP as the main component. It retains the branch that performs a single convolution and adds an SPP module to the other branch. Finally, the two branches are concatenated to integrate all features, enhancing the effect of feature fusion while ensuring model lightweight.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>
<bold>(A)</bold> Basic SPP module, <bold>(B)</bold> CSP structure, effectively reducing the number of model calculation parameters.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g007.tif"/>
</fig>
<p>To further improve the model&#x2019;s computational speed, group convolution (<xref ref-type="bibr" rid="B29">Zhang et al., 2018</xref>) is applied to the SPPCSPC module in this model. Group convolution aims to process feature maps in groups, with each convolution kernel divided into groups and convolved within the corresponding group, as shown in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8</bold>
</xref>. The resulting feature maps are then concatenated together. Group convolution can increase the diagonal correlation between feature maps and significantly reduce training parameters, making it less prone to overfitting, similar to the effect of regularization.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>Group convolution operation procedure. The feature maps are processed in group, and each convolution kernel is divided into group accordingly.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g008.tif"/>
</fig>
<p>Therefore, we propose an effective improved module SPPCSCP_GC, which combines SPP, CSP, and GC modules, so that the network structure can not only perform well under images of different resolutions but also effectively alleviate the problem of gradient vanishing. The module reduces the size of the model, ensuring inference speed and accuracy with fewer parameters and smaller FLOP values.</p>
</sec>
<sec id="s2_2_4">
<label>2.2.4</label>
<title>Coordinate attention module</title>
<p>Improvements in the lightweighting of the neck network can make the model more lightweight and significantly improve its real-time performance, making it more convenient to deploy on mobile devices. However, these improvements inevitably lead to a decrease in detection accuracy. Therefore, we need to optimize the model&#x2019;s performance on detection accuracy while not significantly affecting the computational cost.</p>
<p>To sum up, we introduce the Coordinate Attention (CA) mechanism (<xref ref-type="bibr" rid="B6">Hou et&#xa0;al., 2021</xref>), as shown in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9</bold>
</xref>, which effectively addresses the issue of the SE attention mechanism&#x2019;s focus only on building inter-channel dependencies and ignoring spatial features (<xref ref-type="bibr" rid="B8">Hu et&#xa0;al., 2018</xref>), as well as the CBAM attention mechanism&#x2019;s introduction of large-scale convolution kernels for extracting spatial features but ignoring long-range dependencies (<xref ref-type="bibr" rid="B26">Woo et&#xa0;al., 2018</xref>). The CA module is flexible and lightweight enough to be easily incorporated into the core module of lightweight networks.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>CA attention mechanism. Firstly, pool the input feature maps in the X(H) and Y(W) directions, respectively. Then, concatenate the output maps and use 1*1 convolution to reduce the dimension. Lastly, along the spatial dimension, perform the dimension raising operation and combine it with the sigmoid function to get the attention vector.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g009.tif"/>
</fig>
<p>As shown in the above figure, the CA attention mechanism first performs pooling along the X and Y directions on the input feature map with size <inline-formula>
<mml:math display="inline" id="im39">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, generating feature maps of size <inline-formula>
<mml:math display="inline" id="im40">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im41">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, respectively. Then, the two feature maps are concatenated to obtain a feature map of size <inline-formula>
<mml:math display="inline" id="im42">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, which is then compressed from the C dimension to the <inline-formula>
<mml:math display="inline" id="im43">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> dimension using a <inline-formula>
<mml:math display="inline" id="im44">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> Conv, resulting in a feature map of size <inline-formula>
<mml:math display="inline" id="im45">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>*</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. Next, the h_wish function is used for non-linear activation, obtaining the intermediate feature representing the encoded information. The intermediate feature is then decomposed into a vertical attention tensor of size <inline-formula>
<mml:math display="inline" id="im46">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and a horizontal attention tensor of size <inline-formula>
<mml:math display="inline" id="im47">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, and each attention tensor is upsampled using a set of 1&#xd7;1 Conv, increasing the number of channels from <inline-formula>
<mml:math display="inline" id="im48">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> to C. Finally, the sigmoid function is used for non-linear activation to produce the corresponding attention weights. The attention weights in the two directions are multiplied with the original feature map from the shortcut to obtain the attention-enhanced feature map.</p>
<p>The CA attention mechanism enables the network to collect information from a wider area rather than being biased toward a specific region, thus significantly improving detection accuracy. In addition, the CA module uses only <inline-formula>
<mml:math display="inline" id="im49">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> Conv kernels, two average pooling layers, and very few matrix transpositions, reducing the number of parameters and computational costs.</p>
<p>In this paper&#x2019;s model, the CA module is embedded behind the intersection of all top-down and bottom-up information fusion in the neck network, effectively improving the detection accuracy loss caused by model compression. This allows the network to have attention to key information without incurring high computational costs, thus maintaining its efficiency and flexibility.</p>
</sec>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results</title>
<p>In this section, the experimental setup, hyperparameter settings, and training strategies are detailed in Section 3.1. Then, Section 3.2 describes the evaluation metrics used to evaluate the performance of the model and their calculation formulas. Finally, in Section 3.3, the results of this article are explored in combination with an ablation experiment and visual analysis.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Implementation and setup</title>
<p>Experiments were conducted on the Featurize cloud server, with an Nvidia RTX 3090 graphics card configured for hardware, with 25.4 GB of graphics memory, and deployed on a Linux operating system. The implementation of the proposed method is based on Python 1.10.1.</p>
<p>The model training strategies are as follows. We set a dynamic learning rate to accelerate the network to the optimal value. The initial learning rate(lr0) was set to 0.1, and the OneCycleLR learning rate (lrf) is set to 0.01. The learning rate was updated every epoch until the final learning rate reached to lr0*lrf. The warmup_epochs was set to 3.0, with the warmup initial momentum set to 0.8 and the warmup initial bias lr set to 0.1. The training epoch was set to 500, and the batch size was set to 16. The weight attenuation was set to 0.0005, and the momentum was set to 0.937. The SGD optimizer is used to optimize the network parameters.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Evaluation metrics</title>
<p>To evaluate the detection precision of the proposed model ALAD-YOLO, we selected accuracy (P), recall rate (R), and average accuracy (mAP) as evaluation metrics. Among them, accuracy represents the ratio of positive samples correctly predicted to positive samples predicted by the model. It mainly measures the accuracy of the network in identifying positive samples. The recall rate represents the proportion of correctly identified positive samples to all positive samples, and it reflects the ability of the model in finding positive samples. mAP refers to the mean value of the average accuracies of all categories, which combines the detection performance between different categories. We calculated these evaluation metrics according to the following formulas:</p>
<disp-formula>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo>+</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>R</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo>+</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:munderover>
<mml:mo>&#x222b;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mn>1</mml:mn>
</mml:munderover>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>R</mml:mi>
</mml:mstyle>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>d</mml:mi>
<mml:mi>R</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>m</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>i</mml:mi>
</mml:mstyle>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>n</mml:mi>
</mml:mstyle>
</mml:msubsup>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>A</mml:mi>
</mml:mstyle>
<mml:msub>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>P</mml:mi>
</mml:mstyle>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>i</mml:mi>
</mml:mstyle>
</mml:msub>
</mml:mrow>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>n</mml:mi>
</mml:mstyle>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where TP represents the number of positive samples correctly classified. TN is the number of correctly identified negative examples. FP represents the number of negative examples incorrectly classified as positive examples. FN represents the number of positive samples incorrectly classified as negative samples.</p>
<p>In addition to the above evaluation metrics to evaluate the detection performance of the model, we also use the number of parameters and FLOPs to evaluate the size and computational cost of the ALAD-YOLO model to select lightweight network to deploy on the mobile devices. Fewer parameters and FLOPs mean that under the same computing resources, the model can run more efficiently, while reducing memory usage and improving computing speed.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Ablation experiment and analysis</title>
<p>Ablation experiments are conducted to investigate the contribution of various modules in ALAD-YOLO, which improve the detection performance and reduce computational costs. The YOLOv5s model was used as the benchmark model. First, a lightweight network architecture is introduced to verify its impact on detection performance. Then, based on the lightweight network architecture selected, the improvement of network accuracy is verified by introducing modules, such as DWC3_Ghost module and CA module.</p>
<p>Firstly, different lightweight networks to improve the network performance are verified on the test set. The benchmark model YOLOv5s was compared with improved models, such as YOLOv5s GhostNet (Experiment 2) and YOLOv5s MobileNet (Experiment 3). Compared with the benchmark model YOLOv5s, the mAP50-95 in experiment 2 and experiment 3 decreased by 1.6% and 1.7%, respectively. However, in terms of parameter quantity, the FLOPs in experiment 2 and experiment 3 increased greatly by 49.4% and 54.4%, respectively. Specifically, the quantity decreased from 7 million to 3 million and 4 million. The results show that the backbone network composed of MobileNet basic blocks can significantly reduce the computing cost and the accuracy of mAP50-95 is only 0.1% lower than that of the Ghost module. Considering multiple factors, the experiment finally selects the backbone network composed of MobileNet basic blocks, which can significantly reduce the computing cost, and its impact on network accuracy is also acceptable.</p>
<p>The lightweight network structure can significantly reduce the size of the model and improve detection speed, with the cost of reducing the detection accuracy of the network. Therefore, some improved methods that can improve accuracy without introducing high computational costs are essential.</p>
<p>Secondly, based on the lightweight YOLO network composed of MobileNet basic blocks, the model performance changes caused by introducing different modules are verified. Experiment 6 introduces the SPPCSPC module to replace the original SPP module, better extracting and fusing feature maps. Experiment 7 uses the SPPCSPC_ GC module, which replaces the ordinary convolution in the SPPCSPC structure with group convolutions, effective feature fusion is ensured while achieving lightweight. Experiment 8 and experiment 9 respectively replace the C3 module in the neck part of the original network with DWC3 Host and DWC3 Faster structures. Based on experiment 9, experiment 10 and experiment 11 introduce an attention module to effectively extract key information about the detection results in the network. The improvements of the lightweight YOLO model on detection performance by introducing different improvement methods are shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>.</p>
<list list-type="order">
<list-item>
<p>Experiment 6 and Experiment 7 show that using the SPPCSPC module and SPPCSPC_GC to replace the SPP module in the original network with the GC module can effectively improve the model detection accuracy, with improvements on the mAP50-95 by 3.0% and 6.4% on the test set, respectively. The addition of the SPPCSPC module in experiment 6 resulted in a significant increase of 5.2 G in the FLOPs of the model, which was somewhat outweighed by the 3% mAP increase. However, in experiment 7, after replacing the convolutions in the SPPCSPC module with group convolutions, SPPCSPC_GC was proposed. The mAP50-95 on the test set is increased by 6.4% with only a 0.9 G increase in FLOPs, significantly improving detection accuracy compared with the original benchmark network and meeting the lightweight requirements.</p>
</list-item>
<list-item>
<p>Experiments 8 and 9 show that using the DWC3-Ghost module and DWC3-Faster module instead of the C3 module in the original network neck section can effectively reduce parameter quantity and computational costs and has a certain improvement in detection accuracy. Compared with the benchmark network, experiment 8 and experiment 9 reduced FLOPs by 62.0% and 60.8%, respectively, and improved mAP50-95 on the test set by 0.1% and 0.5%, respectively. The results show that replacing the C3 module with a lightweight structure DWC3-Ghost can effectively compress the model size and computational cost and efficiently fuse and extract features to improve detection accuracy.</p>
</list-item>
<list-item>
<p>Experiments 9 and 11 show that adding a CA module to the lightweight YOLO network can effectively improve the detection accuracy of the network, increasing the mAP50-95 on the test set by 2.7%, while reducing the FLOPs by 0.1 G. Compared with the CBAM module added in experiment 10, its performance in the test set increased by 0.3% and the FLOPs also decreased by 0.1 G. Therefore, both have a slight increase in parameter quantity. The results show that compared with the CBAM module, embedding the CA module can better highlight information that is helpful for disease spot detection. Although it slightly increases the parameters of the model, it can well suppress useless information to improve the accuracy of the model, while reducing the FLOPs of the model.</p>
</list-item>
</list>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Comparison of experimental results based on the lightweight YOLOv5s model with different modules.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center"/>
<th valign="middle" align="left">Model</th>
<th valign="middle" align="center">mAP50-95 (%)</th>
<th valign="middle" align="center">mAP50 (%)</th>
<th valign="middle" align="center">Parameters</th>
<th valign="middle" align="center">FLOPs(G)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">1</td>
<td valign="middle" align="left">YOLOv5s</td>
<td valign="middle" align="center">82.3</td>
<td valign="middle" align="center">97.7</td>
<td valign="middle" align="center">7,018,216</td>
<td valign="middle" align="center">15.8</td>
</tr>
<tr>
<td valign="middle" align="center">2</td>
<td valign="middle" align="left">YOLOv5s-GhostNet</td>
<td valign="middle" align="center">80.7</td>
<td valign="middle" align="center">97.0</td>
<td valign="middle" align="center">3,681,120</td>
<td valign="middle" align="center">8.0</td>
</tr>
<tr>
<td valign="middle" align="center">3</td>
<td valign="middle" align="left">YOLOv5s-MobileNet</td>
<td valign="middle" align="center">80.6</td>
<td valign="middle" align="center">97.6</td>
<td valign="middle" align="center">4,632,840</td>
<td valign="middle" align="center">7.2</td>
</tr>
<tr>
<td valign="middle" align="center">6</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC</td>
<td valign="middle" align="center">83.6</td>
<td valign="middle" align="center">98.2</td>
<td valign="middle" align="center">11,031,144</td>
<td valign="middle" align="center">12.4</td>
</tr>
<tr>
<td valign="middle" align="center">7</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC_GC</td>
<td valign="middle" align="center">87.0</td>
<td valign="middle" align="center">98.4</td>
<td valign="middle" align="center">5,665,768</td>
<td valign="middle" align="center">8.1</td>
</tr>
<tr>
<td valign="middle" align="center">8</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC_GC-DWC3_Faster</td>
<td valign="middle" align="center">87.1</td>
<td valign="middle" align="center">98.7</td>
<td valign="middle" align="center">4,640,616</td>
<td valign="middle" align="center">6.0</td>
</tr>
<tr>
<td valign="middle" align="center">9</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC_GC-DWC3_Ghost</td>
<td valign="middle" align="center">87.5</td>
<td valign="middle" align="center">98.3</td>
<td valign="middle" align="center">4,703,480</td>
<td valign="middle" align="center">6.2</td>
</tr>
<tr>
<td valign="middle" align="center">10</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC_GC-DWC3_Ghost-CBAM</td>
<td valign="middle" align="center">87.8</td>
<td valign="middle" align="center">98.3</td>
<td valign="middle" align="center">4,719,835</td>
<td valign="middle" align="center">6.1</td>
</tr>
<tr>
<td valign="middle" align="center">11</td>
<td valign="middle" align="left">YOLOv5s-MobileNet-SPPCSPC_GC-DWC3_Ghost-CA</td>
<td valign="middle" align="center">90.2</td>
<td valign="middle" align="center">98.7</td>
<td valign="middle" align="center">4,712,323</td>
<td valign="middle" align="center">6.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, the proposed ALAD-YOLO model reduces the parameter quantity and FLOPs, significantly improving the speed of disease spot detection, whereas the impact on detection accuracy can be negligible. Therefore, the proposed ALAD-YOLO is more suitable for deployment on mobile devices with constrained resource and has the great detection performance required for practical applications.</p>
<p>On the test set with 2,748 images, the ALAD-YOLO network accurately identified three apple leaf diseases and healthy leaves, with mAP50 reaching 98.7% and mAP50-95 reaching 90.2%. The detection performance of each category is shown in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>The detection performance of the ALAD-YOLO model on different categories of apple diseased leaves on the test set.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Class</th>
<th valign="middle" align="center">Instances</th>
<th valign="middle" align="center">P</th>
<th valign="middle" align="center">R</th>
<th valign="middle" align="center">mAP50</th>
<th valign="middle" align="center">mAP50-95</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">All</td>
<td valign="middle" align="center">3364</td>
<td valign="middle" align="center">0.964</td>
<td valign="middle" align="center">0.972</td>
<td valign="middle" align="center">0.987</td>
<td valign="middle" align="center">0.902</td>
</tr>
<tr>
<td valign="middle" align="center">Mosaic</td>
<td valign="middle" align="center">1788</td>
<td valign="middle" align="center">0.961</td>
<td valign="middle" align="center">0.971</td>
<td valign="middle" align="center">0.985</td>
<td valign="middle" align="center">0.894</td>
</tr>
<tr>
<td valign="middle" align="center">Spot wilt</td>
<td valign="middle" align="center">624</td>
<td valign="middle" align="center">0.973</td>
<td valign="middle" align="center">0.971</td>
<td valign="middle" align="center">0.992</td>
<td valign="middle" align="center">0.918</td>
</tr>
<tr>
<td valign="middle" align="center">Leaf blight</td>
<td valign="middle" align="center">952</td>
<td valign="middle" align="center">0.958</td>
<td valign="middle" align="center">0.974</td>
<td valign="middle" align="center">0.985</td>
<td valign="middle" align="center">0.892</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To visualize the detection performance of the proposed method, <xref ref-type="fig" rid="f10">
<bold>Figures&#xa0;10</bold>
</xref>, <xref ref-type="fig" rid="f11">
<bold>11</bold>
</xref> provide detection results of apple diseased leaf images in different scenarios. <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref> visualizes the detection ability of the model in simple scenarios, where the simple environment refers to situations where the shooting is clear, the background is relatively simple and clear, the affected area of the apple is obvious, and the number of apple leaves included in the image is relatively small. From <xref ref-type="fig" rid="f10">
<bold>Figures&#xa0;10A&#x2013;C</bold>
</xref>, it can be seen that our model can detect and judge three kinds of apple leaf disease accurately and simultaneously. Due to the proposed CA attention module, our model has a good ability to extract key information from images, resulting in high detection accuracy for different categories of apple leaf diseases. <xref ref-type="fig" rid="f12">
<bold>Figure&#xa0;12</bold>
</xref> indicates the detection effect in difficult cases, where the leaves are at the edge of the figure or partially obscured. From <xref ref-type="fig" rid="f12">
<bold>Figures&#xa0;12A, B, D&#x2013;F</bold>
</xref>, it can be seen that our model can also accurately detect and judge the blades located in the edge region of the image. Also, <xref ref-type="fig" rid="f12">
<bold>Figures&#xa0;12C, G, I</bold>
</xref> indicate that our model can also accurately detect the leaves, which are partially obscured by others. At the same time, it can also grasp edge information in the image well.</p>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>
<bold>(A&#x2013;I)</bold> indicate the detection effect with three samples randomly selected under three categories in the simple condition, i.e., a figure containing only one leaf to be detected. The bounding box shows the predicted label of each detected leaf and the confidence level of the prediction.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g010.tif"/>
</fig>
<fig id="f11" position="float">
<label>Figure&#xa0;11</label>
<caption>
<p>
<bold>(A&#x2013;E)</bold> indicate the complex cases, i.e., a figure containing multiple diseased leaves to be detected, and shading between the leaves is also common due to the high density.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g011.tif"/>
</fig>
<fig id="f12" position="float">
<label>Figure&#xa0;12</label>
<caption>
<p>
<bold>(A&#x2013;I)</bold> indicate the detection effect in difficult cases, i.e., the leaves are at the edge of the figure or partially obscured. <bold>(A, B, D&#x2013;F)</bold> show that ALAD-YOLO captures the edge information well. Also, <bold>(C, G, I)</bold> indicate that ALAD-YOLO has good detection ability for partially obscured diseased leaves.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1204569-g012.tif"/>
</fig>
<p>To better demonstrate the advantages of the proposed model, <xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref> visualizes the detection capabilities of the model in complex scenarios. Due to the fact that most apple leaves are dense and have small spots in real environments, the detection of diseased leaves in dense environments is particularly important. However, our model is still very effective in this area. In the case of densely distributed apple leaves, the environment in which each apple leaf is located may not be the most suitable for testing. From <xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref>, it can be seen that areas with apple leaf diseases can be effectively detected, whether they are shaded areas, strong light exposure areas, shooting edge areas, or leaf stacking areas.</p>
<p>Our conclusion is that ALAD-YOLO can achieve the maximum mAP50-95 of 90.2% on the test set, which is 7.9% higher than the benchmark model YOLOv5s, while maintaining the minimum level of parameter quantity and FLOPs. Compared with other models, our model has the best recognition accuracy and achieves the fastest calculating speed and the smallest model size, meeting the requirements of real-time object detection for embedded mobile devices.</p>
</sec>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>Based on the YOLO-V5 architecture, we updated the original backbone network with Mobilenet-V3 and introduced several effective modules, proposing ALAD-YOLO to achieve both good detection accuracy and impressive speed. To further verify the model performance, we conducted many comparative experiments: compared with the regular YOLO-V5s model, ALAD-YOLO achieved a 7.9% improvement in accuracy while reducing FLOPs by 9.7G; for lightweight improvements such as YOLOv5s-GhostNet, YOLOv5s-MobileNet, and YOLOv5s-ShuffleNet, ALAD-YOLO showed a 10% or so increase in accuracy and 2&#x2013;3-G improvement in speed. For networks with SPP, CA, CBAM, and other modules added, such as YOLOv5s-MobileNet-SPPCSPC and YOLOv5s-MobileNet-SPPCSPC_GC-DWC3_Faster, ALAD-YOLO was able to improve accuracy by around 3% while maintaining speed, achieving a balance between accuracy and lightweight and making it more suitable for completing tasks than other models.</p>
<p>In this study, we found some remaining issues in the model, which are common problems in current object detection models. In small object detection, such as detecting too many leaves, small spots, or unavoidable occlusion due to lighting, ALAD-YOLO may still have some missed detection or false detection (<xref ref-type="bibr" rid="B1">Alonso et&#xa0;al., 2020</xref>). We believe that more powerful and comprehensive data augmentation algorithms, such as random masking and noise introduction, can help the model learn more subtle features. Alternatively, taking pictures of leaves from multiple angles may also solve such problems.</p>
<p>We also studied the performance of ALAD-YOLO on mobile devices. In actual farms, data is usually collected through sensors, processed by edge computing devices, and then transmitted to cloud servers for more in-depth data analysis. Obviously, relying solely on smartphone-based applications cannot achieve round-the-clock disease detection on the farm. With the rapid development of artificial intelligence technology, suitable algorithms have been applied to the artificial intelligence of things (AIoT) (<xref ref-type="bibr" rid="B2">Chen et&#xa0;al., 2020</xref>). Our proposed ALAD-YOLO can be integrated into such a real-time observation system that provides farmers with environmental changes, so that disease can be judged more accurately and quickly. In addition, we also plan to compare our model with other SOTA methods on the Raspberry Pi platform to further evaluate its performance and provide more reference for future research.</p>
<p>At the same time, we hope that some high-performance edge computing modules or lightweight AI supercomputers can be applied to the field of agriculture. With such computing resources, the model will perform better than mobile platforms based on CPUs, and the corresponding models and algorithms will also have wider applications.</p>
</sec>
<sec id="s5" sec-type="conclusion">
<label>5</label>
<title>Conclusion</title>
<p>This paper proposes a lightweight apple leaf disease detection network, called ALAD-YOLO, to address the challenge of balancing accuracy and speed in the current detection of apple leaf diseases. Multiple data augmentation techniques are employed to enhance the apple leaf disease detection dataset for training and evaluation. ALAD-YOLO is an improved version of YOLO-V5s, with a more lightweight Mobilenet-V3 network as its backbone. This modification reduces the computational cost of feature extraction while ensuring accuracy. The proposed DWC3-ghost module is applied to the neck of the network, which improves the efficiency of feature fusion while maintaining its expressiveness. Moreover, the application of the SPPCSPC_GC module further enhances the model&#x2019;s performance under different input resolutions. The&#xa0;introduction of the CA attention mechanism strengthens the model&#x2019;s focus on the target, effectively compensating for the accuracy loss caused by previous lightweight operations. Experimental results show that ALAD-YOLO achieves a detection accuracy of 90.2% with 6.1 GFLOPs. Compared with existing models, ALAD-YOLO not only performs better in terms of accuracy but also has higher computational efficiency. Therefore, the proposed method provides excellent technical support for the real-time and accurate detection of apple leaf diseases. In the subsequent research, we will further optimize the performance of ALAD-YOLO in complex scenarios, so that it can have a wider range of applications.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>WX and RW conceived and proposed the main approach. WX contributes to the collection and collation of data. WX conducted the experiment, and WX and RW worked together to analyze and improve the results. RW and WX wrote, revised, and reviewed the paper together. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Alonso</surname> <given-names>R. S.</given-names>
</name>
<name>
<surname>Sitton-Candanedo</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Casado-Vara</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Prieto</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Corchado</surname> <given-names>J.M.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Deep reinforcement learning for the management of software-defined networks in smart farming</article-title>,&#x201d; in <conf-name>Proceedings of the 2nd IEEE International Conference on Omni-Iayer Intelligent Systems</conf-name>. (ELECTR NETWORK: IEEE), <fpage>135</fpage>&#x2013;<lpage>140</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>C. J.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.S.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>C.Y.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>An AIoT based smart agricultural system for pests detection</article-title>. <source>IEEE Access.</source> <volume>8</volume>, <fpage>180750</fpage>&#x2013;<lpage>180761</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2020.3024891</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>Y. Q.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>J. D.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Y. D.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Z. G.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>A novel object detection model based on faster r-CNN for spodoptera frugiperda according to feeding trace of corn leaves</article-title>. <source>AGRICULTURE-BASEL</source> <volume>12</volume>, <elocation-id>2</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/agriculture12020248</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dutot</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>L.M.</given-names>
</name>
<name>
<surname>Tyson</surname> <given-names>R.C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Predicting the spread of postharvest disease in stored fruit, with application to apples</article-title>. <source>Postharvest. Technol.</source> <volume>85</volume>, <fpage>45</fpage>&#x2013;<lpage>56</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.postharvbio.2013.04.003</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Han</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>C. J.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.Y.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>GhostNet: more features from cheap operations</article-title>,&#x201d; in <conf-name>Proceedings of the 2020 IEEE/CVF conference on computer vision and pattern recognition (CVPR)</conf-name>. <fpage>1577</fpage>&#x2013;<lpage>1568</lpage> (ELECTR NETWORK: IEEE). doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00165</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hou</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>J. S.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Coordinate attention for efficient mobile network design</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. (ELECTR NETWORK: IEEE), <fpage>13708</fpage>&#x2013;<lpage>13717</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01350</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Howard</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sandler</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chu</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>)<article-title>Searching for MobileNetV3</article-title> (<publisher-name>IEEE</publisher-name>) (Accessed <access-date>Proceedings of the 2019 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION</access-date>).</citation>
</ref>
<ref id="B8">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Albanie</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Squeeze-and-Excitation networks</article-title>,&#x201d; in <conf-name>Proceedings of the 31st IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <fpage>7132</fpage>&#x2013;<lpage>7141</lpage> (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2018.00745</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>L. Y.</given-names>
</name>
<name>
<surname>Yue</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J. Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y. T. S.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>X. Q.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>K.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Overexpression of MdMIPS1 enhances drought tolerance and water-use efficiency in apple</article-title>. <source>J. Integr. Agricul.</source> <volume>21</volume>, <fpage>7</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/S2095-3119(21)63822-4</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>He</surname> <given-names>D.J.</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>C.Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Real-time detection of apple leaf diseases using deep learning approach based on improved convolutional neural networks</article-title>. <source>IEEE Access</source> <volume>7</volume>, <fpage>59069</fpage>&#x2013;<lpage>59080</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2019.2914929</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Kaur</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Devendran</surname>
</name>
<name>
<surname>Verma</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Kavita</surname>
</name>
<name>
<surname>Jhanjhi</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>De-noising diseased plant leaf image</article-title> (<publisher-name>ICCIT</publisher-name>) (Accessed <access-date>Proceedings of the 2022 2nd International Conference on Computing and Information Technology</access-date>).</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LI</surname> <given-names>D. W.</given-names>
</name>
<name>
<surname>Ahmed</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>N. L.</given-names>
</name>
<name>
<surname>Sethi</surname> <given-names>A.I.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>YOLO-JD: a deep learning network for jute diseases and pests detection from images</article-title>. <source>Plants</source> <volume>11</volume>, <elocation-id>7</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/plants11070937</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y. H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C. F.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>X. L.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z. X.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>S. D.</given-names>
</name>
<name>
<surname>Lan</surname> <given-names>Y. B.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Detection of the foreign object positions in agricultural soils using mask-RCNN</article-title>. <source>Bio. Engineering.</source> <volume>16</volume>, <elocation-id>1</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.25165/j.ijabe.20231601.7173</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>J. W.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>D. Z.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>R. Y.</given-names>
</name>
<name>
<surname>Cai</surname> <given-names>J. T.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L. C.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Improved YOLOX-tiny network for detection of tobacco brown spot disease</article-title>. <source>Front. Plant Sci.</source> <volume>14</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2023.1135105</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>He</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.X.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Identification of apple leaf diseases based on deep convolutional neural networks</article-title>. <source>Symmetry-BASEL</source> <volume>10</volume>, <elocation-id>1</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/sym10010011</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Pallathadka</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>)<article-title>Application of machine learning techniques in rice leaf disease detection</article-title> (<publisher-name>ELECTR NETWORK: ICAMS</publisher-name>) (Accessed <access-date>Proceedings of the 2022 International Conference on Advances in Materials Science</access-date>).</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Raman</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Chougule</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Chamola</surname> <given-names>V.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A low power consumption mobile based IoT framework for real-time classification and segmentation for apple disease</article-title>. <source>Micro. Microsystems.</source> <volume>94</volume>, <elocation-id>104656</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.micpro.2022.104656</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Ren</surname> <given-names>S.</given-names>
</name>
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>Faster r-CNN: towards real-time object detection with region proposal networks</article-title> (<publisher-name>NIPS</publisher-name>) (Accessed <access-date>Proceedings of the 29th Annual Conference on Neural Information Processing Systems</access-date>).</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sozzi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Kayad</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Marinello</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Cantalamessa</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Cogato</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Automatic bunch detection in white grape varieties using YOLOv3, YOLOv4, and YOLOv5 deep learning algorithms</article-title>. <source>AGRONOMY. BASEL.</source> <volume>12</volume>, <elocation-id>2</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/agronomy12020319</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sujatha</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Chatterjee</surname> <given-names>J.M.</given-names>
</name>
<name>
<surname>Jhanjhi</surname> <given-names>N.Z.</given-names>
</name>
<name>
<surname>Brohi</surname> <given-names>S.N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Performance of deep learning vs machine learning in plant leaf disease detection</article-title>. <source>Micro. Micro</source> <volume>80</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.micpro.2020.103615</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Thapa</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Kai</surname>
</name>
</person-group>. (<year>2020</year>). <article-title>The plant pathology challenge 2020 data set to classify foliar disease of apples</article-title>. <source>Appl. Plant Sci</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/aps3.11390</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>Z. Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Apple detection during different growth stages in orchards using the improved YOLO-V3 model</article-title>. <source>Comput. Agricul.</source> <volume>157</volume>, <fpage>417</fpage>&#x2013;<lpage>426</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2019.01.012</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wadhawan</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Situ</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Zouridakis</surname> <given-names>G.</given-names>
</name>
<etal/>
</person-group>. (<year>2012</year>). <article-title>A portable library for melanoma detection on handheld devices</article-title> (<publisher-name>IEEE</publisher-name>) (Accessed <access-date>Proceedings of the 8TH IEEE international symposium on biomedical imaging: from nano to macro</access-date>).</citation>
</ref>
<ref id="B24">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>C. Y.</given-names>
</name>
<name>
<surname>Alexey</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <source>YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors</source>. Available at: <uri xlink:href="https://arxiv.org/pdf/2207.02696.pdf">https://arxiv.org/pdf/2207.02696.pdf</uri>.</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>C. Y.</given-names>
</name>
<name>
<surname>Bochkovskiy</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>H. Y. M.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Scaled-YOLOv4: scaling cross stage partial network</article-title>,&#x201d; in <conf-name>Proceedings of the 2021 IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <fpage>13024</fpage>&#x2013;<lpage>13033</lpage> (ELECTR NETWORK: IEEE). doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01283</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Woo</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Kweon</surname> <given-names>I. S.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J.Y.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>CBAM: convolutional block attention module. lecture notes in computer science</article-title>,&#x201d; in <conf-name>Proceedings of the 15th European Conference on Computer Vision</conf-name>. <fpage>3</fpage>&#x2013;<lpage>19</lpage> (<publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>ECCV</publisher-name>). doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_1</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>P. F.</given-names>
</name>
<name>
<surname>Lo</surname> <given-names>E. R.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>B. T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>High performance depthwise and pointwise convolutions on mobile devices</article-title>. <source>Proc. 34th AAAI Conf. Artif. Intell.</source>, <fpage>6795</fpage>&#x2013;<lpage>6802</lpage>. doi: <pub-id pub-id-type="doi">10.1609/aaai.v34i04.6159</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Empowering things with intelligence: a survey of the progress, challenges, and opportunities in artificial intelligence of things</article-title>. <source>IEEE Internet Things J.</source> <volume>8:10</volume>, <fpage>7789</fpage>&#x2013;<lpage>7817</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JIOT.2020.3039359</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>M. X.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>ShuffleNet: an extremely efficient convolutional neural network for mobile devices</article-title>,&#x201d; in <conf-name>Proceedings of the 2018 IEEE/CVF conference on computer vision and pattern recognition (CVPR)</conf-name>. (<publisher-loc>Salt Lake City, UT</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>6848</fpage>&#x2013;<lpage>6856</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2018.00716</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>