<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2023.1231903</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>FOTCA: hybrid transformer-CNN architecture using AFNO for accurate plant leaf disease image recognition</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Hu</surname>
<given-names>Bo</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2327927"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Jiang</surname>
<given-names>Wenqian</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zeng</surname>
<given-names>Juan</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Cheng</surname>
<given-names>Chen</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>He</surname>
<given-names>Laichang</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/211668"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Information Engineering, Nanchang University</institution>, <addr-line>Nanchang</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Radiology, the First Affiliated Hospital of Nanchang University</institution>, <addr-line>Nanchang</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Second Clinical Medical College, Nanchang University</institution>, <addr-line>Nanchang</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>School of Mathematics and Statistics, Huazhong University of Science and Technology</institution>, <addr-line>Wuhan</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Muhammad Fazal Ijaz, Melbourne Institute of Technology, Australia</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Yongqiang Zheng, Southwest University, China; Nisha Pillai, Mississippi State University, United States</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Laichang He, <email xlink:href="mailto:laichang_he@163.com">laichang_he@163.com</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>09</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1231903</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>08</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Hu, Jiang, Zeng, Cheng and He</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Hu, Jiang, Zeng, Cheng and He</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Plants are widely grown around the world and have high economic benefits. plant leaf diseases not only negatively affect the healthy growth and development of plants, but also have a negative impact on the environment. While traditional manual methods of identifying plant pests and diseases are costly, inefficient and inaccurate, computer vision technologies can avoid these drawbacks and also achieve shorter control times and associated cost reductions. The focusing mechanism of Transformer-based models(such as Visual Transformer) improves image interpretability and enhances the achievements of convolutional neural network (CNN) in image recognition, but Visual Transformer(ViT) performs poorly on small and medium-sized datasets. Therefore, in this paper, we propose a new hybrid architecture named FOTCA, which uses Transformer architecture based on adaptive Fourier Neural Operators(AFNO) to extract the global features in advance, and further down sampling by convolutional kernel to extract local features in a hybrid manner. To avoid the poor performance of Transformer-based architecture on small datasets, we adopt the idea of migration learning to make the model have good scientific generalization on OOD (Out-of-Distribution) samples to improve the model&#x2019;s overall understanding of images. In further experiments, Focal loss and hybrid architecture can greatly improve the convergence speed and recognition accuracy of the model in ablation experiments compared with traditional models. The model proposed in this paper has the best performance with an average recognition accuracy of 99.8% and an F1-score of 0.9931. It is sufficient for deployment in plant leaf disease image recognition.</p>
</abstract>
<kwd-group>
<kwd>plant leaf disease image recognition</kwd>
<kwd>hybrid architecture</kwd>
<kwd>transformer-based models</kwd>
<kwd>adaptive Fourier Neural Operator</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="3"/>
<equation-count count="15"/>
<ref-count count="30"/>
<page-count count="12"/>
<word-count count="6932"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Technical Advances in Plant Science</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>As one of the main sources of economic growth in many developing countries, it is necessary to produce agricultural crops in a stable and efficient manner (<xref ref-type="bibr" rid="B11">Iqbal et&#xa0;al., 2018</xref>). However, the expansion of crops, combined with the overuse of pesticides and the exacerbation of global climate change, has led to an increase in the occurrence and spread of agricultural pests and diseases. Therefore, controlling these pests and diseases is becoming increasingly challenging. Early detection and treatment of such pests and diseases have unique advantages (<xref ref-type="bibr" rid="B21">Sanju and VelammaL, 2021</xref>). In the early stages of pest infestation, it is difficult to distinguish the leaves of affected plants from those of normal because of the high interclass variation in colour and profile and the low intraclass variation.</p>
<p>Traditional methods for identifying agricultural pests typically involve visual inspection of crops by farmers and agricultural experts. However, these approaches are costly, time-consuming, highly subjective, and non-transferable (<xref ref-type="bibr" rid="B19">Ouppaphan, 2017</xref>). Although some success has been achieved in classifying agricultural images using traditional image processing techniques, these methods face several challenges. First, they often require manual feature extraction, increasing workload. Second, the manually extracted features may not adequately represent the characteristics of agricultural images, leading to semantic gaps. Lastly, the variability and complexity of agricultural images, combined with factors such as image quality and shooting angle, can significantly affect the final recognition results, rendering these techniques unsuitable for large-scale applications (<xref ref-type="bibr" rid="B2">Bisen, 2021</xref>; <xref ref-type="bibr" rid="B20">Pan et&#xa0;al., 2022</xref>).</p>
<p>In light of relevant research on plant leaf disease identification and fine-grained recognition, we investigated issues associated with the relatively coarse recognition of algorithms in current crop pest identification approaches and their inadequate performance on datasets containing multiple, similar pests and diseases. We recognized the potential of the Transformer-based model and applied it to this domain. However, using the original patch and position embedding(Pap Embedding) would impose limitations on the experiment outcomes. Concurrently, utilizing the adaptive Fourier basis function to convert images to the frequency domain and deploying a CNN-Transformer architecture to separately extract local and global features could enhance the training upper limit. As a solution, we propose a novel hybrid architecture for plant leaf disease image recognition, termed FOTCA (where F and O signify Adaptive Fourier Neural Operator (AFNO) (<xref ref-type="bibr" rid="B8">Guibas et&#xa0;al., 2021</xref>) and TCA represents the Transformer-CNN Architecture). This approach addresses and optimizes the convergence issue and generalization capability of the ViT model, further improving training outcomes. Additionally, we evaluate the performance of FOTCA using the Plant-village dataset, a plant leaf disease dataset, to assess its scope and effectiveness.</p>
<p>In summary, this work makes the following three main contributions.</p>
<list list-type="bullet">
<list-item>
<p>This article applies an operator called Adaptive Fourier Neural Operator (AFNO) and learnable Fourier features which can replace traditional position encoding. Compared to traditional self-attention, AFNO maps images to frequency domain for better performance.</p>
</list-item>
<list-item>
<p>A model architecture that integrates both global and local features has been proposed, connecting CNNs and Transformers through inter-level concatenation to achieve coupling of global and local receptive fields. This hybrid architecture, which blends global and local features, can better utilize the extracted features and improve the performance and robustness of the model.</p>
</list-item>
<list-item>
<p>This article proposes that using Focal Loss as the loss function can effectively enhance the model&#x2019;s ability to train on difficult samples.</p>
</list-item>
</list>
<p>The rest of this article is organized as follows. The Section 3 mainly elucidates the details of the dataset composition and model composition structure, in Section 4 we specify the optimization scheme for settings other than the model.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Convolutional Neural Networks</title>
<p>Convolutional Neural Networks (CNNs) have become a popular and powerful tool for image recognition in various fields. In 1998, <xref ref-type="bibr" rid="B16">L&#xe9;cun et&#xa0;al. (1998)</xref> pioneering introduction of the LeNet-5 model laid the fundamental framework for CNNs. In 2012, the AlexNet model (<xref ref-type="bibr" rid="B13">Krizhevsky et&#xa0;al., 2012</xref>) used more convolutional layers and a larger parameter space to fit large-scale datasets in ILSVRC (ImageNet Large-Scale Visual Recognition Challenge). It was a groundbreaking demonstration of the advantages of deep neural networks over shallow neural networks, achieving significantly higher accuracy than the runner-up in recognition accuracy. This breakthrough established the status of CNNs in computer vision and brought new opportunities for plant leaf disease identification.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Fine-grained image recognition</title>
<p>State-of-the-art deep learning algorithms for image recognition predominantly rely on publicly available datasets such as ImageNet. Although utilizing these datasets enhances a model&#x2019;s capability to identify objects in images, including contour features (<xref ref-type="bibr" rid="B25">Wang D. et al., 2021</xref>; <xref ref-type="bibr" rid="B29">Zhang and Wang, 2016</xref>), discerning different types of pests affecting the same plant leaves in plant leaf disease images remains challenging due to their similar contour features (<xref ref-type="bibr" rid="B12">Kong et al., 2021</xref>). Merely transferring a model trained on other data categories to the study of plant leaf disease might not yield the anticipated accuracy. Consequently, models should be encouraged to learn fine-grained features of objects for improved results (<xref ref-type="bibr" rid="B1">Berg et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B27">Yang et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B5">Chen et&#xa0;al., 2019</xref>)</p>
<p>To tackle this issue, numerous researchers have explored convolutional neural networks in fine-grained recognition. For instance, <xref ref-type="bibr" rid="B28">Zhang et&#xa0;al. (2014)</xref> developed a model that surmounts these constraints by utilizing deep convolutional features derived from bottom-up region proposals for fine-grained recognition. Multi-proposal Net (<xref ref-type="bibr" rid="B30">Zhang et&#xa0;al., 2016</xref>) obtains image blocks through the Edge Box Crop method and incorporates an output layer of key points and visual features to further reinforce the local feature positional relationship between local features and overall information. Deep LAC (<xref ref-type="bibr" rid="B18">Lin et&#xa0;al., 2015</xref>) performs part localization, alignment and classification in the same network, and the VLF (valve linkage function) function is proposed for back propagation in Deep LAC, which is able to adaptively reduce the classification and alignment errors and update the localization results.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Transformer-based models for fine-grained visual recognition</title>
<p>As a decoding and encoding architecture model based on attention mechanism, Transformer (<xref ref-type="bibr" rid="B23">Vaswani et&#xa0;al., 2017</xref>) has been widely applied in the field of natural language processing. Inspired by this significant achievement, many scholars have transferred the Transformer structure to computer vision tasks and have achieved considerable success (<xref ref-type="bibr" rid="B4">Carion et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B6">Dosovitskiy et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B24">Wang Y. et&#xa0;al., 2021</xref>). The attention mechanism primarily identifies and classifies objects through different parts of an object, thus making it possible to focus on key features of the subject for datasets of plant leaf diseases with small differences between inter-class. As a result, the Transformer structure completes more successful recognition compared to CNN structures. Focusing on this area of study, <xref ref-type="bibr" rid="B3">Cai et&#xa0;al. (2021)</xref> introduced a ViT with adaptive attention that adds attention weakening and strengthening modules. This improves the performance of key features while capturing more feature information. <xref ref-type="bibr" rid="B7">Fu et&#xa0;al. (2017)</xref> used multiple scales of recurrent attention links to learn the target feature region and used within-scale classification loss and between-scale ranking loss to make the model more focused on the finer-grained features of the target object. <xref ref-type="bibr" rid="B9">He et&#xa0;al. (2022)</xref> proposed the TranFG framework, which integrates all original attention weights of the Transformer into one attention map, enabling the model to recognize image blocks and calculate their relationships while utilizing contrastive loss to expand the distance between confusion class feature representations. <xref ref-type="bibr" rid="B22">Sun (2019)</xref> improved the loss function in order to enable the classification model to learn features with greater distinction between more difficult to distinguish classes, and used feature map inhibition methods to enable the model to learn subtle differences.</p>
</sec>
</sec>
<sec id="s3" sec-type="materials|methods">
<label>3</label>
<title>Materials and methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset and data pre-processing</title>
<p>In this study, the base dataset was selected from the publicly available plant-village dataset on the web. This is a dataset specifically designed to study the work of various plant leaf disease recognition models, with a total of 54,303 images containing a total of 38 different species of 13 plant species (including apple, blueberry, cherry, corn, grape, orange, bell pepper, potato, raspberry, soybean, pumpkin, strawberry, tomato). Plant images, including non-plant leaf species, were inserted to train the model to recognize non-plant leaf images. The ratio of the training and validation sets was 8:2.</p>
<p>Before inputting the images into the model, the images need to be quantified uniformly. In the first step, the training set is expanded by applying some image enhancement techniques to each image independently superimposed, and the main data enhancement methods used in this study are RandomFlip, RandomCrop, RandomHorizontalFlip and RandomResizeCrop, mainly to eliminate the shift of the final recognition result by the change of the shooting angle in real life, and to prevent the neural network from overfitting phenomenon. In addition, the image size is uniformly adjusted to a square of size 224*224 pixels to facilitate further processing by the model. <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref> shows a portion of the dataset images and the pre-processing process of the images.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Plant Village dataset presentation and data augmentation examples. <bold>(A)</bold> A partial sample images of the dataset (containing healthy and sick). <bold>(B)</bold> Data augmentation operation examples.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g001.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Model</title>
<p>In this paper, we study applying deep transfer learning to all models in the model and comparison experiments. The Pre-train and Fine-tuning approach in deep transfer learning is the most convenient and reliable method for deep model learning. By putting a pre-trained model with some generalization ability onto a new similar dataset, the model will show superior performance than training a model from scratch after a simple and short training process. Similarly, the pre-trained model is adjusted for a specific dataset by a loss calculation function called model fine-tuning, and the loss calculation function widely used in deep neural networks today is the CrossentropyLoss (CE) (Shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref>), which is calculated as follows:</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mtext>Loss&#xa0;</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>,</mml:mo>
<mml:mover accent="true">
<mml:mi>p</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mi>log</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mi>p</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>log</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mover accent="true">
<mml:mi>p</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im1">
<mml:mover accent="true">
<mml:mi>p</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> is the predicted probability and <inline-formula>
<mml:math display="inline" id="im2">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula> is the actual outcome category. A remarkable feature of this loss calculation is that the losses are also treated consistently for simple and easily scored samples ( <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mo>&#x226b;</mml:mo>
<mml:mn>0.5</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>). This result may also occur when the losses of a large number of simple samples accumulate and the small loss may swamp the sparse classes, or when there are differences in the loss of different classes in a near-saturated training process. There are two intuitive ways to deal with this type of problem, one is to add weighting factors directly to the loss function, and the other is to add adjustable factors represented by focal loss. In this paper, focal loss is introduced into the model by reducing the loss function to the weights of the easy samples, which can help the loss function to favour the difficult samples and improve the accuracy of the difficult samples (Shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2C</bold>
</xref>). The calculation procedure is as follows:</p>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:msub>
<mml:mi>s</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:msup>
<mml:mi>log</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> reflects the proximity to the ground truth. Larger <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> means closer to category y, representing more accurate recognition. <inline-formula>
<mml:math display="inline" id="im6">
<mml:mi>&#x3b3;</mml:mi>
</mml:math>
</inline-formula> is the adjustable factor.</p>
<p>Due to the wide recognition of the ImageNet dataset, as well as the excellent performance and superior model fitting ability of the pre-trained models, all the experiments cited in this paper are based on the pre-trained models of the ImageNet dataset loaded with the corresponding models, effectively speeding up the training process of the models by means of transfer learning.</p>
<p>The FOTCA model studied in this article is mainly composed of a shortcut module and a Transformer module. The proposed overall model structure is shown in the <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2A</bold>
</xref>, which is mainly divided into three steps: Patch and Position Embedding, Transformer architecture based on adaptive Fourier Neural Operators and Classifier.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>FOTCA model architecture and two loss functions for fine-tuning <bold>(A)</bold> The figure provides an overview of the detailed structure of the entire FOTCA model&#x2019;s neural network. <bold>(B)</bold> Cross-Entropy loss for fine-tuning. <bold>(C)</bold> Focal loss for fine-tuning.belflow chart.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g002.tif"/>
</fig>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Patch and position embedding</title>
<p>The module consists of two sub-modules: Patch Embedding and Positional encoding based on learnable Fourier features. Compared with the regular position embedding, a primary advantage of Positional Embedding based on learnable Fourier features is that it provides richer positional information, especially for input from large datasets and long sequences. It can more finely encode each position through learnable Fourier basis functions, and this learnability makes it more flexible and adaptable. Additionally, it has better local connectivity and translational invariance when performing convolution operations, which can further enhance the model&#x2019;s performance.</p>
<p>For the input image <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, Assuming we use a patch size of <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, we can divide the input image into <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mrow>
<mml:mo mathsize="4.2">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>W</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo mathsize="4.2">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo mathsize="4.2">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mi>H</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo mathsize="4.2">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> local regions <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> according to the patch size. ( <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo mathsize="4.2">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mi>H</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo mathsize="4.2">]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mtext>&#x2003;</mml:mtext>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo mathsize="4.2">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mi>W</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo mathsize="4.2">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>). And expand it into a one-dimensional vector <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, whose size is D, where D is an adjustable parameter representing the vector dimension after patch embedding. For each matrix, we use a learnable weight matrix <inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of size <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> to map it to a low dimensional space:</p>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>patch&#x2009;</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mtext>patch</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>There, <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>patch&#x2009;</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the feature vector obtained through patch embedding.</p>
<p>Next, we generate position encoding vectors for each patch. For each coordinate binary (i, j), use a learnable Fourier feature function to generate a vector <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> with a size of 2 <inline-formula>
<mml:math display="inline" id="im17">
<mml:mi>K</mml:mi>
</mml:math>
</inline-formula>:</p>
<disp-formula>
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi>cos</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mo>&#x2329;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x232a;</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>sin</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mo>&#x2329;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x232a;</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mi>cos</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mo>&#x2329;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mi>&#x232a;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>sin</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mo>&#x2329;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#x232a;</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>H</mml:mi>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>W</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the normalized representation of input coordinates, <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are learnable parameters, and <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:mi>&#x2329;</mml:mi>
<mml:mo>&#xb7;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mo>&#xb7;</mml:mo>
<mml:mo>&#x232a;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> represents inner product between vectors. Finally, we add the position encoding vector to the feature vector obtained from patch embedding, resulting in the final feature vector <inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>:</p>
<disp-formula>
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Then, we can use encoders to encode this sequence for further processing and recognition tasks.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Adaptive Fourier Neural Operators</title>
<p>This module maps the input image to the frequency domain and uses adaptive Fourier basis functions to map each small block in the time domain to the frequency domain. This allows better extraction of frequency domain features and offers advantages in processing periodic as well as regular images, while providing better robustness to transformations such as rotation and scaling of the input data. Specifically, for each small block <inline-formula>
<mml:math display="inline" id="im23">
<mml:mi>x</mml:mi>
</mml:math>
</inline-formula>, its frequency-domain representation <inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:msup>
<mml:mtext mathvariant="bold-italic">x</mml:mtext>
<mml:mrow>
<mml:mtext>freq</mml:mtext>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is obtained by first mapping it from the time domain to the frequency domain through Fast Fourier Transform (FFT). The <inline-formula>
<mml:math display="inline" id="im25">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> element represents the frequency-domain representation at position <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> on the frequency domain.</p>
<disp-formula>
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msubsup>
<mml:mtext mathvariant="bold-italic">X</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>freq</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
</mml:mstyle>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi>&#x3c0;</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mi>P</mml:mi>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Next, we will map each patch from the time domain to the frequency domain using adaptive Fourier basis functions. Specifically, for each small block <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold-italic">X</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, we calculate the inner product between it and the adaptive Fourier basis functions, and use the inner product value as coefficients.</p>
<disp-formula>
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>p</mml:mi>
</mml:munderover>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>p</mml:mi>
</mml:munderover>
</mml:mstyle>
<mml:msub>
<mml:mtext mathvariant="bold-italic">X</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The coefficient of small block <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold-italic">X</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> under the <inline-formula>
<mml:math display="inline" id="im29">
<mml:mrow>
<mml:msup>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> adaptive Fourier basis function is <inline-formula>
<mml:math display="inline" id="im30">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im31">
<mml:mi>p</mml:mi>
</mml:math>
</inline-formula> is the size of the small block, and <inline-formula>
<mml:math display="inline" id="im32">
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the value of the <inline-formula>
<mml:math display="inline" id="im33">
<mml:mrow>
<mml:msup>
<mml:mi>k</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> adaptive Fourier basis function at position <inline-formula>
<mml:math display="inline" id="im34">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. The specific calculation formula is as follows.</p>
<disp-formula>
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mi>B</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>cos</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#xb7;</mml:mo>
<mml:mi>sin</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>k</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where <inline-formula>
<mml:math display="inline" id="im35">
<mml:mrow>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the spatial location vector of the center point of the <inline-formula>
<mml:math display="inline" id="im36">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> patch in the input image and <inline-formula>
<mml:math display="inline" id="im37">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the adaptive basis vector.</p>
<p>Expand all small adaptive Fourier basis functions into a large adaptive Fourier basis function matrix <inline-formula>
<mml:math display="inline" id="im38">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x2102;</mml:mi>
<mml:mrow>
<mml:msup>
<mml:mi>p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im39">
<mml:mi>K</mml:mi>
</mml:math>
</inline-formula> is the number of adaptive Fourier basis functions. At the same time, expand all small adaptive Fourier coefficients into a large adaptive Fourier coefficient matrix <inline-formula>
<mml:math display="inline" id="im40">
<mml:mrow>
<mml:mtext mathvariant="bold-italic">A</mml:mtext>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x2102;</mml:mi>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, where N is the number of small blocks. Then, by performing global average pooling on A, obtain the weight <inline-formula>
<mml:math display="inline" id="im41">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> of each basis function, and the calculation formula is:</p>
<disp-formula>
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>W</mml:mi>
<mml:mo stretchy="false">/</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:munderover>
</mml:mstyle>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The original data is ultimately mapped to the frequency domain through basis functions and basis function weights.</p>
<disp-formula>
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:msub>
<mml:mtext>X</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
</mml:mstyle>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mtext>B</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where <inline-formula>
<mml:math display="inline" id="im42">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>:</mml:mo>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the flattened vector of patch (i,j). Finally, we merge the frequency-domain results of all patches into a large frequency-domain tensor <inline-formula>
<mml:math display="inline" id="im43">
<mml:mrow>
<mml:mtext>F</mml:mtext>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x2102;</mml:mi>
<mml:mrow>
<mml:mfrac>
<mml:mi>H</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mi>W</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. For channel mixing, it is done through an MLP on this frequency-domain tensor to extract high-level features.</p>
<disp-formula>
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:msubsup>
<mml:mi>Z</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mtext>MLP&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mi>H</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mi>W</mml:mi>
<mml:mi>P</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Finally, we concatenate all frequency domain features <inline-formula>
<mml:math display="inline" id="im44">
<mml:mi>z</mml:mi>
</mml:math>
</inline-formula> together to form the final feature vector.</p>
<disp-formula>
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mi>Z</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>z</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>z</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>z</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The feature vector can be fed into any type of neural network for further training or inference operations. This process has strong expressiveness and interpretability in data processing.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Classifier</title>
<p>When using the Transformer-based model, its output vector is often passed to a classifier for image recognition. The traditional approach is to use the hidden layers of MLP to collect features and perform classification. However, there are many localized feature points in plant leaf disease images, and these localized feature points can usually be extracted and represented more efficiently by means of convolutional operations. Therefore, the use of CNNs may be more appropriate for such problems. Here, we consider using a basic block as a classifier, which includes convolutional layers and batch normalization layers, as well as shortcut operations. By upsampling to a higher dimension to obtain local feature values, feature transfer and information fusion are achieved. Also, shortcut operations can avoid gradient disappearance and model overfitting problems caused by excessive stacking of convolutional layers.</p>
<p>After downsampling or upsampling the feature maps using convolutional layers, the downsampled or upsampled feature maps are added to the output feature map using element-wise addition, which implements shortcut. This operation allows the model to better learn details and local features in the input, thereby improving the model&#x2019;s performance. In this article&#x2019;s classifier design, a <inline-formula>
<mml:math display="inline" id="im45">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> convolutional layer is embedded in the shortcut to adjust the depth of the feature maps so they can be added to the feature maps from the next layer. This helps preserve lower-level features, enhance inter-channel communication, and improve the model&#x2019;s performance. Specifically, the input vector undergoes two <inline-formula>
<mml:math display="inline" id="im46">
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> convolutional layers consecutively, with a certain amount of non-linear activation function (such as ReLU) inserted between them, as shown below:</p>
<disp-formula>
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext>ReLU</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>x</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mtext>ReLU</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>At the same time, downsampling is applied to the input vector to promote the underlying features towards the final recognition result. This operation can be achieved by using a <inline-formula>
<mml:math display="inline" id="im47">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> convolution layer in the shortcut. Finally, add the output of twoconsecutive convolution layers to the output of the shortcut, as shown below:</p>
<disp-formula>
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo>=</mml:mo>
<mml:mtext>ReLU&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mo>=</mml:mo>
<mml:mtext>ReLU&#xa0;</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mtext>ReLU&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>x</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Among them, <inline-formula>
<mml:math display="inline" id="im48">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the output of the shortcut, and it needs to ensure that its depth is the same as the output depth of the second convolution layer.</p>
<p>In the Classifier of this article, it has been demonstrated that shortcuts effectively alleviate the gradient vanishing problem, resulting in a decent convergence rate even when applied to very deep models. Using two <inline-formula>
<mml:math display="inline" id="im49">
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> convolutional layers instead of a larger kernel increases nonlinearity, allows for faster collection and extraction of local features, and reduces the number of parameters, making the model easier to train and generalize.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments and discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental environment</title>
<p>The models we studied was developed based on the open source of deep learning framework pytorch1.11.0 with the following experimental equipment: CPU is Intel(R) Xeon(R) Platinum 8255C, GPU is single card V100-SXM2-32GB, CUDA version is 11.3, and programming language is Python3.8 (ubuntu20.04).</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental parameters, evaluation indicators</title>
<p>Based on earlier work by scholars (<xref ref-type="bibr" rid="B14">Kumar et&#xa0;al., 2023a</xref>; <xref ref-type="bibr" rid="B15">Kumar et&#xa0;al., 2023b</xref>; <xref ref-type="bibr" rid="B26">Wei et&#xa0;al., 2023</xref>), we selected DenseNet-169, Inception-v3, VGG19, ResNet-50, ResNet-101, ViT as our comparative experimental models to verify the feasibility and reliability of the improved model plant leaf disease image recognition method in this paper. We all iteratively update the pre-trained model parameters based on each model to accelerate the model convergence during the training process.</p>
<p>The training process uniformly uses the stochastic gradient descent SGD optimization algorithm to optimize its model. The same learning rate adjustment strategy is used for each parameter in the model, and the learning rate is dynamically adjusted using LambdaLR, which is adjusted by adjusting the learning rate according to the number of learning rate updates to linearly decay on top of the original one. The loss calculation function of the FOTCA model uses focal loss, and the rest of the models in the comparison experiments use the cross-entropy loss function. Dropout regularization is also used in the training process. By randomly dropping some neuron connections temporarily during the training process, the purpose is to effectively avoid overfitting of the model during the training process. The generalization ability of the model is also enhanced. The specific hyperparameter settings for the experiments are shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Hyperparameter tuning for the model.</p>
</caption>
<table frame="hsides">
<tbody>
<tr>
<td valign="top" colspan="2" align="left">Initial learning rate</td>
<td valign="top" align="left">0.001</td>
</tr>
<tr>
<td valign="top" colspan="2" align="left">Epochs</td>
<td valign="top" align="left">100</td>
</tr>
<tr>
<td valign="top" colspan="2" align="left">Batch size</td>
<td valign="top" align="left">8</td>
</tr>
<tr>
<td valign="top" colspan="2" align="left">Image size(for all)</td>
<td valign="top" align="left">224*224</td>
</tr>
<tr>
<td valign="top" colspan="2" align="left">Image size(for Inception-v3)</td>
<td valign="top" align="left">299*299</td>
</tr>
<tr>
<td valign="top" rowspan="4" align="left">Transformer-based models</td>
<td valign="top" align="left">Embedding dimension</td>
<td valign="top" align="left">768</td>
</tr>
<tr>
<td valign="top" align="left">Patch size</td>
<td valign="top" align="left">16</td>
</tr>
<tr>
<td valign="top" align="left">Head number</td>
<td valign="top" align="left">8</td>
</tr>
<tr>
<td valign="top" align="left">Depth</td>
<td valign="top" align="left">12</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>This table presents the various hyperparameters and their selected values used of the proposed FOTCA model and compared models.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>We choose accuracy, adjustment time for accuracy, loss, adjustment time for loss, F1-score, parameters and FLOPS. We define for the first time the adjustment time for accuracy (loss). Adjustment time is the number of iterations required during model training to bring the model performance from the initial performance to 95% difference from the final converged performance. This metric reflects the ability of the model in terms of speed of convergence as well as training efficiency. The concept of adjustment time helps to deeply evaluate and compare the speed of the model training process. A shorter adjustment time means that the model reaches its potential performance faster within a limited number of iterations. Thus, with this metric, researchers can more intuitively assess the performance gap between different models.</p>
<p>In addition, adjustment time complements other commonly used evaluation metrics (e.g., accuracy, loss, etc.) to provide researchers with a more comprehensive view of how a model performs during training. In the case of similar model performance, shorter adjustment times may be an advantage due to greater savings in computational resources and time.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparative experiments</title>
<p>Firstly, the performance of Vision Transformer (ViT) is typically affected by the size of the training dataset due to the fact that ViT is based on transformers methods, which typically require a large amount of data for effective training. This is because transformers models, including ViT, are variants of the self-attention mechanism, which allows the model to capture the global relationships of the input data, but this mechanism also requires a large amount of data to support.</p>
<p>For the specific minimum sample requirement, it may vary as it depends on several factors, including the size of the model (e.g., the number of layers of the model, the number of hidden units, etc.), the complexity of the task (<xref ref-type="bibr" rid="B17">Lee et&#xa0;al., 2021</xref>) and the distribution of the data. Therefore, it can only be explored by trying the following experimental procedure.</p>
<p>We chose to compare the ViT model with the ResNet101 model. This decision was made because ResNet101 has similar flops and comparable convergence capability to ViT on large-scale datasets. In preliminary experiments, we selected training set proportions of 6%, 8%, 10%, 20%, and 80%, respectively, to investigate the convergence capability and accuracy of ViT and ResNet101 on medium-sized datasets. The experimental results are shown in <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3A&#x2013;C</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>
<bold>(A)</bold> Accuracy of the ViT model on small and medium-sized datasets: This plot demonstrates the training performance of the ViT model on datasets with different sizes. <bold>(B)</bold> Accuracy of the ResNet101 model on small and medium-sized datasets: This plotdemonstrates the training performance of the ResNet101 model on datasets with different sizes. <bold>(C)</bold> Comparison of the accuracy of ViT and ResNet101 for small and medium-sized datasets. <bold>(D)</bold> Scatter plot comparison of model accuracy and FLOPs(floating-point operations per second) for the models.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g003.tif"/>
</fig>
<p>We observed that ViT&#x2019;s training performance deteriorated rapidly on medium-sized datasets when the size of the training set was small. In contrast, ResNet101 demonstrated smaller variations and more stable curves, with minimal decreases in accuracy. Particularly, there was a significant turning point in ViT&#x2019;s performance when the training set was at 8%. However, these experiments were not able to establish a definitive standard for determining the minimum sample requirements. On the contrary, the optimal dataset size may vary depending on specific applications and data characteristics. In many cases, if feasible, using larger datasets often yields better results.</p>
<p>Next, we compared six mainstream image recognition models, using the same backbone training model and dataset to ensure the fairness of testing. The only exception is that the input image size must be <inline-formula>
<mml:math display="inline" id="im50">
<mml:mrow>
<mml:mn>299</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>299</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> for the Inception-v3 model. We show the accuracy and loss values of each model during the iteration process in the chart. Due to the significant difference between the initial and final iteration results, the observation of the later iteration effect is not very clear and accurate. Therefore, we re-drew the chart showing the changes in the accuracy and loss values of later iterations between 20 and 100 iterations to achieve deeper analysis and optimization. This approach can enhance our understanding of the model&#x2019;s performance, which helps us further improve and optimize our algorithms. The final experimental effect comparison is shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. Specifically shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3D</bold>
</xref>. The plot offers a visual representation of the balance between performance and computational efficiency among different models. The final result of the iteration is shown in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Performance metrics of the proposed FOTCA model and compared models.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Model</th>
<th valign="top" align="left">Accuracy(%)</th>
<th valign="top" align="left">Adjustment time</th>
<th valign="top" align="left">F1 score(%)</th>
<th valign="top" align="left">Loss</th>
<th valign="top" align="left">Adjustment time</th>
<th valign="top" align="left">Params<italic>(M)</italic>
</th>
<th valign="top" align="left">FLOPS<italic>(G)</italic>
</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DenseNet169</td>
<td valign="top" align="left">98.3</td>
<td valign="top" align="left">13</td>
<td valign="top" align="left">91.85</td>
<td valign="top" align="left">0.082</td>
<td valign="top" align="left">14</td>
<td valign="top" align="left">14.15</td>
<td valign="top" align="left">3.44</td>
</tr>
<tr>
<td valign="top" align="left">Inceptionv3</td>
<td valign="top" align="left">96.8</td>
<td valign="top" align="left">28</td>
<td valign="top" align="left">94.45</td>
<td valign="top" align="left">0.15</td>
<td valign="top" align="left">29</td>
<td valign="top" align="left">23.83</td>
<td valign="top" align="left">2.86</td>
</tr>
<tr>
<td valign="top" align="left">ResNet50</td>
<td valign="top" align="left">99.5</td>
<td valign="top" align="left">21</td>
<td valign="top" align="left">98.23</td>
<td valign="top" align="left">0.103</td>
<td valign="top" align="left">27</td>
<td valign="top" align="left">25.56</td>
<td valign="top" align="left">4.13</td>
</tr>
<tr>
<td valign="top" align="left">ResNet101</td>
<td valign="top" align="left">97.6</td>
<td valign="top" align="left">32</td>
<td valign="top" align="left">95.69</td>
<td valign="top" align="left">0.107</td>
<td valign="top" align="left">32</td>
<td valign="top" align="left">44.55</td>
<td valign="top" align="left">7.87</td>
</tr>
<tr>
<td valign="top" align="left">VGG-19</td>
<td valign="top" align="left">97.8</td>
<td valign="top" align="left">28</td>
<td valign="top" align="left">91.42</td>
<td valign="top" align="left">0.206</td>
<td valign="top" align="left">30</td>
<td valign="top" align="left">78.14</td>
<td valign="top" align="left">29.96</td>
</tr>
<tr>
<td valign="top" align="left">ViT</td>
<td valign="top" align="left">99.4</td>
<td valign="top" align="left">9</td>
<td valign="top" align="left">98.03</td>
<td valign="top" align="left">0.026</td>
<td valign="top" align="left">9</td>
<td valign="top" align="left">58.07</td>
<td valign="top" align="left">11.28</td>
</tr>
<tr>
<td valign="top" align="left">FOTCA</td>
<td valign="top" align="left">99.8</td>
<td valign="top" align="left">11</td>
<td valign="top" align="left">99.31</td>
<td valign="top" align="left">0.005</td>
<td valign="top" align="left">9</td>
<td valign="top" align="left">59.14</td>
<td valign="top" align="left">11.87</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Model Accuracy(%)Adjustment F1-Loss Adjustment Params(M) FLOPS(G).</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Showing training accuracy and loss comparison between the proposed FOTCA model and compared models in the study. <bold>(A)</bold> Training accuracy curve for all models. <bold>(B)</bold> Training loss curve for all models. <bold>(C)</bold> Partial training accuracy curve for all models. <bold>(D)</bold> Partial training loss curve for all models.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g004.tif"/>
</fig>
<p>According to the chart, the six mainstream image recognition models showed good feature extraction and representation abilities in the first <inline-formula>
<mml:math display="inline" id="im51">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mn>20</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> epochs, with an accuracy of over 95% and a loss of below 0.35 for all of them. However, we observe a strange status quo demonstrated in the figure: the growth curve of the Transformer-based model is more stable, but the curve of the CNNs model fluctuates up and down consistently at 1% (in accuracy) and 0.03 (in loss), and DenseNet-169 even overfitted, with its accuracy and loss at the <inline-formula>
<mml:math display="inline" id="im52">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mn>29</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> epoch values are the highest (99.8%) and lowest (0.072) of the whole iteration, respectively. Subsequently, the accuracy gradually declined and finally stabilized at about 98.3%. This phenomenon may be due to thedifferent learning patterns and model depths of CNNs and Transformer-based models(containing FOTCA and ViT) during the training process.The convolutional operation of CNN networks requires a large number of parameters. In contrast, the ViT network uses aself-attention layer instead of a fixed convolutional layer, which is better able to cope with long sequences of image features. The training error of ViT networks is usually less jittery due to the relatively simple training process of the self-attention layer. Meanwhile, the characteristics of the attention link are more suitable for training fine-grained plant leaf disease image recognition tasks, which is further demonstrated inside the section4.4. At the same time, overfitting is sooner or later as long as the network is large and deep enough, as indicated in (<xref ref-type="bibr" rid="B10">Hinton et&#xa0;al. (2012)</xref>. The large depth of DenseNet-169 (model depth of 169) reflects the overfitting of the model more quickly.</p>
<p>Meanwhile, comparing the adjustment time of various models, ViT&#x2019;s training results and FOTCA model performance are optimal, both achieving near-maximum effects at the 9<italic>
<sup>th</sup>
</italic> epoch. Other CNN models are relatively slower and require higher computational costs to achieve the same training effect. Therefore, FOTCA model exhibits efficiency and accuracy in feature extraction, learning ability, and model convergence efficiency, making it a more excellent image recognition model.</p>
<p>Two categories of performance can be identified in the comparison experiments. One is represented by models such as VGG-19, DenseNet-169 and Inception-v3, whose accuracy did not reach 98% but has already approached the best performance, so the trend is no longer rising. The other category is represented by models such as ViT, ResNet-50, and ResNet-101, which are based on FOTCA proposed in this article. During the iteration process, these models can approach an accuracy of 99.9%, which means that almost perfect image recognition results can be achieved through limited training. Additionally, FOTCA has leading positions in both accuracy and F1-score indicators compared to other models, and it is the only model with an F1-score of 0.9931. On the basis of efficient recognition using ViT, FOTCA further improves accuracy, loss, and F1-score without increasing parameter quantity and FLOPS compared to other models. This makes its performance nearly perfect, as shown in all the final parameters, parameter quantities, and FLOPS charts for all models in the experiment. Therefore, the FOTCA model proposed in this article achieves the best performance in plant leaf disease recognition.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Model visualization analysis</title>
<p>We proceeded to discuss the differences between the models in terms of their focus on image features by selecting 11 more representative photographs from the dataset that contained features that were essentially the basic features under the category. One photo of a healthy cherry leaf was taken with a global angle of view of the whole leaf and with a portion of the background included (this allows surveying the model&#x2019;s ability to segment objects and backgrounds and verifying whether the model can focus more on the objects themselves), with the diseased and healthy traits shown on the leaf as curling at the edges of the diseased leaf and wilting of the whole leaf, and one partial image of a leaf with grey leaf spot A partial image of a maize leaf with grey leaf spot, which does not include the background (this verifies the model&#x2019;s ability to focus on fine-grained features of the object and the size of the receptive field), and the trait of grey leaf spot on the maize showing random grey-brown rectangular long or irregular long spots and brown-black pockmarks throughout the leaf, with additional longitudinal correlation images having similar features. The graph shows the attention of the seven models to the same image, where the warmer the colour, the more attention the model pays to that feature, and the colour distribution gives an indication of the model&#x2019;s ability to accurately identify the image features. Each model feature concern map is showing in the <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Attention heatmaps for the models on representative plant leaf disease images. The attention heatmaps visually demonstrate the regions within the image that the respective models focus on, highlighting their effectiveness at capturing relevant features for accurate recognition.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g005.tif"/>
</fig>
<p>First of all, both FOTCA and ViT models focus more accurately on the details of cherry leaf edges, but in the results, it was found that the ViT model focuses more on the whole leaf surface, which generates more errors and retains unnecessary information, thus reducing the recognition accuracy. At the same time, none of the CNNs can accurately and carefully focus on the correct plant leaf disease feature points, and the fluctuating anomalies of the metrics (including accuracy and loss values) during the training process are proved accordingly, and it also sideways proves the results that the training of Convolutional Networks is not as good as the ViT effect in the recognition of plant leaf disease images.</p>
<p>In addition, during the recognition of maize grey leaf spot leaves, the FOTCA and ViT models focused on the brown-black pockmarks on the leaves, and the ViT model focused more evenly on the background colour similar to the subject colour, Inception-v3 was able to observe the rectangular long stripes on the image more accurately, while the rest of the models focused on looser areas and could not reflect a more convincing pattern. In turn, the rest of the images show that the FOTCA model focuses more on the fine-grained features of the edges and less on the details of the whole leaf than the ViT model, and the five CNN models, on the other hand, have a more mixed focus, tending to focus on the global information and main features of the images. From this we can conclude that the Transformer-based models focuses more on the global features of the image, forming a global attention to the image, while the CNNs model focuses more on the subtle features, forming a local attention to the image.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Ablation experiment</title>
<p>To evaluate the effectiveness of each improvement module in the FOTCA model, we conducted ablation experiments using the original dataset. We compared four models: the FOTCA model, the FOTCA model with original Pap Embedding, the FOTCA model with MLP-Classifier, and the FOTCA model lacking data augmentation. We plotted the accuracy and loss curves for each model as a function of epoch value in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>. Detailed training data is presented in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. Partial iteration accuracy and loss values are not shown here because the comparison of the individual models in the results of this experiment can be shown more clearly in Figure.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Showing training accuracy and loss comparison between the proposed FOTCA model and ablation experiment models with different parts removed. <bold>(A)</bold> Training accuracy curve for all models. <bold>(B)</bold> Training loss curve for all models.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1231903-g006.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Performance metrics of the proposed FOTCA model and ablation study models with different parts removed.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left">Model</th>
<th valign="top" align="left">Accuracy(%)</th>
<th valign="top" align="left">Adjustment time</th>
<th valign="top" align="left">F1-score(%)</th>
<th valign="top" align="left">loss</th>
<th valign="top" align="left">Adjustment time</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">FOTCA</td>
<td valign="top" align="left">99.8</td>
<td valign="top" align="left">12</td>
<td valign="top" align="left">99.31</td>
<td valign="top" align="left">0.005</td>
<td valign="top" align="left">9</td>
</tr>
<tr>
<td valign="top" align="left">Using the original Pap Embedding</td>
<td valign="top" align="left">99.1</td>
<td valign="top" align="left">12</td>
<td valign="top" align="left">96.31</td>
<td valign="top" align="left">0.029</td>
<td valign="top" align="left">11</td>
</tr>
<tr>
<td valign="top" align="left">Using MLP-Classifier</td>
<td valign="top" align="left">96.1</td>
<td valign="top" align="left">12</td>
<td valign="top" align="left">85.91</td>
<td valign="top" align="left">0.0341</td>
<td valign="top" align="left">15</td>
</tr>
<tr>
<td valign="top" align="left">Missing data augmentation</td>
<td valign="top" align="left">99.7</td>
<td valign="top" align="left">11</td>
<td valign="top" align="left">98.94</td>
<td valign="top" align="left">0.019</td>
<td valign="top" align="left">12</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4_5_1">
<label>4.5.1</label>
<title>Between FOTCA model and the model using original pap embedding</title>
<p>Using Fourier transform can better capture texture and detail information in the image. This is because the Fourier transform can convert spatial-domain information into frequency-domain information, allowing the model to better process this information. In addition, using Fourier transform can greatly reduce the dimensionality of the vector and thus reduce computation complexity.</p>
<p>However, Fourier transformation may cause some spatial information loss, thereby affecting the performance of image recognition tasks. From the experimental results, it can be seen that using Pap Embedding with Fourier operator can partially improve the accuracy and convergence speed of the model, thus achieving better performance. Compared with using original Pap Embedding, the accuracy was improved by 0.7% and the loss decreased by 0.014.</p>
</sec>
<sec id="s4_5_2">
<label>4.5.2</label>
<title>Between FOTCA model and the model using MLP &#x2212; classifier</title>
<p>The classifier used in this article contains shortcut connections across layers, which can help information transmit more effectively, avoid gradient vanishing and information bottleneck problems, and has better explainability. In contrast, using MLP as the classifier will lose some important spatial structure information because MLP is a fully connected layer that flattens all features together and ignores their relationships. The classifier in this article can reconstruct spatial structure information of images more accurately, retain more feature information, and improve the model&#x2019;s recognition accuracy. At the same time, as a classifier, it can help reduce the risk of overfitting by helping the model learn more robust features that have similar responses for different input data, thus improving the model&#x2019;s generalization ability, which is also reflected in the training results.</p>
</sec>
<sec id="s4_5_3">
<label>4.5.3</label>
<title>Between FOTCA model and the model lacking data augmentation</title>
<p>From the iteration graph, it can be observed that the model without data augmentation behaves extremely similarly to the original model during the process of iteration (the accuracy of the original model is 99.8%, and the accuracy of the comparative model is 99.7%). Even in the early stages of training, it produces better results than the original model, possibly due to the large amount of pest and disease data in this study, which allows for a sufficient sample size to collect fine-grained features without the need for data augmentation. However, continuous data augmentation during the initial stages of model training causes the model to learn orientation features while learning the original features of the photos, increasing the computational load of the model, resulting in a less effective performance on the test set compared to the model without data augmentation.</p>
<p>For the model using the original Pap Embedding, its learning ability is constrained for the same linkages, which adds limitations to the model architecture. As for the model using MLP-Classifier, its performance change is the largest among the four models, indicating significant contributions of the classifier used in this paper, which can significantly improve the fitting effect of the model. At the same time, the adjustment time of accuracy (loss) for the four experiments is roughly the same, indicating that making minor changes to the model architecture under the same model does not have a significant effect on convergence efficiency. The FOTCA model is still a fast-converging model that is leading in plant leaf disease image recognition.</p>
</sec>
</sec>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusions</title>
<p>Based on the above experimental results, the Transformer-based models outperforms the CNNs in the field of plant leaf disease image recognition, mainly because it allows a more detailed and accurate focus on image features., and can surpass the CNNs in terms of accuracy, loss, model matching speed, convergence efficiency, etc. The FOTCA model proposed in this paper can further improve its feature observation and extraction ability, and has a tendency to improve in accuracy, loss and F1-score, while demonstrating the enormous potential and application of adaptive Fourier operators. We will continue to extend our model in the future to combine diverse approaches for plant crop disease identification and detection in complex contexts in complex future contexts.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>BH contributed to the conceptualization and design of the study and constructed improved experiments. WJ and JZ organized data organization and analysis validation. CC and LH wrote some manuscripts. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="funding-information">
<title>Funding</title>
<p>This article is supported by National Natural Science Foundation of China (Grant/Award Number: &#x201c;81460329&#x201d;) and the Natural Science Foundation of Jiangxi Province (Grant/Award Numbers: &#x201c;20192ACBL20039/202310454/202310456&#x201d;).</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>We would like to express our sincere gratitude to all the individuals who supported me throughout my research project. Their encouragement, moral support, and companionship were invaluable and made this work possible.</p>
</ack>
<sec id="s9" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Berg</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J. K.</given-names>
</name>
<name>
<surname>Woo Lee</surname> <given-names>S. A.</given-names>
</name>
<name>
<surname>Alexander</surname> <given-names>M. L. M.</given-names>
</name>
<name>
<surname>Jacobs</surname> <given-names>D. W.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Birdsnap: Largescale fine-grained visual categorization of birds</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR))</conf-name>, <publisher-name>IEEE</publisher-name>, <publisher-loc>Columbus, OH, USA</publisher-loc>. <fpage>2011</fpage>&#x2013;<lpage>2018</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bisen</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep convolutional neural network based plant species recognition through features of leaf</article-title>. <source>Multimed. Tools Appl.</source> <volume>80</volume>, <fpage>6443</fpage>&#x2013;<lpage>6456</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/S11042-020-10038-W</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Cai</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Weng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>A transformer architecture with adaptive attention for fine-grained visual classification</article-title>,&#x201d; in <conf-name>International Conference on Computer and Communications</conf-name> (<publisher-name>IEEE</publisher-name>), <publisher-loc>Chengdu, China</publisher-loc>. <fpage>863</fpage>&#x2013;<lpage>867</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Carion</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Massa</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Synnaeve</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Usunier</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Kirillov</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Zagoruyko</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>). End-to-end object detection with transformers</article-title>,&#x201d; in <conf-name>Proceedings of the European Conference on Computer Vision</conf-name>, <publisher-name>Springer</publisher-name>, <publisher-loc>Glasgow, UK</publisher-loc>. <fpage>213</fpage>&#x2013;<lpage>229</lpage>, Springer.</citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bai</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Mei</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Destruction and construction learning for fine-grained image recognition</article-title>,&#x201d; in <conf-name>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name> (<publisher-name>IEEE</publisher-name>), <publisher-loc>Long Beach, CA, USA</publisher-loc>. <fpage>5152</fpage>&#x2013;<lpage>5161</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2019.00530</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dosovitskiy</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Beyer</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Kolesnikov</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Weissenborn</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhai</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Unterthiner</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). &#x201c;<article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>,&#x201d; in <conf-name>Proceedings of the International Conference on Learning Representations</conf-name>. <uri xlink:href="https://OpenReview.net">https://OpenReview.net</uri>, virtual, Formerly Addis Ababa ETHIOPIA</citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Mei</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Look closer to see better:recurrent attention convolutional ne ural network for fine-grained image recognition</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <publisher-name>IEEE</publisher-name>, <publisher-loc>Hawaii, USA</publisher-loc>. <fpage>4438</fpage>&#x2013;<lpage>4446</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guibas</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Mardani</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Anandkumar</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Catanzaro</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Adaptive fourier neural operators: Efficient token mixers for transformers</article-title>. <source>arXiv preprint</source>. <italic>arXiv:2111.13587</italic>. doi:&#xa0;<pub-id pub-id-type="doi">10.4505/arXiv.2111.13587</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>J.-N.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Transfg: A transformer architecture for fine-grained recognition</article-title>,&#x201d; in <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>, <publisher-name>AAAI</publisher-name>, <publisher-loc>Pomona, California</publisher-loc>. <fpage>852</fpage>&#x2013;<lpage>860</lpage>.</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hinton</surname> <given-names>G. E.</given-names>
</name>
<name>
<surname>Srivastava</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Krizhevsky</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Krizhevsky</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Salakhutdinov</surname> <given-names>R. R.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>. <source>arXiv</source> preprint arXiv:1207.0580. doi:&#xa0;<pub-id pub-id-type="doi">10.9774/GLEAF.978-1-909493-38-42</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Iqbal</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Sharif</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Shah</surname> <given-names>J.</given-names>
</name>
<name>
<surname>ur Rehman</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Javed</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>An automated detection and classification of citrus plant diseases using image processing techniques: A review</article-title>. <source>Comput. Electron. Agric.</source> <volume>153</volume>, <fpage>12</fpage>&#x2013;<lpage>32</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2018.07.041</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kong</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Jin</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multi-stream hybrid architecture based on cross-level fusion strategy for finegrained crop species recognition in precision agriculture</article-title>. <source>Comput. Electron. Agric.</source> <volume>185</volume>, <elocation-id>106134</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2021.106134</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname> <given-names>G. E.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Commun. ACM</source> <volume>60</volume>, <fpage>84</fpage>&#x2013;<lpage>90</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1145/3065386</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>S. K.</given-names>
</name>
<name>
<surname>Kaur</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>U.</given-names>
</name>
</person-group> (<year>2023</year>a). <article-title>Vi-net: A hybrid deep convolutional neural network using vgg and inception v3 model for copy-move forgery classification</article-title>. <source>J. Vis. Comun. Image Represent.</source> <volume>89</volume>, <fpage>103644</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.jvcir.2022.103644</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Pal</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>V. P.</given-names>
</name>
<name>
<surname>Jaiswal</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2023</year>b). <article-title>Performance evaluation of resnet model for classification of tomato plant disease</article-title>. <source>Epidemiologic Methods</source> <volume>12</volume>, <fpage>20210044</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1515/em-2021-0044</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>L&#xe9;cun</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bottou</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Haffner</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc. IEEE</source> <volume>86</volume>, <fpage>2278</fpage>&#x2013;<lpage>2324</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/5.726791</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>S. H.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>B. C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Vision transformer for small-size datasets</article-title>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2112.13492</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Deep lac: Deep localization, alignment and classification for fine-grained recognition</article-title>,&#x201d; in <conf-name>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>, <conf-loc>Boston, MA, USA</conf-loc>: <publisher-name>IEEE</publisher-name>. <fpage>1666</fpage>&#x2013;<lpage>1674</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2015.7298775</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ouppaphan</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Corn disease identification from leaf images using convolutional neural networks</article-title>,&#x201d; in <conf-name>Proceedings of the 2017 21st International Computer Science and Engineering Conference (ICSEC)</conf-name> (<publisher-name>IEEE</publisher-name>), <publisher-loc>Bangkok, Thailand</publisher-loc>. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICSEC.2017.8443919</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Plant and animal species recognition based on dynamic vision transformer architecture</article-title>. <source>Remote Sens.</source> <volume>14</volume>, <elocation-id>5242</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14205242</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sanju</surname> <given-names>S.</given-names>
</name>
<name>
<surname>VelammaL</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>An automated detection and classification of plant diseases from the leaves using image processing and machine learning techniques: A state-of-the-art review</article-title>. <source>Ann. Rom. Soc Cell Biol.</source> <volume>25</volume>, <fpage>15933</fpage>&#x2013;<lpage>15950</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fine-grained recognition: Accounting for subtle differences between similar classes</article-title>. <conf-name>Proceedings of the AAAI conference on artificial intelligence</conf-name>. (<publisher-loc>New York, USA</publisher-loc>: <publisher-name>AAAI</publisher-name>) <volume>34</volume> (<issue>07</issue>), <fpage>12047</fpage>&#x2013;<lpage>12054</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1609/aaai.v34i07.6882</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shazeer</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Parmar</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Uszkoreit</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Gomez</surname> <given-names>A. N.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). &#x201c;<article-title>Attention is all you need</article-title>,&#x201d; in <conf-name>In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS&#x2019;17)</conf-name> (<publisher-name>Curran Associates Inc.</publisher-name>), <publisher-loc>Long Beach California USA</publisher-loc>. <fpage>6000</fpage>&#x2013;<lpage>6010</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition</article-title>,&#x201d; in <conf-name>Proceedings of the Neural Information Processing Systems</conf-name>. <publisher-name>NeurIPS</publisher-name>, virtual</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>T-cnn: Trilinear convolutional neural networks model for visual detection of plant diseases</article-title>. <source>Comput. Electron. Agric.</source> <volume>190</volume>, <fpage>106468</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compag.2021.106468</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lyu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>A skin disease classification model based on densenet and convnext fusion</article-title>. <source>Electronics</source> <volume>12</volume>, <elocation-id>438</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/electronics12020438</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Learning to navigate for fine-grained classification</article-title>,&#x201d; in <source>Computer vision &#x2013; ECCV 2018</source>, vol. <volume>11218</volume> . Ed. <person-group person-group-type="editor">
<name>
<surname>Ferrari</surname> <given-names>V.</given-names>
</name>
</person-group> (<publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>). H. M. S. C. W. Y. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-01264-9\26</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Donahue</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R. B.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Part-based R-CNNs for fine-grained category detection</article-title>. <source>Computer Vision--ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13</source>. (<publisher-loc>Zurich, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>834</fpage>&#x2013;<lpage>849</lpage>.  doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-319-10590-1_54</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Cucumber disease recognition based on global-local singular value decomposition</article-title>. <source>Neurocomputing</source> <volume>205</volume>, <fpage>341</fpage>&#x2013;<lpage>348</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.neucom.2016.04.034</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Qiao</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Joint face detection and alignment using multitask cascaded convolutional networks</article-title>. <source>IEEE Signal Process. Lett.</source> <volume>23</volume>, <fpage>1499</fpage>&#x2013;<lpage>1503</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LSP.2016.2603342</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>