<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2023.1196634</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Crop classification in high-resolution remote sensing images based on multi-scale feature fusion semantic segmentation model</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Lu</surname>
<given-names>Tingyu</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Gao</surname>
<given-names>Meixiang</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2264206"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Lei</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>College of Geographical Sciences, Harbin Normal University</institution>, <addr-line>Harbin</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Geography and Spatial Information Techniques, Ningbo University</institution>, <addr-line>Ningbo</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>School of Civil and Environmental Engineering and Geography Science, Ningbo University</institution>, <addr-line>Ningbo</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Department of Surveying Engineering, Heilongjiang Institute of Technology</institution>, <addr-line>Harbin</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Daobilige Su, China Agricultural University, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Zhaoyu Zhai, Nanjing Agricultural University, China; Xiaolei Zhang, Nanjing Agricultural University, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Meixiang Gao, <email xlink:href="mailto:gmx1002@sina.com">gmx1002@sina.com</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1196634</elocation-id>
<history>
<date date-type="received">
<day>23</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Lu, Gao and Wang</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Lu, Gao and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The great success of deep learning in the field of computer vision provides a development opportunity for intelligent information extraction of remote sensing images. In the field of agriculture, a large number of deep convolutional neural networks have been applied to crop spatial distribution recognition. In this paper, crop mapping is defined as a semantic segmentation problem, and a multi-scale feature fusion semantic segmentation model MSSNet is proposed for crop recognition, aiming at the key problem that multi-scale neural networks can learn multiple features under different sensitivity fields to improve classification accuracy and fine-grained image classification. Firstly, the network uses multi-branch asymmetric convolution and dilated convolution. Each branch concatenates conventional convolution with convolution nuclei of different sizes with dilated convolution with different expansion coefficients. Then, the features extracted from each branch are spliced to achieve multi-scale feature fusion. Finally, a skip connection is used to combine low-level features from the shallow network with abstract features from the deep network to further enrich the semantic information. In the experiment of crop classification using Sentinel-2 remote sensing image, it was found that the method made full use of spectral and spatial characteristics of crop, achieved good recognition effect. The output crop classification mapping was better in plot segmentation and edge characterization of ground objects. This study can provide a good reference for high-precision crop mapping and field plot extraction, and at the same time, avoid excessive data acquisition and processing.</p>
</abstract>
<kwd-group>
<kwd>remote sensing</kwd>
<kwd>crop classification</kwd>
<kwd>deep learning</kwd>
<kwd>convolutional neural network</kwd>
<kwd>multi-scale feature</kwd>
</kwd-group>
<counts>
<fig-count count="16"/>
<table-count count="4"/>
<equation-count count="2"/>
<ref-count count="44"/>
<page-count count="16"/>
<word-count count="6005"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Technical Advances in Plant Science</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>With the rapid development of remote sensing technology, the quality and updating speed of remote sensing data have been significantly improved, and multi-source remote sensing data has been widely applied in agriculture, forestry, Marine, environmental protection and other fields (<xref ref-type="bibr" rid="B32">Sun, 2020</xref>). Remote sensing image classification has always been a very active research topic in the application of remote sensing technology, which refers to the use of remote sensing data to make land use or land cover maps (<xref ref-type="bibr" rid="B26">Luo, 2011</xref>).</p>
<p>At present, the application based on artificial intelligence model and algorithm has become very common. Machine learning and deep learning are the methods to realize artificial intelligence. With the continuous innovation of deep learning, the field of computer vision has developed rapidly in the past few years and made breakthroughs constantly (<xref ref-type="bibr" rid="B12">Hopfield, 1982</xref>; <xref ref-type="bibr" rid="B3">Bengio and Delalleau, 2011</xref>). The development of computer vision is driven by the innovation of algorithms, the increase in the amount of visual data and the improvement of computing power. In image classification, target detection and location, image segmentation and other tasks, deep learning algorithms surpass traditional statistical methods on a large number of benchmarks, and even exceed human beings in image and target recognition (<xref ref-type="bibr" rid="B2">Bengio, 2009</xref>; <xref ref-type="bibr" rid="B17">LeCun et&#xa0;al., 2015</xref>).</p>
<p>In the field of agriculture, using remote sensing data to classify crops is an important research content. Timely and accurate acquisition of spatial distribution and planting area of crops by utilizing spatio-temporal scale advantages of remote sensing images is of great significance for ensuring food security and promoting sustainable agricultural development (<xref ref-type="bibr" rid="B16">Kussul et&#xa0;al., 2017</xref>). High resolution remote sensing image has the characteristics of high background complexity, rich detail information and diversified spatial structure, so the classification accuracy is often low when the traditional machine learning classification algorithm is applied to the classification of high resolution remote sensing image. In recent years, many researchers have tried to build semantic segmentation network through deep learning algorithm and applied it in pixel-level ground object fine classification. Remote sensing image classification based on artificial neural network has become a development trend (<xref ref-type="bibr" rid="B44">Zhong et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B29">Rustowicz et&#xa0;al., 2019</xref>).</p>
<p>For traditional machine learning models and popular deep learning models, the architecture design of the model itself and super-parameter fine-tuning determine the feature extraction capability of the model, and the strength of the feature extraction capability is a decisive factor affecting the model performance. The high efficiency of deep learning algorithm is reflected in its independent dependence on highly complex feature engineering, and its high performance is reflected in its powerful feature extraction ability. Therefore, how to enhance the feature extraction ability of the algorithm is the essential problem of deep learning model architecture design (<xref ref-type="bibr" rid="B15">Kawaguchi et&#xa0;al., 2017</xref>).</p>
<p>Multi-scale refers to the sampling processing of signals with different granularity. In deep learning algorithm, it means that the model learns different features at different scales, such as fine features and rough features, as well as the combination of the two features. This method has been proved to effectively improve the performance of the model. The idea of multi-scale feature fusion technology is to extract image features under different sensory fields. At present, there are mainly two types of multi-scale feature network design paradigms, one is skip connection architecture based on deep convolutional neural network (DCNN), such as UNet, VNet (<xref ref-type="bibr" rid="B27">Milletari et&#xa0;al., 2016</xref>), FCN series (<xref ref-type="bibr" rid="B25">Long et&#xa0;al., 2015</xref>), RefineNet (<xref ref-type="bibr" rid="B20">Lin et&#xa0;al., 2017</xref>), etc. This kind of network is characterized by the use of pre-training weights or DCNN (represented by residual network) in the coding stage, and the acceptance of low-level features through skip connections in the decoding stage, and the fusion of low-level and abstract features, so as to achieve multi-scale feature extraction. The other type adopts parallel multi-branch structure design, such as PSPNet (<xref ref-type="bibr" rid="B43">Zhao et&#xa0;al., 2017</xref>), GoogleNet, DeepLab series (<xref ref-type="bibr" rid="B6">Chen et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B5">Chen et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B7">Chen et&#xa0;al., 2018</xref>), etc., which is characterized by using hollow convolution or convolution kernel of various sizes to extract features from different receptor fields, and finally merging multiple channels to form multi-scale features.</p>
<p>Multi-scale feature fusion network is widely used in computer vision tasks such as target detection and image classification. <xref ref-type="bibr" rid="B34">Varadarajan et&#xa0;al. (2021)</xref> designed an object detection network composed of 22 convolution layers. By using multi-scale feature fusion technology, the network can well identify objects of different sizes and shapes from images. <xref ref-type="bibr" rid="B30">Sang et&#xa0;al. (2022)</xref> proposed a target tracking network MTTNet based on multi-scale global retrieval and spatial-temporal consistency matching, and used spatial pyramid pool to solve the problem of multi-scale feature extraction. The experimental results show that the network has stable performance and can effectively perform long-term target tracking tasks. <xref ref-type="bibr" rid="B37">Wu et&#xa0;al. (2022)</xref> integrated the hierarchical pyramid pooling module into the full convolutional neural network, and the improved network was able to collect multi-scale context information. The good performance of the network was verified in the robot object grabbing experiment. In the study of fine-grained image classification, Liu et&#xa0;al (<xref ref-type="bibr" rid="B24">Liu et&#xa0;al., 2021</xref>). fused the attention module with multi-scale feature expression in order to distinguish the subtle differences between the subcategories of the main category. The improved network can learn a list of accurate feature maps. <xref ref-type="bibr" rid="B38">Xie et&#xa0;al. (0000)</xref> proposed a multi-scale densely connected convolutional neural network MS-DenseNet when studying hyperspectral image classification. By learning multi-scale patches around each pixel, they made full use of multi-scale information. <xref ref-type="bibr" rid="B36">Wang et&#xa0;al. (2022)</xref> proposed a multi-scale convolutional neural network point cloud filtering algorithm based on attention mechanism to solve the problem of low accuracy of traditional filtering methods when processing lidar data, combining channel and spatial attention module with multi-scale convolution kernel. The lidar point cloud feature maps output by the algorithm at different scales can adjust the weights of each channel layer and different spatial regions adaptively, so that the network pays extra attention to important information, thus improving the classification performance of the model.</p>
<p>In recent years, many deep semantic segmentation networks using multi-scale feature fusion methods have been applied to pixel level classification of remote sensing images or scene classification of remote sensing images. A large number of research results show that it is an effective method to obtain better image classification results. When conducting large-scale land cover classification, <xref ref-type="bibr" rid="B10">Gao et&#xa0;al. (2021)</xref> found that the traditional sliding window convolutional neural network has a large computational overhead and the classification results are not precise enough. To solve this problem, a new object-oriented deep learning framework was proposed, which uses residual networks to learn features on different adjacent scales and achieves a balance between weak semantics and strong features. When studying the classification of complex remote sensing scenes, <xref ref-type="bibr" rid="B4">Bi et&#xa0;al. (2021)</xref> found that the complex spatial arrangement and object size changes in large-scale aerial images were challenging for classification models. In order to enhance the feature expression ability of remote sensing scenes, a multi-scale expanded convolution operator was designed. To solve the &#x201c;small sample&#x201d; problem of hyperspectral image classification, <xref ref-type="bibr" rid="B11">Gong et&#xa0;al. (2021)</xref> designed a lightweight multi-scale attention pyramid pooling network MSPN, whose core components included a multi-scale three-dimensional CNN module and a squeezing excitation attention module. The network learns and fuses deeper spatial spectral features with fewer training samples, and verifies MSPN&#x2019;s good performance on publicly available hyperspectral data sets. <xref ref-type="bibr" rid="B19">Liao et&#xa0;al. (2022)</xref> developed multi-scale object-driven convolutional neural network multi-OCNNs, which can capture the depth and context information contained in the reference samples, and has achieved good results in land cover classification based on multi-source high-resolution images such as SPOT-6, Gaofen-2 and ALOS.</p>
<p>In summary, a large number of deep learning models represented by convolutional neural networks have been applied to intelligent information extraction tasks of remote sensing images. However, there are few researches on exploiting the potential of multi-spectral remote sensing in crop mapping by using multi-scale feature fusion semantic segmentation model. Models that use pre-trained networks combined with multiple jump connections to achieve feature fusion tend to be deeper, resulting in larger model parameters and a significant increase in computational overhead. Skipping connections can only alleviate the problem of single feature size to a certain extent, and the requirements of fine-grained segmentation cannot be met in the initial stage of training, resulting in insufficient segmentation results. In order to solve this problem, this paper proposes a semantic segmentation model with residual network as the backbone and receiving field module for multi-scale and multi-scale feature fusion. The performance of the model was evaluated in the experiment of crop classification based on Sentinel-2 high-resolution remote sensing image.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Multi-scale feature extraction method based on deep convolutional neural network</title>
<p>On the premise of effectively alleviating the problem of gradient disappearance, the residuals network (ResNet) improves the performance of the model by adding considerable depth. In addition to the common residuals network of 18 layers, 34 layers and 50 layers, there are ResNet-101 and ResNet-152 at a deeper level. A modest increase in the depth of the network is beneficial to the performance of the model, and a large number of experiments have shown that changing the width of the network can achieve the same purpose. Multi-scale feature extraction modules based on parallel multi-branch structure design have been proposed one after another. In this chapter, several important convolutional modules are described in detail.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Inception</title>
<p>The design concept of Inception is to use convolution kernels of different sizes to realize the perception of multi-level features, and finally fusion to obtain better representation of images. Inception module is the core component of the GoogleNet network. From Inception V1 to Inception V4, each version is the optimization of the previous version, with the number of parameters decreasing and the running speed and accuracy gradually improving. The different versions of the Inception model structure are shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Different versions of the Inception module architecture schematic. <bold>(A)</bold> Inception V1; <bold>(B)</bold> Inception V2; <bold>(C)</bold> Inception V3; <bold>(D)</bold> Inception V4.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g001.tif"/>
</fig>
<p>Inception module realizes multi-scale feature space superposition through parallel multi-branch operations and uses intensive operations to maintain model sparsity. Xception to Inception - V3 is improved, and put forward the depth of Separable convolution (Depthwise Separable Convolutions), Inception - V3 will channel is divided into four groups respectively carry out 1 x 1 convolution computation, Xception performs a 1 &#xd7; 1 convolution calculation for each channel&#x2019;s feature graph and concatenates the feature forces. Completely decouple channel and spatial dependencies. Xception has the same number of parameters as Inception-V3, but with better performance and more efficient use of network parameters, as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref> for its structure.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Perform the 3 &#xd7; 3 convolution on each channel of the 1&#xd7;1 convolution.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g002.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Atrous spatial pyramid pooling</title>
<p>Atrous Spatial Pyramid Pooling (ASPP) uses dilated convolution with different dilation coefficients to extract multi-scale features. Dilated convolutions (<xref ref-type="bibr" rid="B40">Yu and Koltun, 2016</xref>) add dilation into the standard convolutions to increase the field of perception (See <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>).</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Atrous convolution increases the receptive field without losing information. <bold>(A)</bold> atrous_rate=1; <bold>(B)</bold> atrous_rate=2; <bold>(C)</bold> atrous_rate=4.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g003.tif"/>
</fig>
<p>The calculation process of dilated convolution is shown in the following formula, where <italic>H<sub>in</sub>
</italic> and <italic>H<sub>out</sub>
</italic> respectively represent the height of the input and output feature graphs, and <italic>W<sub>in</sub>
</italic> and <italic>W<sub>out</sub>
</italic> respectively represent the width of the input and output feature graphs.</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext>padding</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>atrous</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>kernel</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext>padding</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>atrous</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>kernel</mml:mtext>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo stretchy="false">[</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mfrac>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Some semantic segmentation algorithms based on full convolutional neural networks, such as FCN-8S and FCN-16S, need continuous up-sampling in order to achieve the same resolution of input and output images. However, this process cannot recover the loss of detail information caused by previous pooling. Dilated convolution can reduce such loss to a certain extent.</p>
<p>ASPP module first appeared in the semantic segmentation algorithm DeepLab V2, consisting of 3 &#xd7; 3 convolution of four different expansion coefficients. Subsequently, ASPP was applied to many image classifications tasks as an independent module (<xref ref-type="bibr" rid="B42">Yuan et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B21">Liu et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B28">Pedrayes et&#xa0;al., 2021</xref>). The branch structure inside ASPP is not invariable, and designers often adopt different parameter configurations according to different application scenarios. <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref> shows an ASPP Block containing four branches. By setting four different dilation, the module is capable of feature extraction from four different scales.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>An ASPP Block capable of extracting four scale features.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g004.tif"/>
</fig>
<p>ASPP uses filters with multiple sampling rates and effective field of view to detect the incoming convolutional feature map, so as to capture objects and image context information at multiple scales, keep image resolution unchanged, obtain more intensive feature response, and better restore the details of the original image.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Receptive field block</title>
<p>Receptive Field Block (RFB) refers to the design concept of Inception. In each branch structure, conventional convolution of convolution kernel of specific size is first used, then dilated convolution is added, and multi-scale features are extracted by group convolution. Dilated convolution increases receptive field. The RFB module considers the relationship between the receptive field center and the target region to enhance the feature recognition and robustness (<xref ref-type="bibr" rid="B23">Liu et&#xa0;al., 2018</xref>). The structure of RFB is shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>. The extracted multi-scale features are fused and input to the next layer by adding 1 &#xd7; 1 convolution and identity shortcut connection. RFB is a lightweight feature extraction module, which can be conveniently configured in convolutional neural networks. Especially in some target detection tasks, RFB has brought significant performance gains to detection networks (<xref ref-type="bibr" rid="B18">Li et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B22">Liu et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B41">Yuan et&#xa0;al., 2021</xref>).</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>The architecture of RFB.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g005.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Multi-scale feature fusion network-MSSNet</title>
<sec id="s3_1">
<label>3.1</label>
<title>Study area</title>
<p>The study area is Yunshan Farm and 850 Farm located in Hulin City, Heilongjiang Province. The longitude range is 132&#xb0;35 &#x2018;21 &#x201c;~132&#xb0;51&#x2019; 46&#x201d; E, and the latitude range is 45&#xb0;47 &#x2018;21 &#x201c;~45&#xb0;58&#x2019; 11&#x201d; N. Located in the famous Sanjiang Plain, this area is a temperate continental monsoon climate with an effective accumulated temperature of 2501&#xb0;C. 80% of the cultivated land is low-wet land with fertile soil and abundant water resources. It mainly grows corn, rice and soybeans, and is an important commercial grain base in China. The location and scope of the research area are shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Study area with its RGB image composite derived from Sentinel-2 imagery.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g006.tif"/>
</fig>
<p>The remote sensing data uses the high-resolution multi-spectral image of Sentinel-2 satellite developed by ESA. Sentinel-2 is divided into 2A and 2B satellites with a revisit period of 5 days. Sentinel-2 can cover 13 spectral bands and provide multi-spectral images with spatial resolution of 10 meters, 20 meters and 60 meters (<xref ref-type="bibr" rid="B35">Verrelst et&#xa0;al., 2012</xref>). Widely used in agricultural resources monitoring and crop yield estimation, geological survey, land use dynamic monitoring and other fields. Sentinel-2 remote sensing data used in this paper is Level-1C data that was imaged on July 28, 2020, and is derived from Sentinel Hub. Sen2cor 2.11 is used to preprocess the data. Firstly, radiometric calibration and atmospheric correction were carried out for multi-spectral images to eliminate radiation errors caused by atmospheric scattering, etc. Then SNAP 9.0 and ENVI 5.3 platforms were used to generate reflectance Level 2A data at the bottom of the atmosphere. Finally, the resolution of the resampling remote sensing image was 10 meters, and the study area was about 423 square kilometers.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Features extraction and training set construction</title>
<p>In this study, based on Sentinel-2 multispectral images, we designed 12 features (<xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>), including blue (band 2), green (band 3), red (band 4), visible light and near infrared (band 5-band 8a), short wave and infrared (band 11-band 12). The other two features are Normalized Differential Vegetation Index (NDVI) and Enhanced Vegetation Index (EVI). <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref> shows the calculation methods of the two planting cover index data. NIR in the formula is band 8. Among different vegetation indexes, NDVI and EVI are important measurement parameters of surface vegetation cover and vegetation growth (<xref ref-type="bibr" rid="B14">Immitzer et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B13">Huang et&#xa0;al., 2019</xref>), which have been proved to be helpful to improve the classification accuracy (<xref ref-type="bibr" rid="B9">Ferrant et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B31">Silveira et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B1">Belgiu and Csillik, 2018</xref>). The average reflectance spectra of each type of crop are shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>. Principal Components Analysis (PCA) is used to extract the first three principal components of the image after principal component transformation as characteristic variables to participate in the classification.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Features designed in this study.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Band</th>
<th valign="middle" align="center">Description</th>
<th valign="middle" align="center">Central wavelength(nm)</th>
<th valign="middle" align="center">Spatial resolution(m)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center">band 2</td>
<td valign="top" align="center">Blue</td>
<td valign="top" align="center">490</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 3</td>
<td valign="top" align="center">Green</td>
<td valign="top" align="center">560</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 4</td>
<td valign="top" align="center">Red</td>
<td valign="top" align="center">665</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 5</td>
<td valign="top" align="center">Vegetation Red Edge</td>
<td valign="top" align="center">705</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 6</td>
<td valign="top" align="center">Vegetation Red Edge</td>
<td valign="top" align="center">740</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 7</td>
<td valign="top" align="center">Vegetation Red Edge</td>
<td valign="top" align="center">783</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 8</td>
<td valign="top" align="center">NIR</td>
<td valign="top" align="center">842</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 8a</td>
<td valign="top" align="center">Vegetation Red Edge</td>
<td valign="top" align="center">865</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 11</td>
<td valign="top" align="center">SWIR</td>
<td valign="top" align="center">1610</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="top" align="center">band 12</td>
<td valign="top" align="center">SWIR</td>
<td valign="top" align="center">2190</td>
<td valign="top" align="center">10</td>
</tr>
<tr>
<td valign="middle" align="center">NDVI</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mi>I</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mi>I</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">10</td>
</tr>
<tr>
<td valign="middle" align="center">EVI</td>
<td valign="top" align="center">
<inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:mn>2.5</mml:mn>
<mml:mo>*</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mi>I</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mi>I</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>6</mml:mn>
<mml:mi>R</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>7.5</mml:mn>
<mml:mi>B</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="center">&#x2013;</td>
<td valign="middle" align="center">10</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>Mean reflectance spectral of each crop.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g007.tif"/>
</fig>
<p>The training set consists of 216 plots, including 50 plots of corn field, 86 plots of rice field and 80 plots of soybean field, as shown in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8</bold>
</xref>. The marks show the geographical locations of the real land cover sample areas extracted from the study area. The data of these plots were obtained through agricultural census and field survey, which collected a series of ground survey data. Including precise GPS coordinates of plots and crop types, pixel (sample) is the basic unit used for classification. <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> lists crop types and the number of each type of sample in the training set.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>The ground truth used for model training and the effect of partial sample area zoomed.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g008.tif"/>
</fig>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>The number of training dataset per crop class.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Class</th>
<th valign="middle" align="center">Label color</th>
<th valign="top" align="center">Samples size</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">Corn</td>
<td valign="middle" align="left">&#x25ac;</td>
<td valign="middle" align="right">94842</td>
</tr>
<tr>
<td valign="middle" align="left">Rice</td>
<td valign="middle" align="left">&#x25ac;</td>
<td valign="middle" align="right">76964</td>
</tr>
<tr>
<td valign="middle" align="left">Soybean</td>
<td valign="middle" align="left">&#x25ac;</td>
<td valign="middle" align="right">82615</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Comparison of backbone</title>
<p>Backbone is a model containing visual representation capability generated by pre-training upstream data, which is a part of deep learning model. Therefore, Backbone&#x2019;s feature extraction capability directly affects the performance of the algorithm. This paper selects three different backbone networks, ResNet18, VGG19 and ResNet50, as feature extraction models of MSSNet. In the first experiment, T-distributed Random neighbor Embedding (t-SNE) was used to analyze the crop-specific spatial heterogeneity of the original Sentinel-2 data and the data processed by the deep learning model, so as to measure the feature extraction capability of different backbone. Then, 2000 samples each of crop are randomly selected, the original features and extracted features corresponding to these samples are nonlinearly projected to a 2-D plane for visualization using t-SNE. As shown in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9</bold>
</xref>, it is shown that the separability of features extracted by different backbone is significantly better than that of original features among different crop categories. In addition, compared with ResNet18 and VGG19, features extracted by ResNet50 are more separable and samples of the same crop category are more clustered. Therefore, ResNet50 is selected as the backbone network in this paper.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>Two-dimensional plane projection of high-dimensional features learned by different backbone based on t-SNE. <bold>(A)</bold> original feature. <bold>(B)</bold> features extracted by ResNet18. <bold>(C)</bold> features extracted by VGG19. <bold>(D)</bold> features extracted by ResNet50.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g009.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>MSSNet architecture</title>
<p>Inspired by the above multi-scale feature extraction modules, we proposed a semantic Segmentation network MSSNet (Multi Scale Segmentation Net) for fine classification of crops in agricultural areas. This network is based on residual network, receptive field module and skip connection. The architecture is shown in <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref>, and the core contents are summarized as follows:</p>
<list list-type="simple">
<list-item>
<p>&#x2022; The pre-trained residual network (ResNet50) is used as the backbone network to receive global visual features.</p>
</list-item>
<list-item>
<p>&#x2022; Embedded receptive field module (RFB) for multi-scale feature extraction and integration.</p>
</list-item>
<list-item>
<p>&#x2022; Use skip connection to concatenate low-level and high-level features with the same resolution.</p>
</list-item>
</list>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>The architecture of MSSNet, ResNet50 and RFB are core components.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g010.tif"/>
</fig>
<p>The Input of MSSNet model is set as 256 &#xd7; 256 &#xd7; 3, and the backbone network ResNet50 is composed of 5 stages. We select the 39th convolutional layer (activation_39) located at Stage 3 as the output layer, and the size of the feature map is 16 &#xd7; 16 &#xd7; 256. The RFB module consists of four branches, and the identity mapping (arc) directly outputs 16 &#xd7; 16 &#xd7; 256. The second branch consists of a 3 &#xd7; 3 conventional convolution and a dilated convolution with a dilation coefficient of 1, and the output feature map is 16 &#xd7; 16 &#xd7; 256. The third and fourth branches are both composed of two 3 &#xd7; 3 conventional convolution and a dilated convolution. The dilation coefficients of the dilated convolution are different. Since the dilated convolution does not change the parameter number, the output of both branches is 16 &#xd7; 16 &#xd7; 256. RFB uses a 3 &#xd7; 3 convolution to fuse the features extracted from the second to the fourth branches, and outputs the feature graph 16 &#xd7; 16 &#xd7; 768. At this time, the feature graph is added to the output of the identity map, and the final output of RFB is 16 &#xd7; 16 &#xd7; 768. After that, the feature dimension is reduced to 512, and the output feature graph is 16&#xd7;16&#xd7;512 for three consecutive 3 &#xd7; 3 convolution. After the first quadruple up-sampling operation, the spatial resolution of the image is expanded to four times the original one, and the feature graph is 64 &#xd7; 64 &#xd7; 512. Considering the importance of course-scale features for semantic segmentation of fine-grained images, we fused low-level features with high-level features, and used a skip connection to achieve this in MSSNet. The convolutional layer activation_9 was located at Stage 0 of ResNet50, and its output was 64 &#xd7; 64 &#xd7; 64. The skip connection fuses 64 &#xd7; 64 &#xd7; 512 of the first up-sampling feature with 64 &#xd7; 64 &#xd7; 64 &#xd7; 64 of the lower-level features, and outputs 64 &#xd7; 64 &#xd7; 576. After three 3 &#xd7; 3 convolutions, the feature dimension is reduced to 256, and the spatial resolution of the image is restored to 256 &#xd7; 256 by a second quad up-sampling. Finally, the probability that the output pixels of four 3 &#xd7; 3 convolution layers and one convolutional layer using Softmax activation function belong to a certain class is obtained.</p>
<p>In this paper, Python language is used to implement the MSSNet semantic segmentation network based on Keras API (Tensorflow as the back-end), and the network is used to mine spatial features and spectral features from multi-spectral data sets to achieve semantic segmentation. As shown in <xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref>, the reconstructed multispectral data was reduced from 12 features to 3 features by PCA, and then pixel-level classification results were output by ResNet50, receptive field module and continuous upsampling.</p>
<fig id="f11" position="float">
<label>Figure&#xa0;11</label>
<caption>
<p>Semantic segmentation diagram of MSSNet, intermediate feature mapping represents features extracted at different levels.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g011.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment setting</title>
<sec id="s4_1">
<label>4.1</label>
<title>Classification results and accuracy evaluation</title>
<p>In this paper, four deep learning semantic segmentation networks, including MSSNet, are applied to this classification task, and the experimental setup is shown in <xref ref-type="fig" rid="f12">
<bold>Figure&#xa0;12</bold>
</xref>. UNet++ is a deeply supervised semantic segmentation network where subnetworks of encoder and decoder are connected to each other through a series of nested dense jump paths, and PSPNet and DeepLab V2 are deep learning models for intensive prediction tasks. In the process of model training, the hyperparameters are also configured in the same configuration. The optimizer Adam has a learing rate of 0.001, iteration times (epoch) of 120, and batch size of 16. Input image resolution is set to 256 &#xd7; 256, channel number <italic>C</italic>&#xa0;= 3, and is composed of the first three components after principal component transformation of Sentinel-2 image. Therefore, the input image size of the model is (256,256,3). In order to meet the architecture design of MSSNet, samples need to be extracted from image data. After rearrangement and normalization, the data enhancement strategy was used in the training process.</p>
<fig id="f12" position="float">
<label>Figure&#xa0;12</label>
<caption>
<p>Experimental design.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g012.tif"/>
</fig>
<p>In order to quantitatively and accurately assess the influence of different classifiers on crop extraction accuracy, an area of about 9 square kilometers in the research area is selected as the test set, as shown in <xref ref-type="fig" rid="f13">
<bold>Figure&#xa0;13</bold>
</xref>. The marked part is the test set, and the red box is the research area. 61 plots are marked in the test set, including 24 corn fields, 11 paddy fields and 26 soybean fields. The corresponding test samples are 35,368, 9711 and 24,933 respectively.</p>
<fig id="f13" position="float">
<label>Figure&#xa0;13</label>
<caption>
<p>The test set is located in the study area and covers only one training sample.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g013.tif"/>
</fig>
<p>We evaluate the performance of the classifier using Mean Intersection over Union (MIoU), overall accuracy (OA), and the Kappa coefficient shown by the confusion matrix, which is the most commonly used metric for semantic segmentation tasks. MIoU calculates the IOU (the intersection of the real label and the predicted result) for each class separately, and then averages the IOU for all classes. MIoU is the standard accuracy measure. Among them, the overall accuracy can reflect the overall performance of the classifier. Each classification algorithm is trained five times repeatedly, that is, the same classification algorithm will make five predictions on the test set. The combined statistical results of the repeatedly generated confusion matrix are shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. The crop classification diagram generated by different classification algorithms is shown in <xref ref-type="fig" rid="f14">
<bold>Figure&#xa0;14</bold>
</xref>.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Classification accuracies of different algorithms, bold values show the best performance.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Class</th>
<th valign="middle" align="center">UNet++<break/>Mean &#xb1; SD</th>
<th valign="middle" align="center">PSPNet<break/>Mean &#xb1; SD</th>
<th valign="middle" align="center">DeepLab V2<break/>Mean &#xb1; SD</th>
<th valign="middle" align="center">MSSNet<break/>Mean &#xb1; SD</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">Corn</td>
<td valign="middle" align="center">91.13 &#xb1; 0.69%</td>
<td valign="middle" align="center">91.64 &#xb1; 1.26%</td>
<td valign="middle" align="center">
<bold>92.55 &#xb1; 2.36%</bold>
</td>
<td valign="middle" align="center">92.41 &#xb1; 1.68%</td>
</tr>
<tr>
<td valign="middle" align="left">Rice</td>
<td valign="middle" align="center">90.26 &#xb1; 1.42%</td>
<td valign="middle" align="center">91.06 &#xb1; 1.88%</td>
<td valign="middle" align="center">91.38 &#xb1; 2.35%</td>
<td valign="middle" align="center">
<bold>91.58 &#xb1; 2.46%</bold>
</td>
</tr>
<tr>
<td valign="middle" align="left">Soybean</td>
<td valign="middle" align="center">80.41 &#xb1; 2.31%</td>
<td valign="middle" align="center">86.44 &#xb1; 1.95%</td>
<td valign="middle" align="center">79.28 &#xb1; 2.48%</td>
<td valign="middle" align="center">
<bold>88.19 &#xb1; 4.30%</bold>
</td>
</tr>
<tr>
<td valign="middle" align="left">OA (%)</td>
<td valign="middle" align="center">82.68 &#xb1; 1.46</td>
<td valign="middle" align="center">86.90 &#xb1; 1.15</td>
<td valign="middle" align="center">89.56 &#xb1; 1.89</td>
<td valign="middle" align="center">
<bold>90.77 &#xb1; 1.31</bold>
</td>
</tr>
<tr>
<td valign="middle" align="left">MIoU&#xd7;100</td>
<td valign="middle" align="center">74.24 &#xb1; 0.60</td>
<td valign="middle" align="center">74.78 &#xb1; 0.37</td>
<td valign="middle" align="center">75.82 &#xb1; 2.91</td>
<td valign="middle" align="center">
<bold>76.59 &#xb1; 0.21</bold>
</td>
</tr>
<tr>
<td valign="middle" align="left">Kappa &#xd7; 100</td>
<td valign="middle" align="center">74.96 &#xb1; 2.16</td>
<td valign="middle" align="center">81.13 &#xb1; 1.21</td>
<td valign="middle" align="center">79.85 &#xb1; 2.78</td>
<td valign="middle" align="center">
<bold>83.53 &#xb1; 2.29</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values mean the highest classification accuracy.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="f14" position="float">
<label>Figure&#xa0;14</label>
<caption>
<p>The segmentation result of each algorithm. <bold>(A)</bold> Sentinel-2 MSI image; <bold>(B)</bold> The ground truth; <bold>(C)</bold> UNet++; <bold>(D)</bold> PSPNet; <bold>(E)</bold> DeepLab V2; <bold>(F)</bold> MSSNet.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g014.tif"/>
</fig>
<p>Overall accuracy (OA) refers to the ratio between the total number of correctly classified samples of all categories and the total ground truth value. As can be seen from <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>, among the four semantic segmentation algorithms, the multi-scale feature fusion network MSSNet proposed by us achieves higher classification accuracy. The overall accuracy is 8%, 3.87% and 1.21% higher than UNet++, PSPNet and DeepLab V2, respectively. The average classification accuracy of corn and rice reached 90%, but the classification accuracy of soybean was relatively low, and there was obvious misclassification between corn and soybean. MSSNet has the highest MIoU, which means that the model has the best segmentation for various categories, in addition, through qualitative analysis of the classification map, it can be seen that MSSNet is obviously superior to the other three algorithms in the detail characterization ability of image segmentation. The boundary of the block is clearer, the classification results of the block interior are more continuous (blue circular area in <xref ref-type="fig" rid="f14">
<bold>Figure&#xa0;14</bold>
</xref>, and it can extract the small block area more accurately (blue oval area in <xref ref-type="fig" rid="f14">
<bold>Figure&#xa0;14</bold>
</xref>). These performance gains are due to the multi-scale feature extraction and multi-level feature fusion capabilities of the multi-RFB module. Specifically, the convolutional kernel and cavity convolution of different sizes of the RFB module extract rich multi-scale features. The classification results of UNet++ and PSPNet are relatively rough and significantly weaker than the other two algorithms in terms of image details. There are large pixel blocks on the classification map, and the segmentation results cannot restore the details of the input image, which is also the main reason for their low classification accuracy.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Traditional machine learning classification algorithm</title>
<p>Traditional machine learning classifiers such as random forest (RF) and support vector machine (SVM) have been widely applied to classification tasks with their good performance. In this study, we compared four traditional machine learning classification algorithms on the same data set, which are RF, SVM, kernel SVM and XGBoost. RF adopts an integration algorithm with high accuracy and can maintain accuracy even if there is a large amount of missing data. Over-fitting does not occur easily owing to randomly selected samples&#x2019; characteristics and some features&#x2019; random extraction in the training process (<xref ref-type="bibr" rid="B8">Cutler et&#xa0;al., 2004</xref>). Support vector machine (SVM), first proposed by Corinna Cortes and Vapnik et&#xa0;al. in 1995, is a statistical theory specifically for small samples (<xref ref-type="bibr" rid="B33">Vapnik, 1995</xref>). Its unique advantage lies in dealing with small samples, nonlinear, and high-dimensional data problems, and many scholars have applied it to remote sensing image classification tasks.</p>
<p>XGBoost is an open source machine learning project, which effectively implements GBDT algorithm with a lot of improvements, and has a wide range of applications in computer vision tasks such as image classification and object extraction. We use the &#x201c;random search&#x201d; method to optimize the main hyperparameters of the model, and select the best combination of hyperparameters from the candidate values according to the classification accuracy of each model on the test set. The optimization results are shown in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>. The bold characters in the candidate values represent the optimal parameters. Among these traditional machine learning classification algorithms, XGBoost has the best performance with 81.78% OA, Kernel SVM, RF and SVM 81.43%, 81.42% and 78.56%, respectively. According to the statistical data in <xref ref-type="table" rid="T3">
<bold>Tables&#xa0;3</bold>
</xref>, <xref ref-type="table" rid="T4">
<bold>4</bold>
</xref>, the classification accuracy of the deep learning algorithm is better than that of the traditional machine learning algorithm on the whole. Even the classification accuracy of the rough FCN-32S model is slightly higher than that of XGBoost, while the highest classification accuracy of MSSNet is 90.68%, which is obviously higher than that of the traditional machine learning classifier. <xref ref-type="fig" rid="f15">
<bold>Figure&#xa0;15</bold>
</xref> and <xref ref-type="fig" rid="f16">
<bold>Figure&#xa0;16</bold>
</xref> show the crop classification in the study area of XGBoost and MSSNet, respectively.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Comparison of traditional machine learning classifiers.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Classifier</th>
<th valign="middle" align="left">Hyperparameters</th>
<th valign="middle" align="center">OA</th>
<th valign="middle" align="right">Kappa</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="4" align="left">Random<break/>Forest</td>
<td valign="middle" align="left">n_estimators: 30, 50, 100, <bold>200</bold>, 300</td>
<td valign="middle" rowspan="4" align="center">81.42%</td>
<td valign="middle" rowspan="4" align="right">73.58%</td>
</tr>
<tr>
<td valign="middle" align="left">max_depth: 5, 10, <bold>20</bold>, 30, None</td>
</tr>
<tr>
<td valign="middle" align="left">min_samples_split: <bold>3</bold>, 5, 10, 30, 100</td>
</tr>
<tr>
<td valign="middle" align="left">min_samples_leaf: <bold>1</bold>, 3, 5, 7, 10</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="left">SVM</td>
<td valign="middle" align="left">C: 0.01, 0.05, 0.1, 0.5, <bold>1</bold>, 5, 10, 100</td>
<td valign="middle" rowspan="2" align="center">78.56%</td>
<td valign="middle" rowspan="2" align="right">69.55%</td>
</tr>
<tr>
<td valign="middle" align="left">Kernel: &#x201c;linear&#x201d;</td>
</tr>
<tr>
<td valign="middle" rowspan="3" align="left">Kernel SVM</td>
<td valign="middle" align="left">C: 0.01, 0.05, 0.1, 0.5, <bold>1</bold>, 5, 10, 100</td>
<td valign="middle" rowspan="3" align="center">81.43%</td>
<td valign="middle" rowspan="3" align="right">73.40%</td>
</tr>
<tr>
<td valign="middle" align="left">Kernel: &#x201c;rbf&#x201d;</td>
</tr>
<tr>
<td valign="middle" align="left">Gamma:0.1, <bold>0.2</bold>, 0.3, 0.4, 0.5, 0.6, &#x201c;auto&#x201d;</td>
</tr>
<tr>
<td valign="middle" rowspan="8" align="left">XGBoost</td>
<td valign="middle" align="left">learning _rate: 0.01, <bold>0.02</bold>, 0.05, 0.1, 0.2</td>
<td valign="middle" rowspan="8" align="center">81.78%</td>
<td valign="middle" rowspan="8" align="right">74.07%</td>
</tr>
<tr>
<td valign="middle" align="left">gamma:0.05, <bold>0.1</bold>, 0.2, 0.5, 0.7, 1</td>
</tr>
<tr>
<td valign="middle" align="left">max_depth:5, 7, 9, <bold>15</bold>, 17, 21, 25</td>
</tr>
<tr>
<td valign="middle" align="left">min_child_weight:1, <bold>5</bold>, 7, 9, 11</td>
</tr>
<tr>
<td valign="middle" align="left">subsamples:0.5, <bold>0.6</bold>, 0.8, 1</td>
</tr>
<tr>
<td valign="middle" align="left">colsample_bytree:0.5, <bold>0.6</bold>, 0.8, 1</td>
</tr>
<tr>
<td valign="middle" align="left">reg_labda:0.01, <bold>0.1</bold>, 1</td>
</tr>
<tr>
<td valign="middle" align="left">reg_alpha:0, <bold>0.</bold>1, 0.3, 0.5, 1</td>
</tr>
<tr>
<td valign="middle" align="left">MSSNet</td>
<td valign="middle" align="left">
<bold>-</bold>
</td>
<td valign="top" align="center">
<bold>90.68%</bold>
</td>
<td valign="top" align="right">
<bold>86.75%</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values mean the highest classification accuracy.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="f15" position="float">
<label>Figure&#xa0;15</label>
<caption>
<p>Crop classification in the study area using XGBoost algorithm.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g015.tif"/>
</fig>
<fig id="f16" position="float">
<label>Figure&#xa0;16</label>
<caption>
<p>Crop classification in the study area using MSSNet model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1196634-g016.tif"/>
</fig>
</sec>
</sec>
<sec id="s5" sec-type="discussion">
<label>5</label>
<title>Discussion and conclusion</title>
<p>Crop classification is the basis of large-scale crop acreage estimation. Currently, advanced Earth observation technology can identify the spatial distribution of crops on the plot scale. In this study, Sentinel-2 multi-spectral image with a single time phase and deep learning algorithm were used to make an attempt on the task of crop fine classification. In this paper, a multi-scale feature potential representation network MSSNet is proposed. Using ResNet50 as the backbone network, the network uses convolution kernel and void convolution of different sizes in the multi-scale feature module, which can frequently merge the features of different scale branches, and then learn more accurate feature maps to assist classification decision. In the experiment, we compared 4 traditional machine learning classification algorithms with 4 deep learning algorithms including MSSNet, and the classification results show that the deep learning algorithm has obvious advantages, especially the algorithm we proposed has obvious improvement compared with other commonly used classification algorithms.</p>
<p>Existing studies have shown that the best time for crop identification is between week 11 and 20 during the growing period (<xref ref-type="bibr" rid="B39">Xu et&#xa0;al., 2021</xref>), In this paper, Sentinel-2 images from the 14th week of crop growth were used to explore the application potential of single phase remote sensing in crop classification. Temporal, spectral and spatial characteristics are the basis of crop classification based on remote sensing technology. The method of crop extraction by using time series image has become an important method to extract crop planting structure by making full use of the characteristics of crop seasonal rhythm. However, it is often difficult to obtain image data of large range and long time series. In this study, a high classification accuracy is achieved by using single-phase optical images with only spectrum-space features. The selection of input features has an important impact on the performance of the model. Since NDVI and EVI can distinguish the phenological differences of different crops, these two artificial features have been introduced into crop classification experiments in large numbers. This practice was followed in feature design in this paper. 12 features such as blue, green, red, near-infrared band, normalized vegetation index and enhanced vegetation index were selected as key features for crop identification. The distribution of the importance of input features is closely related to the model structure, and the distinction of subtle differences between categories is the key to fine-grained image classification. Specifically, the feature extraction ability of the model for local spatial features determines the degree of refinement of classification results. In this paper, a multi-scale feature fusion module is designed in semantic segmentation model based on void convolution technology. By combining feature maps of different scales, the expression of ground object details is enhanced on the premise of ensuring classification accuracy. At the same time, in object detection and semantic segmentation tasks, model performance is highly dependent on features extracted by backbone. Therefore, we believe that it is very necessary to analyze input features when designing deep learning models.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>TL wrote the draft of the manuscript. TL, LW contributed to data curation, analysis. MG contributed to manuscript revision. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Belgiu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Csillik</surname> <given-names>O.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Sentinel-2 cropland mapping using pixel-based and object-based timeweighted dynamic time warping analysis</article-title>. <source>Remote Sens. Environ.</source> <volume>204</volume>, <fpage>509</fpage>&#x2013;<lpage>523</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rse.2017.10.005</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Learning deep architectures for AI</article-title>. <source>Foundations Trends Macine Learn.</source> <volume>2</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>127</lpage>. doi: <pub-id pub-id-type="doi">10.1561/9781601982957</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Delalleau</surname> <given-names>O.</given-names>
</name>
</person-group> (<year>2011</year>). <source>On the expressive power of deep architectures</source> (<publisher-loc>Verlag</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>18</fpage>&#x2013;<lpage>36</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bi</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Qin</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multi-scale stacking attention pooling for remote sensing scene classification</article-title>. <source>Neurocomputing</source> <volume>436</volume>, <fpage>147</fpage>&#x2013;<lpage>161</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neucom.2021.01.038</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Kokkinos</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Murphy</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yuille</surname> <given-names>A. L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>DeepLab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected CRFs</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>40</volume> (<issue>4</issue>), <fpage>834</fpage>&#x2013;<lpage>848</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2699184</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Schroff</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Adam</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Rethinking atrous convolution for semantic image segmentation</article-title>. <source>arXiv:1706.05587</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1706.05587</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Y. K.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Schroff</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Adam</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Encoder-decoder with atrous separable convolution for semanticimage segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>. <fpage>801</fpage>&#x2013;<lpage>818</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Cutler</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Cutler</surname> <given-names>D. R.</given-names>
</name>
<name>
<surname>Stevens</surname> <given-names>J. R.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Random Fforests</article-title>. <source>Mach. Learning</source>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ferrant</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Selles</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Le Page</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Herrault</surname> <given-names>P. A.</given-names>
</name>
<name>
<surname>Pelletier</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Al-Bitar</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Detection of irrigated crops from sentinel-1 and sentinel-2 data to estimate seasonal groundwater use in South India</article-title>. <source>Remote Sens.</source> <volume>9</volume> (<issue>11</issue>), <fpage>1119</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs9111119</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Classification of very-high-spatial-resolution aerial images based on multiscale features with limited semantic information</article-title>. <source>Remote Sens.</source> <volume>13</volume> (<issue>3</issue>), <fpage>364</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs13030364</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gong</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Dai</surname> <given-names>H.</given-names>
</name>
<name>
<surname>He</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Multiscale information fusion for hyperspectral image classification based on hybrid 2D-3D CNN</article-title>. <source>Remote Sens.</source> <volume>13</volume> (<issue>12</issue>), <fpage>2268</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs13122268</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hopfield</surname> <given-names>J. J.</given-names>
</name>
</person-group> (<year>1982</year>). <article-title>Neural networks and physical systems with emergent collective computational abilities</article-title>. <source>Proc. Natl. Acad. Sci. United States America</source> <volume>79</volume>, <fpage>2254</fpage>&#x2013;<lpage>2558</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1073/pnas.79.8.2554</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Atzberger</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>The optimal threshold and vegetation index time series for retrieving crop phenology based on a modified dynamic threshold method</article-title>. <source>Remote Sens.</source> <volume>11</volume> (<issue>23</issue>), <fpage>2725</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs11232725</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Immitzer</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Vuolo</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Atzberger</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>First experience with Sentinel-2 data for crop and tree species classifications in central Europe</article-title>. <source>Remote Sens.</source> <volume>8</volume> (<issue>3</issue>), <fpage>166</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs8030166</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kawaguchi</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Kaelbling</surname> <given-names>L. P.</given-names>
</name>
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Generalization in deep learning</article-title>. <source>arXiv:1710.05468</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1710.05468</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kussul</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Lavreniuk</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Skakun</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Shelestov</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deep learning classification of land cover and crop types using remote sensing data</article-title>. <source>Remote Sens. Lett.</source> <volume>14</volume> (<issue>5</issue>), <fpage>778</fpage>&#x2013;<lpage>782</lpage>. doi: <pub-id pub-id-type="doi">10.1109/LGRS.2017.2681128</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>LeCun</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bengio</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Hinton</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep learning</article-title>. <source>Nature</source> <volume>521</volume> (<issue>7553</issue>), <fpage>436</fpage>&#x2013;<lpage>444</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/NATURE14539</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>X. F.</given-names>
</name>
<name>
<surname>Pu</surname> <given-names>H. B.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>H. X.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Introduce GIoU into RFB net to optimize object detection bounding box</article-title>,&#x201d; in <conf-name>ICCIP 2019: 2019 the 5th International Conference on Communication and Information Processing</conf-name>. <fpage>108</fpage>&#x2013;<lpage>113</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Land cover classification from very high spatial resolution images <italic>via</italic> multiscale object-driven CNNs and automatic annotation</article-title>. <source>J. Appl. Remote Sens.</source> <volume>16</volume> (<issue>1</issue>), <fpage>014513</fpage>. doi: <pub-id pub-id-type="doi">10.1117/1.JRS.16.014513</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Milan</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Reid</surname> <given-names>I.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>RefineNet: Multi-path refinement networks for high-resolution semantic segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <fpage>1925</fpage>&#x2013;<lpage>1934</lpage>.</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>S.</given-names>
</name>
<name>
<surname>He</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Lan</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Comparison of multi-source satellite images for classifying marsh vegetation using DeepLabV3 Plus deep learning algorithm</article-title>. <source>Ecol. Indic.</source> <volume>125</volume> (<issue>11</issue>), <fpage>107562</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ecolind.2021.107562</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Jin</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Object detection in remote sensing image with improved RFB net</article-title>. <source>J. Geomatics Sci. Technol.</source> <volume>2)</volume>, <fpage>179</fpage>&#x2013;<lpage>184</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Receptive field block net for accurate and fast object detection</article-title>,&#x201d; in <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>, Vol. <volume>2018</volume>. <fpage>385</fpage>&#x2013;<lpage>400</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Dual attention guided multi-scale CNN for fine-grained image classification</article-title>. <source>Inf. Sci.</source> <volume>573</volume>, <fpage>37</fpage>&#x2013;<lpage>45</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.ins.2021.05.040</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Long</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Shelhamer</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Fully convolutional networks for semantic segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <fpage>3431</fpage>&#x2013;<lpage>3440</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname> <given-names>X. B.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Intelligent classification of remote sensing images and its application</article-title>. <source>Publishing House Electron. Industry</source>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Milletari</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Navab</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Ahmadi</surname> <given-names>S. A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>V-Net: Fully convolutional neural networks for volumetric medical image segmentation</article-title>. <source>arXiv:1606.04797</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/3DV.2016.79</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pedrayes</surname> <given-names>O. D.</given-names>
</name>
<name>
<surname>Lema</surname> <given-names>D. G.</given-names>
</name>
<name>
<surname>Garc&#xed;a</surname> <given-names>D. F.</given-names>
</name>
<name>
<surname>Usamentiaga</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Alonso</surname> <given-names>&#xc1;.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Evaluation of semantic segmentation methods for land use with spectral imaging using sentinel-2 and PNOA imagery</article-title>. <source>Remote Sens.</source> <volume>13</volume> (<issue>12</issue>), <fpage>2292</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs13122292</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Rustowicz</surname> <given-names>R. M.</given-names>
</name>
<name>
<surname>Cheong</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ermon</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Burke</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Lobell</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Semantic segmentation of crop type in Africa: A novel dataset and analysis of deep learning methods</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <fpage>75</fpage>&#x2013;<lpage>82</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Multi-scale global retrieval and temporal-spatial consistency matching based long-term tracking network</article-title>. <source>Chin. J. Electron.</source> <volume>32</volume>, <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1049/cje.2021.00.195</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Silveira</surname> <given-names>E. M. D.</given-names>
</name>
<name>
<surname>de Menezes</surname> <given-names>M. D.</given-names>
</name>
<name>
<surname>Acerbi</surname> <given-names>F. W.</given-names>
</name>
<name>
<surname>Terra</surname> <given-names>M.</given-names>
</name>
<name>
<surname>de Mello</surname> <given-names>J. M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Assessment of geostatistical features for object-based image classification of contrasted landscape vegetation cover</article-title>. <source>J. Appl. Remote Sensing.</source> <volume>11</volume> (<issue>3</issue>), <fpage>036004</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/1.JRS.11.036004</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>W. W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Development status and literature analysis of earth observation remote sensing satellites in China</article-title>. <source>Natl. Remote Sens. Bull.</source> <volume>5)</volume>, <fpage>479</fpage>&#x2013;<lpage>510</lpage>. doi: <pub-id pub-id-type="doi">10.11834/jrs.20209464</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vapnik</surname> <given-names>V. N.</given-names>
</name>
</person-group> (<year>1995</year>). <source>The Nature of Statistical Learning Theory</source> (<publisher-name>Springer</publisher-name>).</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Varadarajan</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Garg</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Kotecha</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>An efficient deep convolutional neural network approach for object detection and recognition using a multi-scale anchor box in real-time</article-title>. <source>Future Internet.</source> <volume>13</volume> (<issue>12</issue>), <fpage>307</fpage>. doi: <pub-id pub-id-type="doi">10.3390/fi13120307</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Verrelst</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Mu&#xf1;oz</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Alonso</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Delegido</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Rivera</surname> <given-names>J. P.</given-names>
</name>
<name>
<surname>Camps-Valls</surname> <given-names>G.</given-names>
</name>
<etal/>
</person-group>. (<year>2012</year>). <article-title>Machine learning regression algorithms for biophysical parameter retrieval: Opportunities for Sentinel-2 and -3</article-title>. <source>Remote Sens. Environ.</source> <volume>118</volume>, <fpage>127</fpage>&#x2013;<lpage>139</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rse.2011.11.002</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A filtering method for LiDAR point cloud based on multi-scale CNN with attention mechanism</article-title>. <source>Remote Sensing.</source> <volume>14</volume> (<issue>23</issue>), <fpage>6170</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs14236170</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Real-time pixel-wise grasp affordance prediction based on multi-scale context information fusion</article-title>. <source>Ind. Robot</source> <volume>49</volume> (<issue>2</issue>), <fpage>368</fpage>&#x2013;<lpage>381</lpage>. doi: <pub-id pub-id-type="doi">10.1108/IR-06-2021-0118</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname> <given-names>J.</given-names>
</name>
<name>
<surname>He</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ghamisi</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multiscale densely-connected fusion networks for hyperspectral images classification</article-title>. <source>IEEE Trans. Circuits Syst. Video Technol.</source> <volume>31</volume> (<issue>1</issue>), <fpage>246</fpage>&#x2013;<lpage>259</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TCSVT.2020.2975566</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xiong</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Towards interpreting multi-temporal deep learning models in crop mapping</article-title>. <source>Remote Sens. Environ.</source> <volume>264</volume>, <fpage>112599</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rse.2021.112599</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Koltun</surname> <given-names>V.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Multi-scale context aggregation by dilated convolutions</article-title>. <source>arXiv:1511.07122</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1511.07122</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Qi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Object detection in remote sensing images <italic>via</italic> multi-feature pyramid network with receptive field block</article-title>. <source>Remote Sens.</source> <volume>13</volume> (<issue>5</issue>), <fpage>862</fpage>. doi: <pub-id pub-id-type="doi">10.3390/rs13050862</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Remote sensing image classification based on DeepLab-v3+</article-title>. <source>Laser Optoelectronics Prog.</source> <volume>56</volume> (<issue>15</issue>), <fpage>152801</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3788/LOP56.152801</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Qi</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Pyramid scene parsing network</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>. <fpage>2881</fpage>&#x2013;<lpage>2890</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep learning based multi-temporal crop classification</article-title>. <source>Remote Sens. Environ.</source> <volume>221</volume>, <fpage>430</fpage>&#x2013;<lpage>443</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.rse.2018.11.032</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>