<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="review-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Ecol. Evol.</journal-id>
<journal-title>Frontiers in Ecology and Evolution</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Ecol. Evol.</abbrev-journal-title>
<issn pub-type="epub">2296-701X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fevo.2023.1201125</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Ecology and Evolution</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep learning-based semantic segmentation of remote sensing images: a review</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Lv</surname>
<given-names>Jinna</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1978392"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Shen</surname>
<given-names>Qi</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2000880"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Lv</surname>
<given-names>Mingzheng</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Yiran</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Shi</surname>
<given-names>Lei</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2092111"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Peiying</given-names>
</name>
<xref ref-type="aff" rid="aff5">
<sup>5</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1761709"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Information Management, Beijing Information Science and Technology University</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Teacher&#x2019;s College, Beijing Union University</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>School of International Education, Shangqiu Normal University</institution>, <addr-line>Shangqiu</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>State Key Laboratory of Media Convergence and Communication, Communication University of China</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<aff id="aff5">
<sup>5</sup>
<institution>College of Computer Science and Technology, China University of Petroleum (East China)</institution>, <addr-line>Qingdao</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Qianqian Wang, Beijing Institute of Technology, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Shunchun Yao, South China University of Technology, China; Jinping Liu, Hunan Normal University, China; Yujin Zhang, Shanghai University of Engineering Sciences, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Qi Shen, <email xlink:href="mailto:sftshenqi@buu.edu.cn">sftshenqi@buu.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1201125</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Lv, Shen, Lv, Li, Shi and Zhang</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Lv, Shen, Lv, Li, Shi and Zhang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Semantic segmentation is a fundamental but challenging problem of pixel-level remote sensing (RS) data analysis. Semantic segmentation tasks based on aerial and satellite images play an important role in a wide range of applications. Recently, with the successful applications of deep learning (DL) in the computer vision (CV) field, more and more researchers have introduced and improved DL methods to the task of RS data semantic segmentation and achieved excellent results. Although there are a large number of DL methods, there remains a deficiency in the evaluation and advancement of semantic segmentation techniques for RS data. To solve the problem, this paper surveys more than 100 papers in this field in the past 5 years and elaborates in detail on the aspects of technical framework classification discussion, datasets, experimental evaluation, research challenges, and future research directions. Different from several previously published surveys, this paper first focuses on comprehensively summarizing the advantages and disadvantages of techniques and models based on the important and difficult points. This research will help beginners quickly establish research ideas and processes in this field, allowing them to focus on algorithm innovation without paying too much attention to datasets, evaluation indicators, and research frameworks.</p>
</abstract>
<kwd-group>
<kwd>remote sensing</kwd>
<kwd>deep learning</kwd>
<kwd>convolutional neural network</kwd>
<kwd>semantic segmentation</kwd>
<kwd>satellite image</kwd>
</kwd-group>
<counts>
<fig-count count="10"/>
<table-count count="3"/>
<equation-count count="4"/>
<ref-count count="163"/>
<page-count count="22"/>
<word-count count="11179"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Environmental Informatics and Remote Sensing</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Semantic segmentation is one of the most important problems in computer vision (CV) (<xref ref-type="bibr" rid="B91">Mo et&#xa0;al., 2022</xref>). The goal is to determine the class of each pixel in an image, which has great significance to the analysis and understanding of scene images. For RS data, semantic segmentation also plays a key role in a variety of geographic information applications, including urban planning (<xref ref-type="bibr" rid="B159">Zheng et&#xa0;al., 2020b</xref>; <xref ref-type="bibr" rid="B1">Abdollahi et&#xa0;al., 2021</xref>), economic assessment (<xref ref-type="bibr" rid="B111">Song et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B130">Wang et&#xa0;al., 2022a</xref>), land resource management (<xref ref-type="bibr" rid="B120">Tong et&#xa0;al., 2020</xref>), precision agriculture (<xref ref-type="bibr" rid="B135">Weiss et&#xa0;al., 2020</xref>), and environmental protection (<xref ref-type="bibr" rid="B113">Subudhi et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B57">Li and Liu, 2023</xref>).</p>
<p>With the rise and evolution of deep neural networks, deep learning (DL) has made tremendous breakthroughs in artificial intelligence fields such as CV and natural language processing (NLP) (<xref ref-type="bibr" rid="B50">Kitaev et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B105">Ru et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B151">Zhang et&#xa0;al., 2022b</xref>; <xref ref-type="bibr" rid="B93">Noble et&#xa0;al., 2023</xref>). However, there are huge challenges for RS data, such as top-down data acquisition perspective, large image scale, different image resolution, and variable lighting conditions. DL methods on natural images can be directly employed in RS data. How to efficiently and accurately use the original RS data to obtain the required information and obtain more accurate segmentation results is difficult.</p>
<p>DL is gradually being applied to RS semantic segmentation (<xref ref-type="bibr" rid="B101">Piramanayagam et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B116">Sun et&#xa0;al., 2019</xref>). Early methods based on sliding windows and candidate regions are time-consuming and have a lot of redundant calculations (<xref ref-type="bibr" rid="B25">Davis et&#xa0;al., 1975</xref>; <xref ref-type="bibr" rid="B95">&#xd6;zden and Polat, 2005</xref>; <xref ref-type="bibr" rid="B107">Senthilkumaran and Rajesh, 2009</xref>; <xref ref-type="bibr" rid="B94">Nowozin and Lampert, 2011</xref>; <xref ref-type="bibr" rid="B21">Ciresan et&#xa0;al., 2012</xref>). In recent years, there are more and more DL-based methods in RS, including U-Net methods and its variants (<xref ref-type="bibr" rid="B144">Yue et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B34">Foivos et&#xa0;al., 2020</xref>), multi-scale context aggregation networks (<xref ref-type="bibr" rid="B68">Liu et&#xa0;al., 2018a</xref>; <xref ref-type="bibr" rid="B20">Chen et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B56">Li C. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B139">Xu H. et&#xa0;al., 2022</xref>), and multi-level feature fusion networks (<xref ref-type="bibr" rid="B29">Dong and Chen, 2021</xref>; <xref ref-type="bibr" rid="B66">Li et&#xa0;al., 2021b</xref>). The attention mechanism that pays attention to relevant information and ignores irrelevant information is frequently adopted with its advantages (<xref ref-type="bibr" rid="B65">Li et&#xa0;al., 2021a</xref>; <xref ref-type="bibr" rid="B66">Li et&#xa0;al., 2021b</xref>; <xref ref-type="bibr" rid="B55">Li YC. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B108">Seong and Choi, 2021</xref>). Later, Transformer-based and generative adversarial network (GAN) methods no doubt are getting more and more attention (<xref ref-type="bibr" rid="B109">Shamsolmoali et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B119">Tian et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B23">Cui L. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B40">He et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B128">Wang L. et&#xa0;al., 2022</xref>).</p>
<p>Some works of literature have reviewed the research methods in the field of semantic segmentation of RS data, classified and explored from different perspectives, including RS image analysis on general DL algorithms (<xref ref-type="bibr" rid="B163">Zhu et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B121">Tsagkatakis et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B47">Jiang et&#xa0;al., 2022</xref>), detection field methods (<xref ref-type="bibr" rid="B5">Asokan and Anitha, 2019</xref>; <xref ref-type="bibr" rid="B60">Li et&#xa0;al., 2022a</xref>), Transformer models (<xref ref-type="bibr" rid="B3">Aleissaee et&#xa0;al., 2022</xref>), and image registration methods (<xref ref-type="bibr" rid="B147">Zhang X. et&#xa0;al., 2021</xref>). <xref ref-type="bibr" rid="B106">Sebastian et al. (2022)</xref>. explored and evaluated the strengths and weaknesses of various qualitative and quantitative image segmentation evaluation metrics used in RS applications. <xref ref-type="bibr" rid="B78">Lu et&#xa0;al. (2021)</xref> introduced and analyzed the studies and applications of satellite data from the perspective of semantics, and carried out analysis and discussions from the four research areas of semantic understanding, semantic segmentation, semantic classification, and semantic search. <xref ref-type="bibr" rid="B64">Li et&#xa0;al. (2018)</xref> discussed and performed a comparative analysis of DL models for semantic classification. <xref ref-type="bibr" rid="B121">Tsagkatakis et&#xa0;al. (2019)</xref> comprehensively reviewed DL methods for enhancing RS observations, focusing on key tasks including single- and multi-band super-resolution, denoising, restoration, pan-sharpening, and fusion. Most of the discussed methods are outdated and lack the interpretation of the latest research results and algorithms for semantic segmentation.</p>
<p>To promote the semantic segmentation methods of RS data, we focus on the latest research methods, recent open datasets, and evaluation methods, and look forward to the future direction. First, we describe the definition and research overview of semantic segmentation tasks. Second, we discuss different technical frameworks of DL and summarize their advantages and disadvantages. Then, we illustrate the public RS datasets from data collection and the comparison of experimental results with different methods. Finally, we summarize future challenges and research directions. The main contributions of this paper are as follows:</p>
<list list-type="bullet">
<list-item>
<p>Highlight approaches of DL in RS data for semantic segmentation tasks in the recent 5 years from different technical frameworks, including new architectures, DL components, and advantages and disadvantages.</p>
</list-item>
<list-item>
<p>Detailed summaries of RS datasets. It contains the dataset name, description, classes, channel number, and URL. Statistical results of different methods on the same dataset are also summarized.</p>
</list-item>
<list-item>
<p>This review not only summarizes the existing achievements but also points out some promising research directions for semantic segmentation. From this perspective, it helps potential readers find research points and motivates engineers to develop advanced application patterns.</p>
</list-item>
</list>
</sec>
<sec id="s2">
<label>2</label>
<title>Overview</title>
<p>This section provides an overview of methods for semantic segmentation in the RS field. First, the basic definition of semantic segmentation is described. Second, the traditional and mainstream methods are explained. Third, statistics are made on the semantic segmentation methods of RS images in the past 5 years, and the main published journals, quantity, and keyword visualization of papers are analyzed.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Definition and concept</title>
<p>Semantic segmentation is a very important direction in the CV. Unlike target detection and recognition, semantic segmentation achieves image pixel-level classification. It can divide a picture into multiple blocks according to the similarities and differences of categories. Semantically related pixels are annotated with the same label. The semantic segmentation algorithm can comprehensively complete the recognition, detection, and segmentation of visual elements in the scene, and improve the efficiency and accuracy of image understanding. Compared with image classification and target detection, the semantic segmentation results can provide richer information about image parts and details. Semantic segmentation algorithms have extensive applications and long-term development prospects. For example, in automatic driving technology, semantic segmentation algorithms can assist the automatic driving system to judge road conditions by segmenting roads, vehicles, and pedestrians. For RS images, semantic segmentation plays an irreplaceable role in disaster assessment, crop yield estimation, and land change monitoring.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Research overview</title>
<p>Through the Google Scholar search keys &#x201c;semantic segmentation&#x201d; and &#x201c;remote sensing&#x201d;, the statistics of the number of papers published since 2015 are shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. We can see that semantic segmentation, as an important task in the field of remote sensing, has attracted the attention of researchers. Moreover, more and more new technologies and methods are emerging.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>The number of papers based on remote sensing semantic segmentation since 2015.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g001.tif"/>
</fig>
<p>We collect the works from more than 100 articles based on RS semantic segmentation. This article counts them. The statistics of the research published are shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2A</bold>
</xref> (journals with less than three articles not shown). According to the number of journals published, the top four are, in order, <italic>Remote Sensing, IEEE Transactions on Geoscience and Remote Sensing, Journal of Photogrammetry and Remote Sensing</italic>, and <italic>IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing</italic>. The year&#x2019;s distribution of the research papers is shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2B</bold>
</xref>. It can be found that we pay more attention to the latest studies of the past 2 years, which represent the current advanced technologies for semantic segmentation.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>The distribution map of publications and years of analyzed articles. <bold>(A)</bold> The percentage distribution of publications. <bold>(B)</bold> The percentage distribution of years.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g002.tif"/>
</fig>
<p>We analyze the keywords of the papers in the past 5 years. The keywords have substantial meaning for expressing the central content of the paper and can map the research content and direction in recent years. The statistical results are displayed in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>, through the word cloud diagram. We can see that, in addition to the task keyword &#x201c;semantic segmentation&#x201d; and the data keyword &#x201c;remote sensing&#x201d;, &#x201c;attention&#x201d;, &#x201c;convolutional neural&#x201d;, &#x201c;Transformer&#x201d;, &#x201c;GAN&#x201d;, and &#x201c;unsupervised&#x201d; are used more from the perspective of methods technology.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>The cloud map of keywords statistics of analyzed papers.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g003.tif"/>
</fig>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Method overview</title>
<p>All developments are a long technical accumulation. Early methods of semantic segmentation used traditional methods. With the emergence of DL, more and more new methods are emerging. There are also many excellent DL methods in the field of RS data.</p>
<p>Early semantic segmentation research focused on non-DL models, such as the threshold method (<xref ref-type="bibr" rid="B25">Davis et&#xa0;al., 1975</xref>), the clustering-based method (<xref ref-type="bibr" rid="B95">&#xd6;zden and Polat, 2005</xref>), the edge detection method (<xref ref-type="bibr" rid="B107">Senthilkumaran and Rajesh, 2009</xref>), and conditional random fields (CRFs) (<xref ref-type="bibr" rid="B94">Nowozin and Lampert, 2011</xref>). These traditional methods have low efficiency and accuracy.</p>
<p>With the popularity of DL, many classic semantic segmentation models have been designed. The fully convolutional network (FCN) (<xref ref-type="bibr" rid="B76">Long et&#xa0;al., 2019</xref>) is the first applied DL to a semantic segmentation task. It changed the fully connected layer of CNN to a convolutional layer. However, the receptive field of FCN is fixed, and it is easy to lose detailed information. To solve this problem, the SegNet model (<xref ref-type="bibr" rid="B7">Badrinarayanan et&#xa0;al., 2017</xref>) was proposed. It can reduce the number of parameters by using pooling indices to save the contour information of the image. The U-Net network (<xref ref-type="bibr" rid="B104">Ronneberger et&#xa0;al., 2015</xref>) is an extension of FCN. Its main innovation is utilizing four layers of skip connections in the middle. DeepLab V1 (<xref ref-type="bibr" rid="B15">Chen et&#xa0;al., 2015</xref>) alleviated the down-sampling problem and makes the segmentation boundary clearer by replacing the traditional convolutional layer with porous convolution. It discarded the fully connected layer of the VGG16 and changed the last two pooling steps to one. Then, DeepLab v2 (<xref ref-type="bibr" rid="B16">Chen et&#xa0;al., 2017a</xref>), DeepLab v3 (<xref ref-type="bibr" rid="B17">Chen et&#xa0;al., 2017b</xref>), and DeepLab v3+ (<xref ref-type="bibr" rid="B19">Chen LC. et&#xa0;al., 2018</xref>) were proposed. The contribution of DeepLab v2 lies in the more flexible use of atrous convolution and Atrous Spatial Pyramid Pooling (ASPP). DeepLab v2 abandoned CRF, and improved ASPP, using dilated convolution to deepen the network. DeepLab v3+ modified the main network again and upgraded ResNet-101 to Xception.</p>
<p>In recent years, many improved methods based on classical DL semantic segmentation have been applied to remote sensing images (<xref ref-type="bibr" rid="B162">Zhou et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B28">Ding et&#xa0;al., 2020b</xref>; <xref ref-type="bibr" rid="B96">Pan et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B8">Bai et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B43">Huang et&#xa0;al., 2022</xref>). <xref ref-type="bibr" rid="B8">Bai et&#xa0;al. (2021)</xref> proposed an improved model HCANet and designed two Compact Atrous Spatial Pyramid Pooling (CASPP and CASPP+) modules. <xref ref-type="bibr" rid="B43">Huang et al. (2022)</xref> improved U-Net and U-net++ (<xref ref-type="bibr" rid="B162">Zhou et&#xa0;al., 2018</xref>) connections in one to four layers of U-Net. The advantage of this structure is that the network can learn the significance of characteristics from different depths and fuse them. There are many kinds of research with the latest technologies, for example, attention mechanism (<xref ref-type="bibr" rid="B130">Wang et&#xa0;al., 2022a</xref>), generative confrontation network (<xref ref-type="bibr" rid="B96">Pan et&#xa0;al., 2020</xref>), and Transformer (<xref ref-type="bibr" rid="B28">Ding et&#xa0;al., 2020b</xref>). These new methods have improved the performance of semantic segmentation tasks in RS data and have had an important impact on the evolution of this task.</p>
<p>From the perspective of the DL technology framework, this paper categorizes and outlines the semantic segmentation methods in RS data in the past 5 years by dividing them into six categories, namely, based on CNN, based on attention mechanism, multi-scale strategy, based on Transformer, based on GAN, and fusion-based methods. We display these network models in recent years in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Semantic segmentation methods of remote sensing based on deep learning.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g004.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Semantic segmentation framework</title>
<sec id="s3_1">
<label>3.1</label>
<title>General CNN-based methods</title>
<p>There are many semantic segmentation studies based on the CNN of RS data. This section discusses these papers from the following categories: FCN-based, U-net-based, SegNet-based, DeepLab-based, and other convolutional network methods.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>FCN architecture</title>
<p>In the FCN, 1&#xd7;1 convolution replaces the full connection in the CNN. Then, the probability category value of each pixel is obtained through the softmax layer. FCN introduces the deconvolution shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>. The true category of each pixel is the category with the largest corresponding probability value. Finally, a segmented image is obtained, whose size is the same resolution as the input image. The deconvolution uses known convolution kernels and convolutional output to restore images, thereby obtaining refined features. The reason that FCN is more efficient than CNNs is that computing convolutions are avoided one by one for each pixel block, in which adjacent pixel blocks are repeated.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>The FCN architecture and its deconvolution operator (<xref ref-type="bibr" rid="B76">Long et&#xa0;al., 2019</xref>).  <bold>(A)</bold> FCN architecture. <bold>(B)</bold> Deconvolution.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g005.tif"/>
</fig>
<p>Although the FCN has pushed great breakthroughs in the scene segmentation problem, it relies on a large-scale image recognition network, which is usually trained on a large number of images (<xref ref-type="bibr" rid="B86">Marmanis et&#xa0;al., 2016</xref>). However, in RS domains, label scarcity is a difficult problem. <xref ref-type="bibr" rid="B48">Kemker et&#xa0;al. (2018)</xref> were the first ones to apply the FCN to segment multispectral RS images. It used massive automatically labeled synthetic multispectral images and gained good results. Then, many FCN-based methods (<xref ref-type="bibr" rid="B46">Iglovikov et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B110">Shao et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B134">Wei et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B12">Chen et&#xa0;al., 2022</xref>) are introduced to improve the performance of the segmentation.</p>
<p>Owing to the characteristics of RS data that come from additional spectral bands set by multiple sensors, the commonly used RGB-based pre-training model cannot meet the requirements. According to the characteristics of RS data, some studies have improved the FCN-based method to achieve good semantic segmentation results (<xref ref-type="bibr" rid="B73">Liu et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B18">Chen G. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B10">Chen L. et&#xa0;al., 2021</xref>). <xref ref-type="bibr" rid="B73">Liu et&#xa0;al. (2019)</xref> fused the RGB feature to obtain semantic labels from a DL framework with light detection and ranging (Li-DAR) features. <xref ref-type="bibr" rid="B18">Chen G. et&#xa0;al. (2021)</xref> proposed an improved structure, named SDFCNv2, to optimize the segmentation results of RS data. First, they designed a hybrid model basic convolutional block to obtain a larger receptive field. Second, they develop the spatial channel fusion model to reduce training pressure and improve experimental results. The EFCNet (<xref ref-type="bibr" rid="B10">Chen L. et&#xa0;al., 2021</xref>) was an end-to-end network, which used a depth-variant block to learn the weights of different scale features. Transfer learning with FCN can improve segmentation accuracy (<xref ref-type="bibr" rid="B138">Wurm et&#xa0;al., 2019</xref>). Different convolutional blocks of the FCN network extract multi-scale information without the need for ensemble learning techniques (<xref ref-type="bibr" rid="B99">Pastorino et&#xa0;al., 2022a</xref>). Incorporating the features extracted from the FCN network and spatial information can obtain more accurate results (<xref ref-type="bibr" rid="B100">Pastorino et&#xa0;al., 2022b</xref>).</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>U-Net architecture</title>
<p>The U-Net architecture (<xref ref-type="bibr" rid="B76">Long et&#xa0;al., 2019</xref>) is the most widely utilized model in current semantic segmentation studies. It uses skip connections in addition to the traditional encoder&#x2013;decoder layers to fuse low-level and high-level features in the expansion path to improve localization accuracy. Many variant methods have emerged later, such as Unet++ (<xref ref-type="bibr" rid="B162">Zhou et&#xa0;al., 2018</xref>), DC-Unet (<xref ref-type="bibr" rid="B77">Lou et&#xa0;al., 2021</xref>), and TransUNet (<xref ref-type="bibr" rid="B14">Chen J. et&#xa0;al., 2021</xref>), and methods based on the U-Net structure have been used for RS images (<xref ref-type="bibr" rid="B44">Huang et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B45">Huang et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B118">Tasar et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B88">Maxwell et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B70">Liu Z. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B102">Priyanka et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B126">Wang K. et&#xa0;al., 2022</xref>) and show better performance.</p>
<p>
<xref ref-type="bibr" rid="B88">Maxwell et&#xa0;al. (2020)</xref> experimented with a UNet-based approach on a large dataset of historical land surfaces from the US Geological Survey and reduced manual digitization operations. An incremental learning method (<xref ref-type="bibr" rid="B118">Tasar et&#xa0;al., 2019</xref>), which was a variant of U-Net, included an encoder as the first 13 convolutional layers of VGG16 and a decoder, and two central convolutional layers. ResUNet-a (<xref ref-type="bibr" rid="B34">Foivos et&#xa0;al., 2020</xref>) added residual connections in a U-Net backbone, which solved the problem of gradient disappearance and explosion. It employed multiple parallel atrous convolutions to extract object features at multiple scales. <xref ref-type="bibr" rid="B102">Priyanka et al. (2022)</xref> designed a DIResUNet model to combine the building blocks with the U-Net scheme by integrating initial modules, modified residual blocks, and Dense Global Space Pyramid Pooling (DGSPP). In this way, local and global related scenes are extracted in parallel by a dedicated processing operator, leading to more efficient semantic segmentation. HCANet (<xref ref-type="bibr" rid="B8">Bai et&#xa0;al., 2021</xref>) was similar to the encoder&#x2013;decoder structure of U-Net. In HCANet, there are two modules, CASPP and CASPP+. The CASPP module substitutes the crop operations in U-Net and obtains multi-scale context information from ResNet with multi-scale features. To get aggregated context information, the HCANet method employed the CASPP+ module in the middle layer of the network. <xref ref-type="bibr" rid="B144">Yue et&#xa0;al. (2019)</xref> proposed TreeUNet, which connects a segmentation module and a Tree-CNN block.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>SegNet architecture</title>
<p>The SegNet network includes an encoder network, a symmetric decoder network, and a classification layer pixel-wise. It has 13 convolutional layers that are the same as the VGG16. Up-pooling, which applies the index of Max Pooling, is used in the encoder to the decoder. It improves the recognition effect of the segmentation task on the segmentation boundary. As shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, the positions of the maximum values of the four colors are recorded. In the up-pooling block, these positions are marked, and the other positions are filled with zeros. In this way, the recognition effect of the segmentation task can be improved on the boundary.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>The up-pooling in SegNet architecture (<xref ref-type="bibr" rid="B7">Badrinarayanan et al., 2017</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g006.tif"/>
</fig>
<p>Many studies utilized the method of combining SegNet with other operations to achieve semantic segmentation tasks (<xref ref-type="bibr" rid="B85">Marmanis et&#xa0;al., 2017</xref>), and some researchers used the idea of SegNet to design improved models (<xref ref-type="bibr" rid="B136">Weng et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B159">Zheng et&#xa0;al., 2020b</xref>). <xref ref-type="bibr" rid="B85">Marmanis et&#xa0;al. (2017)</xref> saved memory by adding boundary detection in the SegNet encoder&#x2013;decoder architecture. <xref ref-type="bibr" rid="B136">Weng et&#xa0;al. (2020)</xref> proposed the SR SegNet to accomplish water segmentation. On the one hand, limit the number of parameters by adding an improved residual block and depth separable convolution into the encoder. Meanwhile, the dilated convolution can improve the ability of feature extraction. On the other hand, SR SegNet used more convolution kernels in the encoder network and employed a cascade method to combine different level features of images.</p>
</sec>
<sec id="s3_1_4">
<label>3.1.4</label>
<title>DeepLab architecture</title>
<p>The DeepLab series includes some semantic segmentation algorithms proposed by the Google team. DeepLab v1 was launched in 2014 and achieved second place in the segmentation task on the PASCAL VOC2012 dataset. Then, from 2017 to 2018, DeepLab v2, DeepLab v3, and DeepLab v3+ were successively established. The two innovations of DeepLab v1 are atrous convolution and fully connected CRF. The difference between DeepLab v2 is ASPP. DeepLab v3 further optimizes ASPP, including adding convolution and batch normalization operations. DeepLab v3+ is based on the structure of U-Net and adds an up-sampling decoder module to advance the accuracy of the edge.</p>
<p>In the field of RS, there are also many methods using DeepLab series structural models, such as those based on DeepLab (<xref ref-type="bibr" rid="B11">Chen K. et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B42">Hu et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B123">Venugopal, 2020</xref>; <xref ref-type="bibr" rid="B127">Wang Y. et&#xa0;al., 2021</xref>), and some using DeepLab v3 models (<xref ref-type="bibr" rid="B30">Du et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B51">Kong et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B4">Andrade et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B130">Wang et&#xa0;al., 2022a</xref>; <xref ref-type="bibr" rid="B125">Wang M. et&#xa0;al., 2022</xref>).</p>
<p>A dilated CNN method (<xref ref-type="bibr" rid="B123">Venugopal, 2020</xref>) was proposed based on DeepLab to catch the differences in images. <xref ref-type="bibr" rid="B127">Wang Y. et&#xa0;al. (2021)</xref> proposed a feature-regularized mask DeepLab model to alleviate the overfitting problem caused by small-scale samples. <xref ref-type="bibr" rid="B11">Chen K. et&#xa0;al. (2018)</xref> introduced a shuffling operator based on the DeepLab model to improve the convolutional network.</p>
<p>DeepLabv3+ extends DeepLabv3, which added an effective decoder module to refine segmentation results (<xref ref-type="bibr" rid="B30">Du et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B51">Kong et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B4">Andrade et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B130">Wang et&#xa0;al., 2022a</xref>; <xref ref-type="bibr" rid="B125">Wang M. et&#xa0;al., 2022</xref>). <xref ref-type="bibr" rid="B130">Wang et&#xa0;al. (2022a)</xref> combined features by an attention mechanism based on DeepLabv3+, named CFAMNet. First, a feature module based on attention focused on the correlation between different classes. Then, the multi-parallel space pyramid pool structure extracted features of different scales of the input data. To correctly handle the imbalance problem between different classes, <xref ref-type="bibr" rid="B4">Andrade et&#xa0;al. (2022)</xref> extended the original DeepLabv3+ model, which can improve the depiction quality of forest polygons. <xref ref-type="bibr" rid="B125">Wang M. et&#xa0;al. (2022)</xref> proposed an improved DeepLabv3+ semantic segmentation network, adopting style differences in the generalization RS data in the backbone network ResNet101 using the Instance Batch Normalization (IBN) module.</p>
</sec>
<sec id="s3_1_5">
<label>3.1.5</label>
<title>Other CNN methods</title>
<p>The power of CNN is that its multi-layer structure can automatically learn features (<xref ref-type="bibr" rid="B53">Li Y. et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B90">Mi and Chen, 2020</xref>; <xref ref-type="bibr" rid="B81">Ma et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B148">Zhang Y. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B24">Cui H. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B83">Ma J. et&#xa0;al., 2022</xref>). <xref ref-type="bibr" rid="B24">Cui H. et&#xa0;al. (2022)</xref> proposed a novel method, Hybrid DA Network (MDANet), for patch image adaptation. It reduced the difference in projection distribution of different patch images by placing them into the virtual center of the hybrid domain. <xref ref-type="bibr" rid="B83">Ma J. et&#xa0;al. (2022)</xref> designed a progressive reconstruction block based on the ASPP block and used different proportions of atrous convolutional layers to continuously process features of different resolutions. Super pixel-enhanced deep neural forests (SDNFs) (<xref ref-type="bibr" rid="B90">Mi and Chen, 2020</xref>) achieved better classification accuracy by combining deep convolutional neural networks (DCNNs) with decision forests.</p>
<p>Compared with the large-scale coverage areas of RS images, key objects such as cars and ships in HRS images usually only contain a few pixels. To address this issue, <xref ref-type="bibr" rid="B81">Ma et&#xa0;al. (2021)</xref> designed a semantic segmentation model of small objects, named foreground activation (FA), which is from the perspective of structure and optimization. <xref ref-type="bibr" rid="B53">Li Y. et&#xa0;al. (2020)</xref> coupled CNN and graph neural network (GNN) design models to discover the spatial topological relationship between visual elements. A novel activation function Hard-Swish in (<xref ref-type="bibr" rid="B6">Avenash and Viswanath, 2019</xref>) obtained better accurate results. Some new methods with the CNN network, for example, <xref ref-type="bibr" rid="B142">Yang and Ma (2022)</xref>, proposed a sparse and complete latent structure <italic>via</italic> prototypes to solve the complex context of the background class. The weakly supervised method based on the CNN network can better solve tree species segmentation problems (<xref ref-type="bibr" rid="B2">Ahlswede et&#xa0;al., 2022</xref>).</p>
</sec>
<sec id="s3_1_6" sec-type="discussion">
<label>3.1.6</label>
<title>Discussion</title>
<p>The advantage of the FCN-based method is that it can adapt any input image size. Although the effect of 8 times up-sampling is much better than that of 32 times, the result of up-sampling is still relatively blurred and smooth, and it is not sensitive to the details of images. The classification of each pixel does not fully consider the relationship between pixels. The spatial regularization ignored spatial consistency. Since the model based on the U-Net structure does not add pads during the convolution process, two pixels are reduced after each convolution. The SegNet network uses pooling indices to save the contour features of the input image, reducing parameters. The DeepLab series performs ASPP, which improves the positioning of the target boundary by using DCNN and reduces the positioning accuracy caused by the invariance of DCNN.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Attention mechanism-based methods</title>
<p>The attention mechanism is a prevalent technique in DL methods (<xref ref-type="bibr" rid="B122">Vaswani et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B35">Fu et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B39">Guo et&#xa0;al., 2022</xref>). Excellent semantic segmentation models are often complex and require massive computing resources. In particular, the frequently used FCNs rely on detailed spatial and contextual information, which hinders their practical application. In DANet (<xref ref-type="bibr" rid="B35">Fu et&#xa0;al., 2019</xref>), rich information relations can be obtained through a dot-product operator. Although attention technology greatly improves segmentation accuracy, the requirement of massive computation resources also hinders its application. In recent years, more and more improved methods have emerged, such as the self-attention mechanism and fusion attention mechanism. This section summarizes and discusses linear attention and sub-attention mechanisms, and channel and spatial attention mechanisms.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Self-attention and linear attention</title>
<p>The input received by the neural network is many vectors of different sizes, and there is some relationship among them, but actual training cannot fully utilize these relationships. To solve the problem that the fully connected neural network cannot establish correlations for multiple related inputs, the self-attention operator emerged. It requires the machine to recognize the correlation between different components.</p>
<p>RSANet (<xref ref-type="bibr" rid="B154">Zhao D. et&#xa0;al., 2021</xref>) was a region self-attention mechanism. Compared with the traditional methods, it can decrease the feature noise and the redundant features. <xref ref-type="bibr" rid="B56">Li C. et&#xa0;al. (2021)</xref> employed a layered self-attention embedded neural network with dense connections, which made full use of short- and long-range contextual features. Self-attention models were learned for automatic learning of channel and position weights (<xref ref-type="bibr" rid="B13">Chen Z. et&#xa0;al., 2021</xref>) and built a feature library and extract features of class-constrained (<xref ref-type="bibr" rid="B26">Deng et&#xa0;al., 2021</xref>). <xref ref-type="bibr" rid="B56">Li C. et&#xa0;al. (2021)</xref> proposed the Multi-Scale Context Self-Attention Network (MSCSANet). It combined the benefits of self-attention and the mechanism of CNN to improve the segmentation quality. Through the position and channel attention modules, the correlation within the feature map was calculated as well as the multi-scale contextual feature map and local features.</p>
<p>Linear attention is an optimization genre of self-attention, which can optimize the complexity from O (N<sup>2</sup>) to O (N). The <italic>i</italic>th query feature is <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msubsup>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and the <italic>i</italic>th key feature is <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. The first-order representation of the Taylor expansion is <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msubsup>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2248;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and guarantees <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msubsup>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2265;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> using the L2 normal form for <italic>q<sub>i</sub>
</italic> and <italic>k<sub>j</sub>
</italic>.</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>q</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>k</mml:mtext>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where the sim(.) measures the similarity between <italic>q<sub>i</sub>
</italic> and <italic>k<sub>j</sub>
</italic>. Therefore,</p>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>Q</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>K</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>V</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where Q is the corresponding query matrix, K is the key matrix, and V is the value matrix. The vector form is represented as follows:</p>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>Q</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>K</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>V</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mi>K</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mi>V</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mi>Q</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mi>K</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<xref ref-type="bibr" rid="B65">Li et&#xa0;al. (2021a)</xref> used a linear attention mechanism (LAM). They reconstruct the skip connections of the original U-Net and design a multi-stage method. <xref ref-type="bibr" rid="B67">Li et&#xa0;al. (2021c)</xref> designed a novel Attention Bilateral Context Network (ABCNet), which utilizes a lightweight CNN spatial path and contextual path for semantic segmentation of high-resolution RS images and used a LAM modeling the global contextual information. A2-FPN (<xref ref-type="bibr" rid="B61">Li R. et&#xa0;al., 2022</xref>) was proposed for attention aggregation. The model introduces a LAM and an attention aggregation module for a feature pyramid network to enhance multi-scale feature learning. <xref ref-type="bibr" rid="B129">Wang L. et&#xa0;al. (2021)</xref> utilized stacked convolution to build the texture path and to fuse dependency and texture features. <xref ref-type="bibr" rid="B87">Marsocci et&#xa0;al. (2021)</xref> proposed a combined self-supervised algorithm using an attention mechanism and a semantic segmentation algorithm based on a LAM for the shape of aerial images.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Channel and spatial attention</title>
<p>The semantic segmentation methods in RS data widely used attention mechanisms, such as channel and spatial attention. The channel attention focuses on feature learning in the important channel dimensions and weakens others. The spatial attention module emphasizes key areas and weakens the background. Most of the current research methods combine these two methods to improve the segmentation effect, and some methods use one of them separately.</p>
<sec id="s3_2_2_1">
<label>3.2.2.1</label>
<title>Channel attention</title>
<p>Channel attention generates an attention mask in the channel domain to select important channels. Channel attention focuses on the channel dimension, which is shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7A</bold>
</xref>. A feature detector detected feature maps of each channel. For a feature map, the importance of each channel is calculated, and the weighted feature map is obtained by multiplying weights with the feature maps. <xref ref-type="bibr" rid="B112">Su et&#xa0;al. (2022)</xref> designed architecture similar to U-Net using wavelet frequency channel attention blocks as the attention mechanism. To select the most discriminative features, <xref ref-type="bibr" rid="B97">Panboonyuen et&#xa0;al. (2019)</xref> changed the weights of RS features at each stage to adaptively assign more weight values to important features. CFAMNet (<xref ref-type="bibr" rid="B130">Wang et&#xa0;al., 2022a</xref>) improved the deep DeepLabv3+ network. Its attention module obtained relevance between different categories. A multi-parallel ASPP extracted space relevance and obtained the context features of different scales.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>Channel attention and spatial attention (<xref ref-type="bibr" rid="B137">Woo et al., 2018</xref>). <bold>(A)</bold> Channel Attention. <bold>(B)</bold> Spatial attention. <bold>(C)</bold> The convolutional block attention combines channel attention and spatial attention.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g007.tif"/>
</fig>
</sec>
<sec id="s3_2_2_2">
<label>3.2.2.2</label>
<title>Spatial attention</title>
<p>Spatial attention focuses on the space and which points on each channel are more important (<xref ref-type="bibr" rid="B79">Luo et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B153">Zhao Q. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B63">Li et&#xa0;al., 2022b</xref>); thus, it is necessary to generate a spatial weight, which is shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7B</bold>
</xref>. First, average the values of different channels at the same plane space point (AvgPool) and take the maximum value (MaxPool) to obtain the weight. Then, a convolutional layer and a sigmoid function are used to obtain the final weight, and this weight is multiplied by each channel to achieve a weighted feature map in the spatial dimension. Owing to the size of the convolution kernel and the disappearing gradient, the data extracted from some buildings are inaccurate, and the information on some smaller buildings will be lost as the network deepens. A multi-scale spatial attention module (<xref ref-type="bibr" rid="B63">Li et&#xa0;al., 2022b</xref>) is designed to provide contextual information for the features obtained by this network model. A multi-scale spatial attention module provides contextual information for the features obtained by this network model. <xref ref-type="bibr" rid="B153">Zhao Q. et&#xa0;al. (2021)</xref> used a multi-scale module to advance the accuracy of high-resolution aerial labeling.</p>
</sec>
<sec id="s3_2_2_3">
<label>3.2.2.3</label>
<title>Fusion attention mechanism</title>
<p>Many experiments prove that fusing channel and spatial attention can get better segmentation results (<xref ref-type="bibr" rid="B27">Ding et&#xa0;al., 2020a</xref>; <xref ref-type="bibr" rid="B59">Li H. et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B115">Sun et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B108">Seong and Choi, 2021</xref>; <xref ref-type="bibr" rid="B32">Fan et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B74">Liu R. et&#xa0;al., 2022</xref>), which is shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7C</bold>
</xref>. There are two ways for the integration of channel and spatial attention: (1) parallel model, where channel attention and spatial attention are paralleled, and (2) sequential model. First, let the feature map pass channel attention and then pass spatial attention or vice versa. Most experiments prove that it is better to pass channel attention first. The bilateral segmentation network (BiSeNetV2) (<xref ref-type="bibr" rid="B115">Sun et&#xa0;al., 2020</xref>) includes a detailed branch and a semantic branch. The detailed branch uses wide channels and shallow layers to capture low-level details and generate high-resolution feature representations. It takes the feature map {C<sub>1</sub>, C<sub>2</sub>, C<sub>3</sub>, C<sub>4</sub>, and C<sub>5</sub>} as input. C<sub>1</sub> contains rich spatial location information, which is concatenated with Conv 1&#xd7;1 and C<sub>2</sub> to obtain feature map C<sub>12</sub> through convolution operation. Next, the spatial boundary attention map <italic>A</italic>
<sub>1</sub>=1/1(1+exp(C<sub>12</sub>)) is obtained through the sigmoid operation. The channel attention gate assigns weights according to the importance of each channel, and the spatial attention gate assigns weights according to the importance of each pixel location for the entire channel. <xref ref-type="bibr" rid="B27">Ding et&#xa0;al. (2020a)</xref> represented features in two ways <italic>via</italic> augmentation. On the one hand, the attention module is utilized to enhance embedding attention based on contextual information computed by local stitching. On the other hand, local foci from high-level features are embedded by the attention embedding module. <xref ref-type="bibr" rid="B32">Fan et&#xa0;al. (2022)</xref> fused channel and spatial attention; the attention module is combined with dilated convolutional layers to form a new central region encoding and decoding, which improves the accuracy of river segmentation. <xref ref-type="bibr" rid="B59">Li H. et&#xa0;al. (2020)</xref> proposed an end-to-end semantic segmentation network that integrates lightweight spatial and channel attention modules to adaptively refine features. Global relationships between different spatial positions or feature maps can be learned and reasoned by relation-augmented representations (<xref ref-type="bibr" rid="B92">Mou et&#xa0;al., 2020</xref>).</p>
</sec>
</sec>
<sec id="s3_2_3" sec-type="discussion">
<label>3.2.3</label>
<title>Discussion</title>
<p>Self-attention is the weight given to each input depending on the relationship between the input data. Self-attention has the advantage of parallel computing when calculating. Linear attention is similar to dot-product attention, but it uses less memory and computation. Channel attention focuses on the importance of different channels, while spatial attention gates focus on the importance of different pixel locations. In recent years, to improve semantic segmentation performance, most methods fuse channel and spatial attention mechanisms. However, researchers simply add or connect the attention results of the spatial and channel dimensions. How to identify the semantic segmentation of complex backgrounds is a problem that needs to be solved continuously. Therefore, it is necessary to design efficient fusion models to meet higher accuracy requirements.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Multi-scale strategy-based methods</title>
<p>RS images have high resolution and multi-scale variation characteristics. However, the receptive field size of the CNN is fixed. For the large-scale visual elements in the image, the receptive field can only cover its local area, which can easily cause wrong recognition results, and for the small-scale visual elements in the image. The challenge of exploiting multi-scale segmentation is to automatically select the best consecutive segmentation scale analysis (<xref ref-type="bibr" rid="B150">Zhang et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B160">Zhong et&#xa0;al., 2022a</xref>). Most methods are based on hierarchical structure or parallel structure, combined with an attention mechanism to achieve multi-scale feature fusion. This section discusses multi-scale semantic segmentation methods for RS images from hierarchical and parallel structures.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Hierarchical structure</title>
<p>The algorithm based on the hierarchical structure obtains multi-scale information through different stage features of the CNN, which is shown in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8A</bold>
</xref>. During the forward propagation process of the CNN, the receptive field increases continuously with the convolution and pooling operations. Multi-scale features from channel and spatial can be captured by fusing the features from CNN&#x2019;s different stages (<xref ref-type="bibr" rid="B158">Zheng et&#xa0;al., 2020a</xref>; <xref ref-type="bibr" rid="B52">Li Z. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B69">Liu B. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B80">Luo et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B132">Wang et&#xa0;al., 2022b</xref>; <xref ref-type="bibr" rid="B155">Zhao et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B157">Zheng et&#xa0;al., 2022</xref>).</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>The hierarchical and parallel structures of the multi-scale strategy (<xref ref-type="bibr" rid="B149">Zhang and Li, 2020</xref>).  <bold>(A)</bold> Hierarchical structure. <bold>(B)</bold> Parellel structure.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g008.tif"/>
</fig>
<p>High-resolution RS data have larger dimensions than typical natural images. <xref ref-type="bibr" rid="B92">Mou et al. (2020)</xref> studied a framework for object-specific optimization by identifying and fusing meaningful objects based on line segment tree models representing hierarchical multi-scale segmentation. Nodes in each path originate from leaf nodes. The EaNet model (<xref ref-type="bibr" rid="B159">Zheng et&#xa0;al., 2020b</xref>) is an edge-ware CNN. A kernel pyramid pooling (LKPP) module extracts different scale information. They designed a new loss function to optimize boundaries. <xref ref-type="bibr" rid="B157">Zheng et&#xa0;al. (2022)</xref> used a different scale input convolution module for extracting acceptable local information. <xref ref-type="bibr" rid="B52">Li Z. et&#xa0;al. (2021)</xref> extracted different features at multiple scales; SS AConv cascaded multi-scale structure (SCMS) transforms the SS AConv and residual correction scheme into a cascaded spatial pyramid by integrating different rates of SS AConv.</p>
<p>
<xref ref-type="bibr" rid="B139">Xu H. et&#xa0;al. (2022)</xref> designed the FSHRNet using strong linear separability of high-resolution features to achieve multi-scale object segmentation in VHR images. <xref ref-type="bibr" rid="B67">Li et&#xa0;al. (2021c)</xref> proposed a layered self-attention model with dense connections. The method made full use of short and long contextual features. Inspired by transfer learning, <xref ref-type="bibr" rid="B155">Zhao et&#xa0;al. (2022)</xref> improved a multi-scale network that can advance the network&#x2019;s robustness. It learned scale-invariant and small objects context information. <xref ref-type="bibr" rid="B69">Liu B. et&#xa0;al. (2022)</xref> designed a method that can efficiently extract different scale features and generate maps, which helps to subdivide objects into small and different sizes. <xref ref-type="bibr" rid="B80">Luo et&#xa0;al. (2022)</xref> extracted categorical object representations from multi-scale pixel features. It can identify the similarities and differences between categories. The article by <xref ref-type="bibr" rid="B158">Zheng et&#xa0;al. (2020a)</xref> learned the symbiotic relationship between scenes through the foreground-scene relationship module. Relevant context-associated foreground augments foreground functionality, thereby reducing false positives. <xref ref-type="bibr" rid="B132">Wang et al. (2022b)</xref> used dynamic multi-scale dilated convolution to extract different scale features.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Parallel structure</title>
<p>The parallel structure algorithm connects multiple parallel branches with different receptive fields after the semantic feature map obtained by the convolution module to form a parallel structure to capture features of different scales, which is shown in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8B</bold>
</xref>. <xref ref-type="bibr" rid="B75">Liu et&#xa0;al. (2018b)</xref> automatically learned multi-scale and multi-level features, which are obtained from a deep supervision network to provide comprehensive direct supervision to deal with various scenarios and scales of the road. <xref ref-type="bibr" rid="B68">Liu et&#xa0;al. (2018a)</xref> captured different scale contexts in the output results of CNN encoders, and then continuously aggregate them in a self-cascading manner. <xref ref-type="bibr" rid="B9">Bello et&#xa0;al. (2022)</xref> proposed an efficient dense multi-scale segmentation network for accurate and specialized remote real-time segmentation of RS images. <xref ref-type="bibr" rid="B131">Wang et&#xa0;al. (2022 2022)</xref> designed a new backbone network, taking multi-scale problems as an entry point, which can focus on more important information of multi-scales.</p>
<p>Because of the size of the CNN kernel and the vanishing gradient, the data extracted from buildings are inaccurate, and the information of some smaller buildings will be lost as the network deepens. <xref ref-type="bibr" rid="B31">Duan and Hu (2019)</xref> proposed a new erasure attention module to cooperate with the multi-scale refinement scheme to efficiently perform feature embedding.</p>
</sec>
<sec id="s3_3_3" sec-type="discussion">
<label>3.3.3</label>
<title>Discussion</title>
<p>The multi-scale strategy is a common technique for the semantic segmentation task of RS data. Since high-resolution images contain different object scales, it is necessary to combine the information of different scales of receptive fields to meet the requirements of the accurate segmentation of various objects. The FCN uses the same convolution operation on the entire image, without considering the multi-scale problem of visual elements, which damages the segmentation accuracy of larger-scale and smaller-scale visual elements. The multi-scale model generally builds a multi-scale RS image segmentation network first, then fuses multi-scale features, and finally predicts the results through convolution and up-sampling.</p>
<p>The method based on the hierarchical structure obtains multi-scale information through the features of different stages of the CNN. During the forward propagation process of the CNN, the receptive field expanded continuously using the pooling and convolution parts. The shallower feature map corresponds to a smaller receptive field, and the feature scale is also smaller. While the deep feature map corresponds to a larger receptive field, the feature scale is also larger. Therefore, different scale features can be obtained by fusing feature maps of different stages. The method based on parallel structure connects multiple parallel branches of different receptive fields after the semantic feature map obtained by the convolution module to form a parallel structure to capture features of different scales. These parallel branches are computed from the semantic feature map obtained by the convolution module, compared to the hierarchical algorithms, which are more suitable for learning semantic features.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Transformer-based methods</title>
<p>The Transformer was originally applied in the field of NLP. Each word is called a token in NLP, and in CV, the image is cut into non-overlapping patch sequences that are similar to tokens. SETR (<xref ref-type="bibr" rid="B156">Zheng et&#xa0;al., 2021</xref>) is the first representative model of semantic segmentation based on vision Transformer (ViT), which replaced the CNN encoder with a pure Transformer structure encoder. It drives the development of semantic segmentation in recent years.</p>
<p>Recently, Transformer technology makes significant contributions to improving semantic segmentation performance in the RS field (<xref ref-type="bibr" rid="B54">Li W. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B84">Ma L et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B117">Sun et&#xa0;al., 2022</xref>). However, compared with the words in the text, the pixels in the image have a very high resolution, and the computational complexity of using a Transformer in CV is the square of the image scale, which will lead to an excessively large amount of calculation. To solve the above problems, the Swin Transformer (ST) (<xref ref-type="bibr" rid="B71">Liu et&#xa0;al., 2021</xref>) network was proposed, which is shown in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9A</bold>
</xref>. Its features are learned by moving the window. The moving window not only brings greater efficiency but also greatly reduces the sequence length. The advantage of the hierarchical structure is that it flexibly provides information on various scales. Because self-attention can calculate within the window, its computational complexity increases linearly with the size of the picture rather than quadratic. Therefore, in RS semantic segmentation, it is widely used (<xref ref-type="bibr" rid="B98">Panboonyuen et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B140">Xu et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B33">Feng et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B38">Gu et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B72">Liu Y. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B62">Li X. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B89">Meng et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B141">Xu Y. et&#xa0;al., 2022</xref>). ST first used the module to segment the data into many non-overlapping different patches. The state-of-the-art solutions for segmentation tasks in RS data are usually solved by CNN methods and Transformer technology. A pre-trained ST (SwinTF) (<xref ref-type="bibr" rid="B98">Panboonyuen et&#xa0;al., 2021</xref>) model with ViT was used as the backbone to weight downstream tasks by concatenating task layers on the pre-trained encoder. The original ST as the backbone of the encoder module contains a convolutional layer and attention operator. <xref ref-type="bibr" rid="B62">Li X. et&#xa0;al. (2022)</xref> utilized ST blocks and convolution blocks to advance the segmentation performance. <xref ref-type="bibr" rid="B140">Xu et&#xa0;al. (2021)</xref> argued that Transformer-based architectures usually face two main problems: massive computational and difficulty of edge segmentation. Therefore, the authors proposed a new model based on a Transformer network to achieve accurate edge detection and fewer parameters. Use an efficient Transformer backbone to improve ST to reduce computational load. <xref ref-type="bibr" rid="B72">Liu Y. et&#xa0;al. (2022)</xref> designed UPer head with ST to challenge the land-cover segmentation.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>The Transformer unit and adaptive fusion module. <bold>(A)</bold> Swin Transformer (<xref ref-type="bibr" rid="B71">Liu et&#xa0;al., 2021</xref>). <bold>(B)</bold> Adaptive fusion module (<xref ref-type="bibr" rid="B36">Gao et al., 2021</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g009.tif"/>
</fig>
<p>CNN cannot simulate global semantic correlation, and the Transformer model can be built with global features (<xref ref-type="bibr" rid="B37">Ghali et&#xa0;al., 2021</xref>). Combining CNN and Transformer can improve the performance of semantic segmentation (<xref ref-type="bibr" rid="B152">Zhao X. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B124">Wang H. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B146">Zhang C. et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B145">Zhang et&#xa0;al., 2022a</xref>). CNN obtained local detail features and the Transformer module obtained the global context features. <xref ref-type="bibr" rid="B161">Zhong et&#xa0;al. (2022b)</xref> designed a semantic segmentation network, which combined CNN and Transformer parts. It solved over-segmentation and the inaccurate edge detection problem, which was caused by small differences between lakes and complex texture features. StransFuse (<xref ref-type="bibr" rid="B36">Gao et&#xa0;al., 2021</xref>) was a new method combining both advantages of the Transformer and the CNN model. It can better improve the performance of various RS images, which is shown in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9B</bold>
</xref>. Multi-level Transformers can fuse features in different levels in each modality and high-level cross-modal features (<xref ref-type="bibr" rid="B82">Ma X. et&#xa0;al., 2022</xref>).</p>
<p>The Transformer breaks through the limitation that the CNN model cannot be calculated in parallel and can reasonably utilize GPU resources. The Transformer&#x2019;s ability to acquire local information is not as strong as CNN&#x2019;s. Therefore, combining Transformer and CNN can improve semantic segmentation. The ST improves the ordinary Transformer and can be flexibly modeled at various scales using a layered architecture. The sliding window feature of the ST enables it to compute self-attention in locally non-overlapping windows and allows cross-window connections.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>GAN-based methods</title>
<p>Training neural networks, which largely depend on massive images with precise pixel-level annotation, is labor-intensive, especially for big-scale RS data. Segmenting multispectral images using supervised machine learning algorithms requires numerous pixel-level labeled data, which makes the task extremely challenging.</p>
<p>In recent years, some studies have introduced GAN into RS images for semantic segmentation tasks (<xref ref-type="bibr" rid="B22">Creswell et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B49">Kerdegari et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B41">Hong et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B58">Li D. et&#xa0;al., 2022</xref>). The GAN (<xref ref-type="bibr" rid="B22">Creswell et&#xa0;al., 2018</xref>) consists of generator (G) and discriminator (D) parts. The generator part can generate a fake image to fool the discriminator, and the discriminator distinguishes the fake image from the real image. The generator G transforms a random sample <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mi>d</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> distribution <italic>&#x3b3;</italic> into a generated sample G(z). The discriminator D discriminates them from the training samples from the distribution <italic>&#x3bc;</italic>, while G tries to make the generated samples&#x2019; distribution similar to that of the training samples. The adversarial target loss function is shown below:</p>
<disp-formula>
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>D</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mtext>G</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>:</mml:mo>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext>E</mml:mtext>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>log</mml:mi>
<mml:mtext>D</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>x</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">]</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>E</mml:mtext>
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">[</mml:mo>
<mml:mi>log</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>G</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtext>z</mml:mtext>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<xref ref-type="bibr" rid="B119">Tian et&#xa0;al. (2021)</xref> proposed a combined GAN and FCN network and constructed an FCN-based segmentation network to enhance the deep semantic receptive field of the model. GAN is integrated into an FCN semantic segmentation network to synthesize global image feature information and then accurately segment and sense complex RS images. <xref ref-type="bibr" rid="B41">Hong et&#xa0;al. (2020)</xref> proposed plug-and-play units in two networks: a self-generative GANs module and mutual GANs module, to learn perturbation-insensitive feature representations and eliminate multimodality, yielding more efficient and robust information transfers, respectively. <xref ref-type="bibr" rid="B114">Sun et&#xa0;al. (2021)</xref> proposed a subdivision method based on GANs to reduce intra-class differences. The background and target should be generated separately <italic>via</italic> the Orthogonal GAN (O-GAN). The O-GAN works by adding new loss functions to their discriminators. To better extract architectural features, the drawing is based on the idea of fine-grained image classification through an O-GAN intermediate convolutional layer (SCDA) with selective convolutional descriptor aggregation.</p>
<p>Because of the cumbersome and difficult annotation for RS images, the exploration of unsupervised and semi-supervised models is difficult. The domain adaptive method using the confrontation generation network learns domain-invariant features through the confrontation between the generator and the discriminator, which can effectively reduce the difference between domains. Most methods use GAN to generate RS images and combine them with network models such as CNN for semantic segmentation.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Fusion-based methods</title>
<p>As researchers continue to pursue the accuracy of semantic segmentation, a large number of fusion models of different technologies and structures have emerged, and have shown excellent results. Some fusion methods in recent years are listed in <xref ref-type="fig" rid="T1">
<bold>Table 1</bold>
</xref>. First, the CNN network is the basis of most models. Adding an attention mechanism block is the most common way in the research of fusion models (<xref ref-type="bibr" rid="B97">Panboonyuen et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B109">Shamsolmoali et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B51">Kong et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B70">Liu Z. et&#xa0;al., 2022</xref>). Second, different targets have different scales on the image. Therefore, multi-scale methods often integrate other feature extraction methods to improve models, such as CNN and Transformer (<xref ref-type="bibr" rid="B20">Chen et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B159">Zheng et&#xa0;al., 2020b</xref>; <xref ref-type="bibr" rid="B153">Zhao Q. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B145">Zhang et&#xa0;al., 2022a</xref>). Third, some complex models integrate more modules, such as GAN, ST, and multi-scale (<xref ref-type="bibr" rid="B52">Li Z. et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B87">Marsocci et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B140">Xu et&#xa0;al., 2021</xref>). However, complex models often require massive computing resources; thus, more models that balance computing resources and accuracy are needed.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Different methods that integrated different models.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Papers</th>
<th valign="middle" align="left">Year</th>
<th valign="middle" align="left">Datasets</th>
<th valign="middle" align="left">Methods</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B97">Panboonyuen et&#xa0;al. (2019)</xref>
</td>
<td valign="middle" align="left">2019</td>
<td valign="middle" align="left">ISPRS Vaihingen, Landsat-8 dataset</td>
<td valign="middle" align="left">CNN, Transfer learning, Attention mechanism</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B79">Luo et&#xa0;al. (2019)</xref>
</td>
<td valign="middle" align="left">2019</td>
<td valign="middle" align="left">fg</td>
<td valign="middle" align="left">CNN, Multi-scale, Self-attention</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B159">Zheng et&#xa0;al. (2020b)</xref> </td>
<td valign="middle" align="left">2020</td>
<td valign="middle" align="left">WHU building dataset, Cityscape, ISPRS Vaihingen</td>
<td valign="middle" align="left">CNN, Object-specific optimization, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B65">Li et&#xa0;al. (2021a)</xref>
</td>
<td valign="middle" align="left">2020</td>
<td valign="middle" align="left">ISPRS Vaihingen</td>
<td valign="middle" align="left">Attention mechanism, Multi-scale, ResU-Net</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B31">Duan and Hu (2019)</xref>
</td>
<td valign="middle" align="left">2020</td>
<td valign="middle" align="left">GID</td>
<td valign="middle" align="left">CNN, Multi-scale, Attention mechanism</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B110">Shao et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">2020</td>
<td valign="middle" align="left">WHDLD, DLRSD</td>
<td valign="middle" align="left">FCN, Region convolutional features, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B115">Sun et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">2020</td>
<td valign="middle" align="left">AIR-SEG, ISPRS Vaihingen</td>
<td valign="middle" align="left">FCN, Boundary attention model, Channel-weighted, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B20">Chen et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">ISPRS Potsdam, ISPRS Vaihingen</td>
<td valign="middle" align="left">GAN, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B108">Seong and Choi (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">SpaceNet building datasets, GIS, WHU dataset</td>
<td valign="middle" align="left">CNN, Attention mechanism, ResNet</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B119">Tian et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam, DeepGlobe Road</td>
<td valign="middle" align="left">FCN, GAN</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B98">Panboonyuen et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">ISPRS Vaihingen</td>
<td valign="middle" align="left">Feature pyramid network, CNN, Transformer</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B37">Ghali et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">Corsican Fire dataset</td>
<td valign="middle" align="left">CNN, Transformer, TransUNet, U2Net Architecture</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B13">Chen Z. et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">WHU and Massachusetts Building datasets</td>
<td valign="middle" align="left">U-Net, Self-attention, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B109">Shamsolmoali et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">DeepGlobe Road Extraction Data Set</td>
<td valign="middle" align="left">Feature pyramid, Multi-scale, Attention mechanism</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B56">Li C. et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Dense connection, Self-attention</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B140">Xu et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Swin Transformer</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B71">Kong et&#xa0;al. (2021)</xref> </td>
<td valign="middle" align="left">2021</td>
<td valign="middle" align="left">Sentinel-1 SAR images</td>
<td valign="middle" align="left">Channel spatial Attention mechanism, DeepLabv3+, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B128">Wang L. et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Potsdam, ISPRS Vaihingen</td>
<td valign="middle" align="left">Transformer, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B33">Feng et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">GID</td>
<td valign="middle" align="left">CNN, Swin Transformer, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B38">Gu et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">WHDLD, LoveDA</td>
<td valign="middle" align="left">CNN, Swin Transformer, U-Net, Multi-scale, A deformable adaptive patch merging layer</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B89">Meng et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">FCN, Swin Transformer</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B145">Zhang et&#xa0;al. (2022a)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Potsdam, WHU Building, dataset</td>
<td valign="middle" align="left">CNN, Transformer, Depthwise channel self-attention</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B70">Liu Z. et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">DCNN, Attention mechanism</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B62">Li X. et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">DeepGlobe Land Cover Dataset, ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Transformer, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B66">Li et&#xa0;al. (2021b)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Multi-attention network, Multi-scale,</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B155">Zhao et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Potsdam</td>
<td valign="middle" align="left">Collaborative enhanced fusion, Attention mechanism, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B80">Luo et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Potsdam, GID</td>
<td valign="middle" align="left">Feature pyramid, cross-attention, Transformer, Multi-scale</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B157">Zheng et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">GID</td>
<td valign="middle" align="left">Multi-scale, Transformer, Attention mechanism, semi-supervised, Pyramid scene parsing network</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B72">Liu Y. et al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Swin Transformer, Multi-scale, Dynamic attention pyramid head</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B40">He et al. (2022)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">ISPRS Vaihingen, ISPRS Potsdam</td>
<td valign="middle" align="left">CNN, Swin Transformer, UNet, Spatial interaction module</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B132">Wang et&#xa0;al. (2022b)</xref>
</td>
<td valign="middle" align="left">2022</td>
<td valign="middle" align="left">SSS image datasets</td>
<td valign="middle" align="left">CNN, Attention mechanism, Dynamic Multi-scale Dilated Convolution, Adaptive Receptive Field Mechanism</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B23">Cui L. et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">2023</td>
<td valign="middle" align="left">24 remote sensing city-scale images of Yushu city and Beichuan city after the Yushu and Wenchuan earthquakes</td>
<td valign="middle" align="left">CNN, Swin Transformer, Convolutional block attention module</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Dataset description and experimental discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Dataset description</title>
<p>We describe some public RS datasets for semantic segmentation tasks in this section. The most frequently referenced datasets are the ISPRS Vaihingen and Potsdam datasets, followed by GID and WHDLD. The image samples and classifications of these four datasets are shown in <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref>. We describe the datasets with more papers&#x2019; references, which include the description, classes, channels, and URLs shown in <xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary Table&#xa0;1</bold>
</xref>.</p>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>Visualization of the four common datasets. <bold>(A)</bold> ISPRS Vaihingen. <bold>(B)</bold> ISPRS Postdam. <bold>(C)</bold> WHDLD. <bold>(D)</bold> GID.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-11-1201125-g010.tif"/>
</fig>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Satellite image datasets</title>
<p>In this section, we list a few datasets for semantic segmentation tasks, which are captured by satellites. Satellite images are obtained from the earth observation remote sensing instrument loaded on the satellite.</p>
<sec id="s4_1_1_1">
<label>4.1.1.1</label>
<title>ISPRS Vaihingen</title>
<p>The ISPRS Vaihingen is a comparatively small village where there are many independent small buildings. The dataset contains 33 true orthophoto (TOP) images (GSD ~ 9&#xa0;cm) with 2,500 &#xd7; 2,000 pixels, which are of very high resolution. There are approximately 16 image tiles that are noted with pixel-wise labels. In addition, every pixel is split into one of six land categories, namely, impervious ground, architecture, low vegetation, tree, car, and clutter.</p>
</sec>
<sec id="s4_1_1_2">
<label>4.1.1.2</label>
<title>ISPRS Potsdam</title>
<p>The ISPRS Potsdam dataset covers an area of 3.42 km<sup>2</sup>, which consists of 38 image tiles with a spatial resolution of 5&#xa0;cm. All images are 6,000 &#xd7; 6,000 pixels with four bands for near-infrared (NIR) and red (R), green (G), and blue (B) channels.</p>
<p>Similar to the Vaihingen region, it is also made up of three bands of RS TIFF files and a single band of digital surveying and mapping (DSM). RS images and DSM are defined on the same reference system (UTM WGS84) because of the same coverage size of each RS image. In particular, each image is decomposed into smaller images using radial transform files. The dataset also provides a tiff storage form for different channel combinations of the TOP image so that participants can select their respective desired data.</p>
<p>The dataset label is a semi-dense disparity map obtained by averaging DSM data matched by multiple sets of commercial software based on internal and external orientation elements. The dataset provides a normalized DSM that does not require manual annotation. Accordingly, it is not guaranteed that there is no false data here, which is to help researchers use high data without using absolute DSM.</p>
</sec>
<sec id="s4_1_1_3">
<label>4.1.1.3</label>
<title>GID</title>
<p>GID (<xref ref-type="bibr" rid="B120">Tong et&#xa0;al., 2020</xref>) covers 506-km<sup>2</sup> areas that are captured <italic>via</italic> the satellite Gaofen-2. This dataset includes 150 high-quality Gaofen-2 RS images with 7,200 &#xd7; 6,800 pixels. The dataset has a rich diversity in spectrum, texture, and structure, which is very close to the real feature distribution characteristics. The GID dataset is divided into two parts: a large-scale set of labeling categories (GID-5) and a fine land cover set (GID-15). It contains five classes in GID-5. In addition, 150 image-level labeled Gaofen-2 satellite RS images are offered. Among them, there are 120 images in the training section; meanwhile, 30 images are included in the validation set.</p>
</sec>
<sec id="s4_1_1_4">
<label>4.1.1.4</label>
<title>WHDLD</title>
<p>The Wuhan dense labeling dataset (WHDLD) (<xref ref-type="bibr" rid="B110">Shao et&#xa0;al., 2020</xref>) is captured from an enormous image of the downtown area of Wuhan in the RS field. With a resolution of 2&#xa0;m, this dataset provides 4,940 RGB images with 256 &#xd7; 256 pixels. WHDLD is labeled with six categories. They are building, roads, sidewalks, vegetation, bare soil, and water.</p>
</sec>
<sec id="s4_1_1_5">
<label>4.1.1.5</label>
<title>DeepGlobe Land Cover</title>
<p>The dataset contains a space resolution of 0.5&#xa0;m and is built of red, green, and blue bands. It is generated from a satellite with 2,448 &#xd7; 2,448 pixels. Seven classes have been split into downtown area, farm land, range land, forest, water area, barren, and unknown.</p>
</sec>
<sec id="s4_1_1_6">
<label>4.1.1.6</label>
<title>GF-2</title>
<p>Based on the GF-2 satellite, this dataset has a space resolution of 0.8&#xa0;m, with 2,000 &#xd7; 2,000 pixels. With the help of ENVI, the image of GF-2 is preprocessed. These data are labeled by Matlab software with different colors and diverse image types.</p>
</sec>
<sec id="s4_1_1_7">
<label>4.1.1.7</label>
<title>RSSCN7</title>
<p>RSSCN7 (<xref ref-type="bibr" rid="B103">Qin et&#xa0;al., 2015</xref>) consists of 2,800 RS images. Collected from Google Earth, each class is equipped with 400 images with 400 &#xd7; 400 pixels. This dataset is split into seven different classes, namely, grass land, forest, farm land, parking lots, residential region, industrial region, and rivers/lakes.</p>
</sec>
<sec id="s4_1_1_8">
<label>4.1.1.8</label>
<title>LoveDA</title>
<p>The LoveDA dataset (<xref ref-type="bibr" rid="B133">Wang J. et&#xa0;al., 2021</xref>) collects different images of different cities and villages from Nanjing, Changzhou, and Wuhan, China. Along with a spatial resolution of 3&#xa0;m, this dataset offers 5,987 RS images. Each picture has a resolution of 1,024 &#xd7; 1,024. This dataset provides six categories, namely, building, roads, water, infertile soil, forest, and agriculture.</p>
</sec>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Aerial image datasets</title>
<p>We review a few RS semantic segmentation datasets captured by aircraft in this section. These data have the following characteristics: high definition, large scale, small area, and high visibility.</p>
<sec id="s4_1_2_1">
<label>4.1.2.1</label>
<title>Landcover</title>
<p>The Landcover aerial image labeling dataset consists of images from Poland&#x2019;s rural areas, from which there are 39.51 km<sup>2</sup> with a size of 50 cm/pixel and 176.76 km<sup>2</sup> with a resolution of 25 cm/pixel. These images are labeled with four classes. They are forests, water, building, and others.</p>
</sec>
<sec id="s4_1_2_2">
<label>4.1.2.2</label>
<title>UAVid</title>
<p>UAVid (<xref ref-type="bibr" rid="B143">Ye et&#xa0;al., 2020</xref>) is a UAV semantic segmentation dataset revolving around city street scenes with a resolution of 4,096 &#xd7; 2,160 and 3,840 &#xd7; 2,160. It contains 300 images intensively labeled with eight classes to cope with the semantic labeling task. The eight classes are architecture, urban road, tree, low vegetation, moving car, static car, human, and clutter/background. UAV is a quite challenging field due to the high resolution of images and the elaboration of scenes.</p>
</sec>
<sec id="s4_1_2_3">
<label>4.1.2.3</label>
<title>ISAID</title>
<p>This dataset is designed for instance segmentation (<xref ref-type="bibr" rid="B158">Zheng et&#xa0;al., 2020a</xref>), offering 2,806 high-resolution RS images from approximately 800 &#xd7; 800 pixels to approximately 4,000 &#xd7; 13,000 pixels with 15 foreground classes and 1 background class.</p>
</sec>
<sec id="s4_1_2_4">
<label>4.1.2.4</label>
<title>Massachusetts road datasets</title>
<p>The Massachusetts road dataset covers 2,600 km<sup>2</sup> of Massachusetts. This dataset consists of aerial images with a size of at least 1,500 &#xd7; 1,500 and a resolution of 1&#xa0;m. In addition, this dataset also provides seven pixels of ground segmentation truth collected from OpenStreetMap.</p>
</sec>
<sec id="s4_1_2_5">
<label>4.1.2.5</label>
<title>DLRSD</title>
<p>With a spatial size of 256 &#xd7; 256, DLRSE (<xref ref-type="bibr" rid="B110">Shao et&#xa0;al., 2020</xref>) consists of 2,100 RGB images and a resolution of 0.3&#xa0;m. The dataset is labeled based on the UCMerced LandUse dataset with 17 categories, namely, airplanes, bare soil, architecture, car, chaparral, courthouse, dock, field, grass, mobile house, sidewalk, sand, marine, ship, tank, trees, and water.</p>
</sec>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental comparison</title>
<p>Semantic segmentation methods for RS images are most commonly used for experimental comparisons on datasets ISPRS Vaihingen and ISPRS Potsdam. This paper summarizes the referenced RS semantic segmentation papers in the experimental comparison of the two as shown in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>, using the indicators mF1, mIoU, and OA.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Comparison of different methods on ISPRS Potsdam and Vaihingen datasets.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" rowspan="2" align="left">Methods</th>
<th valign="middle" rowspan="2" colspan="3" align="center">Models</th>
<th valign="middle" colspan="6" align="left">ISPRS Potsdam</th>
<th valign="middle" colspan="6" align="left">ISPRS Vaihingen</th>
</tr>
<tr>
<th valign="middle" colspan="2" align="left">mF1 (%)</th>
<th valign="middle" colspan="2" align="left">mIoU (%)</th>
<th valign="middle" colspan="2" align="left">OA (%)</th>
<th valign="middle" colspan="2" align="left">mF1 (%)</th>
<th valign="middle" colspan="2" align="left">mIoU (%)</th>
<th valign="middle" align="left" colspan="2">OA (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" rowspan="4" align="left">Based on CNN</td>
<td valign="top" colspan="3" align="left">EFCNet (<xref ref-type="bibr" rid="B10">Chen L. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">79.74</td>
<td valign="top" colspan="2" align="left">65.7</td>
<td valign="top" colspan="2" align="left">80.72</td>
<td valign="top" colspan="2" align="left">81.87</td>
<td valign="top" colspan="2" align="left">70.14</td>
<td valign="top" align="left">85.46</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SDFCNv2 (<xref ref-type="bibr" rid="B18">Chen G. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">67.82</td>
<td valign="top" colspan="2" align="left">85.03</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">&#x2013;</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">EGCAN (<xref ref-type="bibr" rid="B70">Liu Z. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">93</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">91.4</td>
<td valign="top" colspan="2" align="left">89.7</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">91</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">HCANet (<xref ref-type="bibr" rid="B8">Bai et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">88.07</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">88.92</td>
<td valign="top" colspan="2" align="left">88.94</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">89.71</td>
</tr>
<tr>
<td valign="top" rowspan="10" align="left">Attention Mechanism</td>
<td valign="top" colspan="3" align="left">MANet (<xref ref-type="bibr" rid="B66">Li et&#xa0;al., 2021b</xref>)</td>
<td valign="top" colspan="2" align="left">92.9</td>
<td valign="top" colspan="2" align="left">86.95</td>
<td valign="top" colspan="2" align="left">91.32</td>
<td valign="top" colspan="2" align="left">90.41</td>
<td valign="top" colspan="2" align="left">82.71</td>
<td valign="top" align="left">90.96</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">ABCNet (<xref ref-type="bibr" rid="B67">Li et&#xa0;al., 2021c</xref>)</td>
<td valign="top" colspan="2" align="left">92.7</td>
<td valign="top" colspan="2" align="left">86.5</td>
<td valign="top" colspan="2" align="left">91.3</td>
<td valign="top" colspan="2" align="left"/>
<td valign="top" colspan="2" align="left"/>
<td valign="top" align="left"/>
</tr>
<tr>
<td valign="top" colspan="3" align="left">A2-FPN (<xref ref-type="bibr" rid="B61">Li R. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">92.4</td>
<td valign="top" colspan="2" align="left">86.1</td>
<td valign="top" colspan="2" align="left">91.1</td>
<td valign="top" colspan="2" align="left">90.1</td>
<td valign="top" colspan="2" align="left">82.2</td>
<td valign="top" align="left">91</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">LANet (<xref ref-type="bibr" rid="B27">Ding et&#xa0;al., 2020a</xref>)</td>
<td valign="top" colspan="2" align="left">91.95</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">90.84</td>
<td valign="top" colspan="2" align="left">88.09</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">89.83</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SCAttNet (<xref ref-type="bibr" rid="B59">Li H. et&#xa0;al., 2020</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">68.31</td>
<td valign="top" colspan="2" align="left">87.97</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">&#x2013;</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">CAM-DFCN (<xref ref-type="bibr" rid="B79">Luo et&#xa0;al., 2019</xref>)</td>
<td valign="top" colspan="2" align="left">89.43</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">90.26</td>
<td valign="top" colspan="2" align="left">88.55</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">90.41</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">DSPCANet (<xref ref-type="bibr" rid="B55">Li YC. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">77.66</td>
<td valign="top" colspan="2" align="left">90.13</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">72.56</td>
<td valign="top" align="left">87.32</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">MARE (<xref ref-type="bibr" rid="B87">Marsocci et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">87.95</td>
<td valign="top" colspan="2" align="left">90.35</td>
<td valign="top" align="left">81.76</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">MAResUNet (<xref ref-type="bibr" rid="B65">Li et&#xa0;al., 2021a</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">90.28</td>
<td valign="top" colspan="2" align="left">83.3</td>
<td valign="top" align="left">90.86</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">EaNet (<xref ref-type="bibr" rid="B159">Zheng et&#xa0;al., 2020b</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">90.3</td>
<td valign="top" colspan="2" align="left"/>
<td valign="top" align="left">90.8</td>
</tr>
<tr>
<td valign="top" rowspan="4" align="left">Multi-scale Strategy</td>
<td valign="top" colspan="3" align="left">FSHRNet (<xref ref-type="bibr" rid="B139">Xu H. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">90.67</td>
<td valign="top" colspan="2" align="left">83.16</td>
<td valign="top" colspan="2" align="left">89.82</td>
<td valign="top" colspan="2" align="left">86.66</td>
<td valign="top" colspan="2" align="left">88.38</td>
<td valign="top" align="left">76.86</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SMAFNet (<xref ref-type="bibr" rid="B20">Chen et&#xa0;al., 2020</xref>)</td>
<td valign="top" colspan="2" align="left">88.18</td>
<td valign="top" colspan="2" align="left">71.31</td>
<td valign="top" colspan="2" align="left">86.77</td>
<td valign="top" colspan="2" align="left">86.91</td>
<td valign="top" colspan="2" align="left">65.28</td>
<td valign="top" align="left">88.45</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">MFNet (<xref ref-type="bibr" rid="B66">Li et&#xa0;al., 2021b</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">91.65</td>
<td valign="top" colspan="2" align="left">88.24</td>
<td valign="top" colspan="2" align="left">77.05</td>
<td valign="top" align="left">91.47</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">DGPRNet (<xref ref-type="bibr" rid="B148">Zhang Y. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">77.05</td>
<td valign="top" colspan="2" align="left">85.69</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">82.36</td>
<td valign="top" align="left">90.43</td>
</tr>
<tr>
<td valign="top" rowspan="8" align="left">Based on Transformer</td>
<td valign="middle" colspan="3" align="left">DC-Swin (<xref ref-type="bibr" rid="B128">Wang L. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">93.25</td>
<td valign="top" colspan="2" align="left">87.56</td>
<td valign="top" colspan="2" align="left">92</td>
<td valign="top" colspan="2" align="left">90.71</td>
<td valign="top" colspan="2" align="left">83.22</td>
<td valign="top" align="left">91.63</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">DHT-E (<xref ref-type="bibr" rid="B145">Zhang et&#xa0;al., 2022a</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">81.7</td>
<td valign="top" colspan="2" align="left">89.3</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">&#x2013;</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">ICTNet (<xref ref-type="bibr" rid="B62">Li X. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">93</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">91.57</td>
<td valign="top" colspan="2" align="left">92.34</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">90.14</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">MAT (<xref ref-type="bibr" rid="B152">Zhao X. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">91.59</td>
<td valign="top" colspan="2" align="left">84.82</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">88.7</td>
<td valign="top" colspan="2" align="left">79.93</td>
<td valign="top" align="left">&#x2013;</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SUDNet (<xref ref-type="bibr" rid="B141">Xu Y. et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">92.57</td>
<td valign="top" colspan="2" align="left">86.4</td>
<td valign="top" colspan="2" align="left">92.98</td>
<td valign="top" colspan="2" align="left">89.49</td>
<td valign="top" colspan="2" align="left">81.26</td>
<td valign="top" align="left">90.95</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">CG-Swin (<xref ref-type="bibr" rid="B89">Meng et&#xa0;al., 2022</xref>)</td>
<td valign="top" colspan="2" align="left">93.29</td>
<td valign="top" colspan="2" align="left">87.61</td>
<td valign="top" colspan="2" align="left">91.93</td>
<td valign="top" colspan="2" align="left">90.81</td>
<td valign="top" colspan="2" align="left">83.39</td>
<td valign="top" align="left">91.68</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SSAtNet (<xref ref-type="bibr" rid="B153">Zhao Q. et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">76.4</td>
<td valign="top" align="left">88.01</td>
</tr>
<tr>
<td valign="top" colspan="3" align="left">SwinTF-PSP (<xref ref-type="bibr" rid="B98">Panboonyuen et&#xa0;al., 2021</xref>)</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">94.83</td>
<td valign="top" colspan="2" align="left">90.98</td>
<td valign="top" align="left">&#x2013;</td>
</tr>
<tr>
<td valign="top" align="left">Based on GAN</td>
<td valign="top" colspan="3" align="left">Semi-supervised GAN (<xref ref-type="bibr" rid="B49">Kerdegari et&#xa0;al., 2019</xref>)</td>
<td valign="top" colspan="2" align="left">88.57</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" colspan="2" align="left">87.89</td>
<td valign="top" colspan="2" align="left">87.08</td>
<td valign="top" colspan="2" align="left">&#x2013;</td>
<td valign="top" align="left">88.34</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The values cannot rank the performance of the methods, because the training set and test size of different papers are different during the experiments. However, according to the comparison of different methods in the table, the overall performance of the method based on the attention mechanism and the Transformer mechanism is better than others.</p>
<p>Attention mechanisms are widely used in RS semantic segmentation, combining channel and spatial attention or multi-scale features to improve segmentation performance. The Transformer can perceive the global information of the input sequence, which is a huge advantage of the Transformer over CNN. In CNN, information can only start locally, and as the number of layers increases, the area that can be perceived gradually increases. However, the Transformer starts from the input, and each layer structure can see all the information and establish the association between the basic units, so it can handle more complex problems.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Discussion</title>
<p>The benefits and drawbacks of typical techniques are analyzed through the experimental outcomes combined with their characteristics, as shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. Researchers can use the strengths and weaknesses of the methods as a research reference to carry out future work.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>The advantages and disadvantages analysis and selection guidance for the existing methods.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Methods</th>
<th valign="middle" align="left">Advantages</th>
<th valign="middle" align="left">Disadvantages</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B18">Chen G. et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">The number of model parameters is small; it can excavate deep generalized features.</td>
<td valign="middle" align="left">Rely on a large number of training datasets.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B12">Chen et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">It takes much less running time.</td>
<td valign="middle" align="left">Need to focus on unsupervised learning.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B1">Abdollahi et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">It takes forward and backward dependencies into account and considers all the information.</td>
<td valign="middle" align="left">Need to do multi-object segmentation from remote sensing data simultaneously.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B34">Foivos et al. (2020)</xref>. </td>
<td valign="middle" align="left">Tanimoto loss results in balanced gradients can be used for regression problems</td>
<td valign="middle" align="left">Due to the original image being reduced, the fine details of the trees cannot be recognized</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B136">Weng et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">It reduces a large number of parameters; The training speed is high.</td>
<td valign="middle" align="left">Missed detections and false alarms, achieved poor water-body extraction results without complete water-body boundaries.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B30">Du et&#xa0;al. (2014)</xref>
</td>
<td valign="middle" align="left">It can alleviate the retention of accurate boundary information on ground objects.</td>
<td valign="middle" align="left">The recognition accuracy of objects with large scale is not high.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B67">Li et&#xa0;al. (2021c)</xref>
</td>
<td valign="middle" align="left">It can obtain detailed spatial and contextual information. It reduces the parameter number.</td>
<td valign="middle" align="left">It is dependent on fully convolutional networks.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B55">Li Y. C. et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">It can extract effective spectral and spatial enhancement features.</td>
<td valign="middle" align="left">Need to focus on the multi-scale convolution in different topologies.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B157">Zheng et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">It combines the advantages of merging Transformer and CNN to get local and global features.</td>
<td valign="middle" align="left">Obtains more refined object information</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B52">Li Z. et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">It draws contextual information and refines objects at dense multi-scales.</td>
<td valign="middle" align="left">It leads to a decreased performance in the recovery of edges of very thin semantics.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B63">Li et&#xa0;al. (2022b)</xref>
</td>
<td valign="middle" align="left">It can better identify dense buildings and small targets.</td>
<td valign="middle" align="left">Need to automatic enhancement of training data.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B84">Ma L. et&#xa0;al. (2022)</xref>
</td>
<td valign="middle" align="left">It has effective attention weight enhancement and edge convolutions for powerful local feature encodings.</td>
<td valign="middle" align="left">Missing validation results on other remote sensing datasets.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B140">Xu et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">It can better solve the problem of high computing load and blurred edges.</td>
<td valign="middle" align="left">Boundary detection is not well resolved.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B36">Gao et&#xa0;al. (2021)</xref>
</td>
<td valign="middle" align="left">Avoiding gradient disappearance and feature map information loss.</td>
<td valign="middle" align="left">The algorithm structure is complex.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B96">Pan et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">The model can generate ground-truth data by controlling the numeral and scope of samples.</td>
<td valign="middle" align="left">Need to use supervised training data to fit the parameters.</td>
</tr>
<tr>
<td valign="middle" align="left">
<xref ref-type="bibr" rid="B41">Hong et&#xa0;al. (2020)</xref>
</td>
<td valign="middle" align="left">It eliminates the gap between modalities and obtains a smoother and more detailed appearance in urban scene parsing.</td>
<td valign="middle" align="left">Massive labeled RS images are required for its training.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusion and future direction</title>
<p>This paper reviews the state-of-the-art progress in semantic segmentation of RS images, which summarizes them from the angle of DL framework and technology. The earliest CNN-based classical methods were applied to the semantic segmentation task in the field of RS data and achieved good experimental results. Next, with burgeoning technologies such as attention mechanism, multi-scale, Transformer, and GAN, the performance of high-pixel semantic segmentation is improved. Integrating multiple techniques is a wise choice for researchers, which enables the progress of both the accuracy and efficiency of the segmentation.</p>
<p>After an in-depth study of semantic segmentation techniques, we found that although researchers have made effective efforts, there are still many challenges in this research, and further efforts are required in future work.</p>
<list list-type="bullet">
<list-item>
<p>High-resolution RS images require manual pixel labeling, which is arduous and labor-intensive. Therefore, the problem of insufficient samples still exists. Future work can be improved in the following aspects: (1) how to construct multi-angle, multi-tone, and other sample analysis models; (2) exploring approaches to achieve more promising performance, rarely using fine annotation or rough brands, and reducing training samples; and (3) merging datasets and combining different optical and SAR datasets. Robust Transformer models can be explored for multi-source RS data, which comprise aerial and satellite images with diverse spatial and spectral resolutions.</p>
</list-item>
<list-item>
<p>Optimize and improve the semantic segmentation models. Semantic segmentation technology can directly promote the development of smart cities, resource monitoring, and other fields. These tasks generate a higher demand for models. (1) How to better capture more differentiated features and context information for its high-resolution images. (2) How to design unsupervised learning models for improving the performance of high-resolution images, including weakly supervised and semi-supervised methods, which do not require a large amount of labeled data. (3) Change the number or types of convolutions in convolutional models. (4) How to replace the edge-guided context aggregation method and use better edge extractors in explicit augmentation methods.</p>
</list-item>
<list-item>
<p>Reduce the computational complexity and improve the robustness of the model. It is important to improve the performance and quality of the existing models, which are large and computationally intensive and hinder their wide application. How to balance the performance and computer power of semantic segmentation is a future research direction. (1) Build real-time semantic segmentation models with less model size and computational complexity. (2) Design a more efficient and concise feature extraction method. (3) Reduce latency.</p>
</list-item>
<list-item>
<p>Research on more complex actual scenarios. Many experiments are only implemented on specific datasets. Therefore, how to design new methods that can be suitable for actual complex scenarios remains to be studied.</p>
</list-item>
<list-item>
<p>Research on small target segmentation. Owing to the small proportion of the pixel area of the small target, a certain amount of detailed information will be lost after multiple down-sampling, which will give rise to an accuracy decrease to a certain extent. In the future, we can start with small targets and improve accuracy with methods such as residual connections, attention mechanisms, and pyramid structures.</p>
</list-item>
</list>
<p>Unfortunately, since semantic segmentation of RS images is a hot research field, a large number of research methods have emerged in recent years and are constantly updated, so it is difficult for us to find all semantic segmentation methods. In the future, researchers&#x2019; attention should be directed to new methods and theories for semantic segmentation of RS images.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>Conceptualization and methodology, JL; investigation and resources, JL, ML, and LS; data curation, YL and PZ; writing&#x2014;original draft preparation, JL, ML, and YL; writing&#x2014;review and editing: QS. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>This work was partially supported by the R&amp;D Program of Beijing Municipal Education Commission (No. KM202211417014), the Academic Research Projects of Beijing Union University (No. ZK20202215), the Natural Science Foundation of Shandong Province under Grant ZR2022LZH015 (ZR2020MF006), the Industry&#x2013;University Research Innovation Foundation of Ministry of Education of China under Grant (2021FNA01001), and the Shandong Provincial Natural Science Foundation, China under Grant ZR2020MF006 and ZR2022LZH015.</p>
</sec>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fevo.2023.1201125/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fevo.2023.1201125/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="DataSheet_1.pdf" id="SM1" mimetype="application/pdf"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abdollahi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Pradhan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Shukla</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Chakraborty</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Alamri</surname> <given-names>A. M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multi-object segmentation in complex urban scenes from high-resolution remote sensing data</article-title>. <source>Remote Sens.</source> <volume>13</volume>, <elocation-id>3710</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/RS13183710</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ahlswede</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Madam</surname> <given-names>N. T.</given-names>
</name>
<name>
<surname>Schulz</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Kleinschmit</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Demir</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Weakly supervised semantic segmentation of remote sensing images for tree species classification based on explanation methods</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS</conf-name>, <conf-loc>Kuala Lumpur, Malaysia</conf-loc>. <publisher-name>IEEE</publisher-name>, Vol. <volume>2022</volume>. <fpage>4847</fpage>&#x2013;<lpage>4850</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS46834.2022.9884676</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aleissaee</surname> <given-names>A. A.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Anwer</surname> <given-names>R. M.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Cholakkal</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>G. S.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Transformers in remote sensing: a survey</article-title>. <source>arXiv</source> <volume>2209</volume>, <fpage>01206</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs15071860</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Andrade</surname> <given-names>R. B.</given-names>
</name>
<name>
<surname>Mota</surname> <given-names>G. L. A.</given-names>
</name>
<name>
<surname>da Costa</surname> <given-names>G. A. O. P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Deforestation detection in the Amazon using DeepLabv3+ semantic segmentation model variants</article-title>. <source>Remote. Sens.</source> <volume>14</volume> (<issue>19</issue>), <elocation-id>4694</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14194694</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Asokan</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Anitha</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Change detection techniques for remote sensing applications: a survey</article-title>. <source>Earth Sci. Inf.</source> <volume>12</volume>, <fpage>143</fpage>&#x2013;<lpage>160</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s12145-019-00380-5</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Avenash</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Viswanath</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Semantic segmentation of satellite images using a modified CNN with hard-swish activation function</article-title>,&#x201d; in <conf-name>Proceedings of the VISIGRAPP 2019</conf-name>. <conf-loc>Prague, Czech Republic</conf-loc>. <fpage>413</fpage>&#x2013;<lpage>420</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.5220/0007469604130420</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Badrinarayanan</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Kendall</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Cipolla</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Segnet: a deep convolutional encoder-decoder architecture for image segmentation</article-title>. <source>IEEE T. Pattern. Anal.</source> <volume>39</volume>, <fpage>2481</fpage>&#x2013;<lpage>2495</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2644615</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bai</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>S. Y.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>C. J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>HCANet: a hierarchical context aggregation network for semantic segmentation of high-resolution remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2021.3063799</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bello</surname> <given-names>I. M.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. Y.</given-names>
</name>
<name>
<surname>Aslam</surname> <given-names>M. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Densely multiscale framework for segmentation of high resolution remote sensing imagery</article-title>. <source>Comput. Geosci.</source> <volume>167</volume>, <elocation-id>105196</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cageo.2022.105196</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Dou</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W. B.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>B. Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H. F.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>EFCNet: ensemble full convolutional network for semantic segmentation of high-resolution remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2021.3076093</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Semantic segmentation of aerial images with shuffling convolutional neural networks</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>15</volume>, <fpage>173</fpage>&#x2013;<lpage>177</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2017.2778181</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>G.</given-names>
</name>
<name>
<surname>He</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>P. Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X. D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A superpixel-guided unsupervised fast semantic segmentation method of remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2022.3198065</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>H. Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Self-attention in reconstruction bias U-net for semantic segmentation of building rooftops in optical remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>13</volume> (<issue>13</issue>), <elocation-id>2524</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13132524</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Transunet: Transformers make strong encoders for medical image segmentation</article-title>. <source>arXiv</source> <volume>2102</volume>, <elocation-id>4306</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2102.04306</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Kokkinos</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Murphy</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yuille</surname> <given-names>A. L.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>Semantic image segmentation with deep convolutional nets and fully connected CRFs</article-title>,&#x201d; in <conf-name>Proceedings of the 2015 ICLR</conf-name>, <conf-loc>San Diego, CA, USA</conf-loc>, Vol. <volume>4</volume>. <fpage>357</fpage>&#x2013;<lpage>361</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Kokkinos</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Murphy</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yuille</surname> <given-names>A. L.</given-names>
</name>
</person-group> (<year>2017</year>a). <article-title>Deeplab: semantic image segmentation with deep convolutional nets,atrous convolution, and fully connected crfs</article-title>. <source>IEEE T. Pattern. Anal.</source> <volume>40</volume>, <fpage>834</fpage>&#x2013;<lpage>848</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2017.2699184</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Schroff</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Adam</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>b). <article-title>Rethinking atrous convolution for semantic image segmentation</article-title>. <source>arXiv</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.1706.05587</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>SDFCNv2: an improved FCN framework for remote sensing images semantic segmentation</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>4902</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13234902</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Schroff</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Adam</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the ECCV 2018</conf-name>, <conf-loc>Munich, Germany</conf-loc>, <publisher-name>Springer</publisher-name>, Vol. <volume>2</volume>. <fpage>801</fpage>&#x2013;<lpage>818</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_49</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J. H.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>SMAF-net: sharing multiscale adversarial feature for high-resolution remote sensing imagery semantic segmentation</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>18</volume>, <fpage>1921</fpage>&#x2013;<lpage>1925</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2020.3011151</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ciresan</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Giusti</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Gambardella</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Schmidhuber</surname>
</name>
</person-group>. (<year>2012</year>). <article-title>Deep neural networks segment neuronal membranes in electron microscopy images</article-title>. <source>NIPS</source> <volume>2012</volume>, <fpage>2825</fpage>&#x2013;<lpage>2860</lpage>.</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Creswell</surname> <given-names>A.</given-names>
</name>
<name>
<surname>White</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Dumoulin</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Arulkumaran</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Sengupta</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Bharath</surname> <given-names>A. A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Generative adversarial networks: an overview</article-title>. <source>IEEE Signal Proc. Mag.</source> <volume>35</volume>, <fpage>53</fpage>&#x2013;<lpage>65</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/MSP.2017.2765202</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cui</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Jing</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Huan</surname> <given-names>Y. X.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Q. Q.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Improved Swin Transformer-based semantic segmentation of postearthquake dense buildings in urban areas using remote sensing images</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>16</volume>, <fpage>369</fpage>&#x2013;<lpage>385</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2022.3225150</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cui</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Qi</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H. F.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>MDANet: unsupervised, mixed-domain adaptation for semantic segmentation of remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13030454</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Davis</surname> <given-names>L. S.</given-names>
</name>
<name>
<surname>Rosenfeld</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Weszka</surname> <given-names>J. S.</given-names>
</name>
</person-group> (<year>1975</year>). <article-title>Region extraction by averaging and thresholding</article-title>. <source>IEEE T. Syst. Man. CY-S</source> <volume>1975</volume>, <fpage>3, 383</fpage>&#x2013;<lpage>3, 388</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/tsmc.1975.5408419</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>M. Z.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>Y. F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>CCANet: class-constraint coarse-to-fine attentional deep network for subdecimeter aerial image semantic segmentation</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2021.3055950</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Bruzzone</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>a). <article-title>LANet: local attention embedding to improve the semantic segmentation of remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>59</volume>, <fpage>426</fpage>&#x2013;<lpage>435</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.2994150</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Bruzzone</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>b). <article-title>Semantic segmentation of large-size VHR remote sensing images using a two-stage multiscale training architecture</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>58</volume>, <fpage>5367</fpage>&#x2013;<lpage>5376</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.2964675</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A multi-level feature fusion network for remote sensing image segmentation</article-title>. <source>Sensors</source> <volume>21</volume>, <elocation-id>1267</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/s21041267</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X. Y.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Incorporating DeepLabv3+ and object-based image analysis for semantic segmentation of very high resolution remote sensing images</article-title>. <source>Int. J. Digit. Earth</source> <volume>14</volume>, <fpage>357</fpage>&#x2013;<lpage>378</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1080/17538947.2020.1831087</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Multiscale refinement network for water-body segmentation in high-resolution satellite imagery</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>17</volume>, <fpage>686</fpage>&#x2013;<lpage>690</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2019.2926412</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname> <given-names>Z. Y.</given-names>
</name>
<name>
<surname>Hou</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Zang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Y. J.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>River segmentation of remote sensing images based on composite attention network</article-title>. <source>Complex</source> <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2022/7750281</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Feng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A semantic segmentation method for remote sensing images based on the Swin Transformer fusion gabor filter</article-title>. <source>IEEE Access</source> <volume>10</volume>, <fpage>77432</fpage>&#x2013;<lpage>77451</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2022.3193248</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Foivos</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Diakogiannis</surname> <given-names>F. W.</given-names>
</name>
<name>
<surname>Peter</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data</article-title>. <source>ISPRS J. Photogramm.</source> <volume>162</volume>, <fpage>94</fpage>&#x2013;<lpage>114</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.01.013</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Bao</surname> <given-names>Y. J.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>Z. W.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). &#x201c;<article-title>Dual attention network for scene segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2019</conf-name>. <conf-loc>Long Beach, CA, USA</conf-loc>, <publisher-name>IEEE</publisher-name>, <fpage>3146</fpage>&#x2013;<lpage>3154</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2019.00326</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wan</surname> <given-names>Y. L.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>Z. Q.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>STransFuse: fusing Swin Transformer and convolutional neural network for remote sensing image semantic segmentation</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens.</source> <volume>14</volume>, <fpage>10990</fpage>&#x2013;<lpage>11003</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2021.3119654</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ghali</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Akhloufi</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Jmal</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Mseddi</surname> <given-names>W. S.</given-names>
</name>
<name>
<surname>Attia</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Wildfire segmentation using deep vision Transformers</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>3527</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13173527</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>H. B.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>C. C.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>H. L.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Adaptive enhanced Swin Transformer with U-net for remote sensing image segmentation</article-title>. <source>Comput. Electr. Eng.</source> <volume>102</volume>, <elocation-id>108223</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.compeleceng.2022.108223</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname> <given-names>M. H.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>T. X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z. N.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>P. T.</given-names>
</name>
<name>
<surname>Mu</surname> <given-names>T. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Attention mechanisms in computer vision: a survey</article-title>. <source>Comput. Visual Media</source> <volume>8</volume>, <fpage>331</fpage>&#x2013;<lpage>368</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s41095-022-0271-y</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Xue</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Swin Transformer embedding UNet for remote sensing image semantic segmentation</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>60</volume>, <fpage>4408715</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2022.3144165</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hong</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Z. B.</given-names>
</name>
<name>
<surname>Chanussot</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multimodal GANs: toward crossmodal hyperspectral&#x2013;multispectral image segmentation</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>59</volume>, <fpage>5103</fpage>&#x2013;<lpage>5113</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.3020823</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Cai</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z. Y.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>A semantic segmentation approach based on deepLab network in high-resolution remote sensing images</article-title>,&#x201d; in <conf-name>Proceedings of the ICIG 2019</conf-name>, <conf-loc>Beijing, China</conf-loc>. <publisher-name>Springer</publisher-name>, <fpage>292</fpage>&#x2013;<lpage>304</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-34113-8_25</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tong</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>H. J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Q. W.</given-names>
</name>
<name>
<surname>Iwamoto</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>UNet 3+: a full-scale connected UNet for medical image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the ICASSP 2020</conf-name>, <conf-loc>Barcelona</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>1055</fpage>&#x2013;<lpage>1059</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICASSP40776.2020.9053405</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>van der Maaten</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Weinberger</surname> <given-names>K. Q.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Densely connected convolutional networks</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2017</conf-name>, <conf-loc>Honolulu, Hawaii, USA</conf-loc>. <fpage>4700</fpage>&#x2013;<lpage>4708</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2017.243</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Audeberr</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Khalel</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Tarabalka</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Malof</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). &#x201c;<article-title>Large-Scale semantic classification: outcome of the first year of inria aerial image labeling benchmark</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS 2018</conf-name>, <conf-loc>Valencia, Spain</conf-loc>. <fpage>6947</fpage>&#x2013;<lpage>6950</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS.2018.8518525</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Iglovikov</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Seferbekov</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Buslaev</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shvets</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Ternausnetv2: fully convolutional network for instance segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2018</conf-name>, <conf-loc>Munich, Germany</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>233</fpage>&#x2013;<lpage>237</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPRW.2018.00042</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>H. F.</given-names>
</name>
<name>
<surname>Hao</surname> <given-names>Z. M.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>J. M.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>A survey on deep learning-based change detection from high-resolution remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>1552</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14071552</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kemker</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Salvaggio</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Kanan</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Algorithms for semantic segmentation of multispectral remote sensing imagery using deep learning</article-title>. <source>ISPRS J. Photogramm.</source> <volume>145</volume>, <fpage>60</fpage>&#x2013;<lpage>77</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2018.04.014</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kerdegari</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Razaak</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Argyriou</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Remagnino</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Urban scene segmentation using semi-supervised GAN</article-title>,&#x201d; in <conf-name>Proceedings of the Image and Signal Processing for Remote Sensing</conf-name>, <conf-loc>Denver, USA</conf-loc>. <fpage>477</fpage>&#x2013;<lpage>484</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/12.2533055</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kitaev</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Klein</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Learned incremental representations for parsing</article-title>,&#x201d; in <conf-name>Proceedings of the ACL 2022</conf-name>, <conf-loc>Xiangcheng, China</conf-loc>. <fpage>3086</fpage>&#x2013;<lpage>3095</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.18653/v1/2022.acl-long.220</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kong</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Leung</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>X. Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A novel deeplabv3+ network for sar imagery semantic segmentation based on the potential energy loss function of gibbs distribution</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>454</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13030454</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z. Q.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z. H.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Cascaded multiscale structure with self-smoothing atrous convolution for semantic segmentation</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2021.3088902</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multi-label remote sensing image scene classification by combining a convolutional neural network and a graph neural network</article-title>. <source>Remote Sens.</source> <volume>12</volume>, <elocation-id>4003</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs12234003</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Momanyi</surname> <given-names>B. M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Unsupervised domain adaptation for remote sensing semantic segmentation with Transformer</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>4942</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14194942</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y. C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H. C.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>W. S.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>H. L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>DSPCANet: dual-channel scale-aware segmentation network with position and channel attentions for high-resolution aerial images</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>14</volume>, <fpage>8552</fpage>&#x2013;<lpage>8565</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2021.3102137</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Lyu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Tong</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Hierarchical self-attention embedded neural network with dense connection for remote-sensing image semantic segmentation</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>126623</fpage>&#x2013;<lpage>126634</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2021.3111899</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>System dynamics simulation and regulation of human-water system coevolution in Northwest China</article-title>. <source>Front. Ecol. Evol.</source> <volume>10</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/FEVO.2022.1106998</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W. H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>A. D.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>W. F.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). &#x201c;<article-title>A dual-fusion semantic segmentation framework with gan for sar images</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS</conf-name>, Vol. <volume>2022</volume>. <fpage>991</fpage>&#x2013;<lpage>994</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS46834.2022.9884931</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Qiu</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Mei</surname> <given-names>X. M.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>SCAttNet: semantic segmentation network with spatial and channel attention mechanism for high-resolution remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>18</volume>, <fpage>905</fpage>&#x2013;<lpage>909</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2020.2988294</pub-id>
</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y. X.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Z. K.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>D. D.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>a). <article-title>Deep learning-based object detection techniques for remote sensing images: a survey</article-title>. <source>Remote Sens. Remote. Sens.</source> <volume>14</volume>, <elocation-id>2385</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14102385</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>S. Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A2-FPN for semantic segmentation of fine-resolution remotely sensed images</article-title>. <source>Remote Sens.</source> <volume>43</volume> (<issue>3</issue>), <fpage>1131</fpage>&#x2013;<lpage>1155. 16</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1080/01431161.2022.2030071</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>R. L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z. Q.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X. Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Encoding contextual information by interlacing Transformer and convolution for remote sensing imagery semantic segmentation</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>4065</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14164065</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L. Q.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Q.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>b). <article-title>HCRB-MSAN: horizontally connected residual blocks-based multiscale attention network for semantic segmentation of buildings in HSR remote sensing images</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>15</volume>, <fpage>5534</fpage>&#x2013;<lpage>5544</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2022.3188515</pub-id>
</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H. K.</given-names>
</name>
<name>
<surname>Xue</surname> <given-names>X. Z.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>). <article-title>Deep learning for remote sensing image classification: a survey</article-title>. <source>WIREs Data Min. Knowl. Discov.</source> <volume>8</volume>, <elocation-id>e1264</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/widm.1264</pub-id>
</citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>S. Y.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>J. L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>a). <article-title>Multistage attention ResU-net for semantic segmentation of fine-resolution remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2021.3063381</pub-id>
</citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>S. Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>J. L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L. B.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>b). <article-title>Multiattention network for semantic segmentation of fine-resolution remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2021.3093977</pub-id>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
</person-group> (<year>2021</year>c). <article-title>ABCNet: attentive bilateral contextual network for efficient semantic segmentation of fine-resolution remotely sensed imagery</article-title>. <source>ISPRS J. Photogramm.</source> <volume>181</volume>, <fpage>84</fpage>&#x2013;<lpage>98.46</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2021.09.005</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Bai</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xiang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>a). <article-title>Semantic labeling in very high resolution images <italic>via</italic> a self-cascaded convolutional neural network</article-title>. <source>ISPRS J. Photogramm.</source> <volume>145</volume>, <fpage>78</fpage>&#x2013;<lpage>95</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2017.12.007</pub-id>
</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>J. W.</given-names>
</name>
<name>
<surname>Bi</surname> <given-names>X. L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W. S.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>X. B.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>PGNet: positioning guidance network for semantic segmentation of very-High-Resolution remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>4219</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14174219</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z. Q.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Edge guided context aggregation network for semantic segmentation of remote sensing imagery</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>1353</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14061353</pub-id>
</citation>
</ref>
<ref id="B71">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Y. T.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y. X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Swin Transformer: hierarchical vision Transformer using shifted windows</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2021</conf-name>, <conf-loc>Kuala Lumpur, Malaysia</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>10012</fpage>&#x2013;<lpage>10022</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00986</pub-id>
</citation>
</ref>
<ref id="B72">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y. H.</given-names>
</name>
<name>
<surname>Mei</surname> <given-names>S. H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>He</surname> <given-names>M. Y.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Semantic segmentation of high-resolution remote sensing images using an improved Transformer</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS 2022</conf-name>. <publisher-loc>Kuala Lumpur, Malaysia</publisher-loc>, <publisher-name>IEEE</publisher-name>, <fpage>3496</fpage>&#x2013;<lpage>3499</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS46834.2022.9884103</pub-id>
</citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Piramanayagam</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Monteiro</surname> <given-names>S. T.</given-names>
</name>
<name>
<surname>Saber</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Semantic segmentation of multisensor remote sensing imagery with deep ConvNets and higher-order conditional random fields</article-title>. <source>J. Appl. Remote Sens.</source> <volume>13</volume>, <elocation-id>1</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/1.JRS.13.016501</pub-id>
</citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>R. R.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>X. T.</given-names>
</name>
<name>
<surname>Na</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Leng</surname> <given-names>H. J.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>J. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>RAANet: a residual ASPP with attention framework for semantic segmentation of high-resolution remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <fpage>3109</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2020.3021098</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y. H.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>X. H.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>M. H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X. B.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>b). <article-title>RoadNet: learning to comprehensively analyze road networks in complex urban scenes from high-resolution remotely sensed images</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>57</volume>, <fpage>2043</fpage>&#x2013;<lpage>2056</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2018.2870871</pub-id>
</citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Long</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Shelhamer</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fully convolutional networks for semantic segmentation</article-title>. <source>IEEE T. Pattern. Anal.</source> <volume>39</volume>, <fpage>640</fpage>&#x2013;<lpage>651</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2016.2572683</pub-id>
</citation>
</ref>
<ref id="B77">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lou</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Loew</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>DC-UNet: rethinking the U-net architecture with dual channel efficient CNN for medical image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the Medical Imaging Kunming</conf-name>, <conf-loc>China</conf-loc>, Vol. <volume>11596</volume>. <fpage>758</fpage>&#x2013;<lpage>768</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/12.2582338</pub-id>
</citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>X. D.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y. H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A survey of semantic construction and application of satellite remote sensing images and data</article-title>. <source>Organ. End User Comput.</source> <volume>33</volume>, <fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.4018/JOEUC.20211101.oa6</pub-id>
</citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname> <given-names>H. F.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>C. C.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>L. N.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>L. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>High-resolution aerial images semantic segmentation using deep fully convolutional network with channel attention mechanism</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>12</volume>, <fpage>3492</fpage>&#x2013;<lpage>3507</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2019.2930724</pub-id>
</citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname> <given-names>Y. Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. N.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>X. K.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>Z. Y.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>Z. X.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Pixel representation augmented through cross-attention for high-resolution remote sensing imagery segmentation</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>5415</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14215415</pub-id>
</citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>A. L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>Y. F.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Factseg: foreground activation-driven small object semantic segmentation in large-scale remote sensing imagery</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2021.3097148</pub-id>
</citation>
</ref>
<ref id="B82">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>X. P.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X. K.</given-names>
</name>
<name>
<surname>Pun</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>MSFNET: multi-stage fusion network for semantic segmentation of fine-resolution remote sensing data</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS 2022</conf-name>. <conf-loc>Kuala Lumpur, Malaysia</conf-loc>, <publisher-name>IEEE</publisher-name>, <fpage>2833</fpage>&#x2013;<lpage>2836</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS46834.2022.9883789</pub-id>
</citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>J. B.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>W. J.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>X. H.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Deep-separation guided progressive reconstruction network for semantic segmentation of remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>5510</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14215510</pub-id>
</citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>L. F.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>H. Y.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>Y. T.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Y. P.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>STN: saliency-guided Transformer network for point-wise semantic segmentation of urban scenes</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2022.3190558</pub-id>
</citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marmanis</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Schindler</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Wegner</surname> <given-names>J. D.</given-names>
</name>
<name>
<surname>Galliani</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Datcu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Stilla</surname> <given-names>U.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Classification with an edge: improving semantic image segmentation with boundary detection</article-title>. <source>ISPRS J. Photogramm.</source> <volume>135</volume>, <fpage>158</fpage>&#x2013;<lpage>172</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2017.11.009</pub-id>
</citation>
</ref>
<ref id="B86">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Marmanis</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Wegner</surname> <given-names>J. D.</given-names>
</name>
<name>
<surname>Galliani</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Schindler</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Datcu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Stilla</surname> <given-names>U.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Semantic segmentation of aerial images with an ensemble of CNSS</article-title>,&#x201d; in <conf-name>ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences</conf-name>, Vol. <volume>3</volume>. <fpage>473</fpage>&#x2013;<lpage>480</lpage>.</citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marsocci</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Scardapane</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Komodakis</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MARE: self-supervised multi-attention REsu-net for semantic segmentation in remote sensing</article-title>. <source>Remote. Sens.</source> <volume>13</volume> (<issue>16</issue>), <fpage>3275.8</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13163275</pub-id>
</citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maxwell</surname> <given-names>A. E.</given-names>
</name>
<name>
<surname>Bester</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Guill&#xe9;n</surname> <given-names>L. A.</given-names>
</name>
<name>
<surname>Ramezan</surname> <given-names>C. A.</given-names>
</name>
<name>
<surname>Carpinello</surname> <given-names>D. J.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>Y. T.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Semantic segmentation deep learning for extracting surface mine extents from historic topographic maps</article-title>. <source>Remote. Sens.</source> <volume>12</volume>, <elocation-id>4145</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs12244145</pub-id>
</citation>
</ref>
<ref id="B89">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meng</surname> <given-names>X. L.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Y. C.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L. B.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Class-guided Swin Transformer for semantic segmentation of remote sensing imagery</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2022.3215200</pub-id>
</citation>
</ref>
<ref id="B90">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mi</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Superpixel-enhanced deep neural forest for remote sensing image semantic segmentation</article-title>. <source>ISPRS J. Photogramm.</source> <volume>159</volume>, <fpage>140</fpage>&#x2013;<lpage>152</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.08.015</pub-id>
</citation>
</ref>
<ref id="B91">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Review the state-of-the-art technologies of semantic segmentation based on deep learning</article-title>. <source>Neurocomputing</source> <volume>493</volume>, <fpage>626</fpage>&#x2013;<lpage>646</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.NEUCOM.2022.01.005</pub-id>
</citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mou</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Hua</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>X. X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Relation matters: relational context-aware fully convolutional network for semantic segmentation of high-resolution aerial images</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>58</volume> (<issue>11</issue>), <fpage>7557</fpage>&#x2013;<lpage>7569</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.2979552</pub-id>
</citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Noble</surname> <given-names>P. J.</given-names>
</name>
<name>
<surname>Seitz</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>S. S.</given-names>
</name>
<name>
<surname>Manoylov</surname> <given-names>K. M.</given-names>
</name>
<name>
<surname>Chandra</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Characterization of algal community composition and structure from the nearshore environment, Lake Tahoe</article-title>. <source>Front. Ecol. Evol.</source> <volume>10</volume>, <elocation-id>1053499</elocation-id>. doi: <pub-id pub-id-type="doi">10.3389/fevo.2022.1053499</pub-id>
</citation>
</ref>
<ref id="B94">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nowozin</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lampert</surname> <given-names>C. H.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Structured learning and prediction in computer vision</article-title>. <source>Found. Trends. Comput.</source> <volume>6</volume>, <fpage>185</fpage>&#x2013;<lpage>365</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1561/0600000033</pub-id>
</citation>
</ref>
<ref id="B95">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>&#xd6;zden</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Polat</surname> <given-names>E.</given-names>
</name>
</person-group> (<year>2005</year>). &#x201c;<article-title>Image segmentation using color and texture features</article-title>,&#x201d; in <conf-name>Proceedings of the EUSIPCO 2005</conf-name>. <conf-loc>Antalya, Turkey</conf-loc>, <publisher-name>IEEE</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Conditional generative adversarial network-based training sample set improvement model for the semantic segmentation of high-resolution remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>59</volume>, <fpage>7854</fpage>&#x2013;<lpage>7870</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.3033816</pub-id>
</citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Panboonyuen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Jitkajornwanich</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lawawirojwong</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Srestasathiern</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Vateekul</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Semantic segmentation on remotely sensed images using an enhanced global convolutional network with channel attention and domain specific transfer learning</article-title>. <source>Remote. Sens.</source> <volume>11</volume>, <elocation-id>83</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs11010083</pub-id>
</citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Panboonyuen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Jitkajornwanich</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Lawawirojwong</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Srestasathiern</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Vateekul</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Transformer-based decoder designs for semantic segmentation on remotely sensed images</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>5100</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13245100</pub-id>
</citation>
</ref>
<ref id="B99">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pastorino</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Moser</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Serpico</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Zerubia</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>Fully convolutional and feedforward networks for the semantic segmentation of remotely sensed images</article-title>. <source>Proc. ICIP 2022</source>, <fpage>1876</fpage>&#x2013;<lpage>1880</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICIP46576.2022.9897336</pub-id>
</citation>
</ref>
<ref id="B100">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Pastorino</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Moser</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Serpico</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Zerubia</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>b). &#x201c;<article-title>Semantic segmentation of sar images through fully convolutional networks and hierarchical probabilistic graphical models</article-title>,&#x201d; in <conf-name>Proceedings of the IGARSS 2022</conf-name>. <conf-loc>Kuala Lumpur, Malaysia</conf-loc>, <publisher-name>IEEE</publisher-name>, <fpage>1047</fpage>&#x2013;<lpage>1050</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IGARSS46834.2022.9883111</pub-id>
</citation>
</ref>
<ref id="B101">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Piramanayagam</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Saber</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Schwartzkopf</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Koehler</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Supervised classification of multisensor remotely sensed images using a deep learning framework</article-title>. <source>Remote. Sens.</source> <volume>10</volume>, <elocation-id>1429</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs10091429</pub-id>
</citation>
</ref>
<ref id="B102">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Priyanka</surname>
</name>
<name>
<surname>Sravya</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Shyam</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Nalini</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chintala</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>Fabio</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>DIResUNet: architecture for multiclass semantic segmentation of high resolution remote sensing imagery data</article-title>. <source>Appl. Intell.</source> <volume>52</volume>, <fpage>15462</fpage>&#x2013;<lpage>15482</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s10489-022-03310-z</pub-id>
</citation>
</ref>
<ref id="B103">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qin</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ni</surname> <given-names>L. H.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Deep learning based feature selection for remote sensing scene classification</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>12</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2015.2475299</pub-id>
</citation>
</ref>
<ref id="B104">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ronneberger</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Fischer</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Brox</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2015</year>). &#x201c;<article-title>U-Net: convolutional networks for biomedical image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the MICCAI 2015</conf-name>, <conf-loc>Munich, Germany</conf-loc>, <publisher-name>Springer</publisher-name>, Vol. <volume>2015</volume>. <fpage>234</fpage>&#x2013;<lpage>241</lpage>.</citation>
</ref>
<ref id="B105">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ru</surname> <given-names>L. X.</given-names>
</name>
<name>
<surname>Zhan</surname> <given-names>Y. B.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>B. S.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Learning affinity from attention: end-to-end weakly-supervised semantic segmentation with Transformers</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2022</conf-name>, <conf-loc>New Orleans, Louisiana</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>16846</fpage>&#x2013;<lpage>16855</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01634</pub-id>
</citation>
</ref>
<ref id="B106">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sebastian</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Rohith</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>L. S.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Significant full reference image segmentation evaluation: a survey in remote sensing field</article-title>. <source>Multim. Tools Appl.</source> <volume>81</volume>, <fpage>17959</fpage>&#x2013;<lpage>17987</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11042-022-12769-4</pub-id>
</citation>
</ref>
<ref id="B107">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Senthilkumaran</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Rajesh</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2009</year>). &#x201c;<article-title>Image segmentation-a survey of soft computing approaches</article-title>,&#x201d; in <conf-name>Proceedings of 2009 International Conference on Advances in Recent Technologies in Communication and Computing. IEEE</conf-name>, Vol. <volume>2009</volume>. <fpage>844</fpage>&#x2013;<lpage>846</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ARTCom.2009.219</pub-id>
</citation>
</ref>
<ref id="B108">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Seong</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Choi</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semantic segmentation of urban buildings using a high-resolution network (HRNet) with channel and spatial attention gates</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>3087</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13163087</pub-id>
</citation>
</ref>
<ref id="B109">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shamsolmoali</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Zareapoor</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Road segmentation for remote sensing images using adversarial spatial pyramid networks</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>59</volume>, <fpage>4673</fpage>&#x2013;<lpage>4688</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.3016086</pub-id>
</citation>
</ref>
<ref id="B110">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shao</surname> <given-names>Z. F.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>W. X.</given-names>
</name>
<name>
<surname>Deng</surname> <given-names>X. Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>M. D.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>Q. M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Multilabel remote sensing image retrieval based on fully convolutional network</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>13</volume>, <fpage>318</fpage>&#x2013;<lpage>328</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2019.2961634</pub-id>
</citation>
</ref>
<ref id="B111">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Song</surname> <given-names>M. X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>B. C.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>P. J.</given-names>
</name>
<name>
<surname>Shao</surname> <given-names>Z. H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>DMF-CL: dense multi-scale feature contrastive learning for semantic segmentation of remote-sensing images</article-title>,&#x201d; in <conf-name>Proceedings of Pattern Recognition and Computer Vision</conf-name>, <conf-loc>Shenzhen, China</conf-loc>. <publisher-name>Springer</publisher-name>, <fpage>152</fpage>&#x2013;<lpage>164</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-031-18916-6_13</pub-id>
</citation>
</ref>
<ref id="B112">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Su</surname> <given-names>Y. C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>T. J.</given-names>
</name>
<name>
<surname>Liuy</surname> <given-names>K. H.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Multi-scale wavelet frequency channel attention for remote sensing image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of IVMSP 2022</conf-name>, <conf-loc>Nafplio, Greece</conf-loc>, <publisher-name>IEEE</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/IVMSP54334.2022.9816247</pub-id>
</citation>
</ref>
<ref id="B113">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Subudhi</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Narayan</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Biswal</surname> <given-names>P. K.</given-names>
</name>
<name>
<surname>Dell'acqua</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A survey on superpixel segmentation as a preprocessing step in hyperspectral image analysis</article-title>. <source>IEEE J-STARS</source> <volume>14</volume>, <fpage>5015</fpage>&#x2013;<lpage>5035</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2021.3076005</pub-id>
</citation>
</ref>
<ref id="B114">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>S. T.</given-names>
</name>
<name>
<surname>Mu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L. Z.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>X. L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y. W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semantic segmentation for buildings of large intra-class variation in remote sensing images with O-GAN</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>475</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13030475</pub-id>
</citation>
</ref>
<ref id="B115">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Mayer</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>BAS4Net: boundary-aware semi-supervised semantic segmentation network for very high resolution remote sensing images</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>13</volume>, <fpage>5398</fpage>&#x2013;<lpage>5413</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JSTARS.2020.3021098</pub-id>
</citation>
</ref>
<ref id="B116">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Problems of encoder-decoder frameworks for high-resolution remote sensing image segmentation: structural stereotype and insufficient learning</article-title>. <source>Neurocomputing</source> <volume>330</volume>, <fpage>297</fpage>&#x2013;<lpage>304</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.neucom.2018.11.051</pub-id>
</citation>
</ref>
<ref id="B117">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname> <given-names>Z. Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>W. P.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Multi-resolution Transformer network for building and road segmentation of remote sensing image</article-title>. <source>ISPRS Int. J. Geo Inf.</source> <volume>11</volume>, <elocation-id>165</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/ijgi11030165</pub-id>
</citation>
</ref>
<ref id="B118">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tasar</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Tarabalka</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Alliez</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Incremental learning for semantic segmentation of large-scale remote sensing data</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Obs. Remote. Sens.</source> <volume>12</volume>, <fpage>3524</fpage>&#x2013;<lpage>3537</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/jstars.2019.2925416</pub-id>
</citation>
</ref>
<ref id="B119">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semantic segmentation of remote sensing image based on GAN and FCN network model</article-title>. <source>Sci. Program.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>11</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2021/9491376</pub-id>
</citation>
</ref>
<ref id="B120">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tong</surname> <given-names>X. Y.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>G. S.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>S.</given-names>
</name>
<name>
<surname>You</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Land-cover classification with high-resolution remote sensing images using transferable deep models</article-title>. <source>Remote Sens. Environ.</source> <volume>237</volume>, <fpage>111322</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.CAGEO.2021.104969</pub-id>
</citation>
</ref>
<ref id="B121">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsagkatakis</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Aidini</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Fotiadou</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Giannopoulos</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Pentari</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Tsakalides</surname> <given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Survey of deep-learning approaches for remote sensing observation enhancement</article-title>. <source>Sensors</source> <volume>19</volume>, <elocation-id>3929</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/s19183929</pub-id>
</citation>
</ref>
<ref id="B122">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shazeer</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Parmar</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Uszkoreit</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Aidan</surname> <given-names>N.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). &#x201c;<article-title>Attention is all you need</article-title>,&#x201d; in <conf-name>Proceedings of the NIPS 2017</conf-name>. <conf-loc>Long Beach, CA, USA</conf-loc>, Vol. <volume>2017</volume>. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</citation>
</ref>
<ref id="B123">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Venugopal</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Automatic semantic segmentation with DeepLab dilated learning network for change detection in remote sensing images</article-title>. <source>Neural Process. Lett.</source> <volume>51</volume>, <fpage>2355</fpage>&#x2013;<lpage>2377</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11063-019-10174-x</pub-id>
</citation>
</ref>
<ref id="B124">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X. Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T. X.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Z. Y.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J. Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>CCTNet: coupled CNN and Transformer network for crop segmentation of remote sensing images</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>1956</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14091956</pub-id>
</citation>
</ref>
<ref id="B125">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>H. B.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>S. Q.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Remote sensing image segmentation of ground objects based on improved Deeplabv3+</article-title>,&#x201d; in <conf-name>Proceedings of the ICIT 2022</conf-name>, <conf-loc>Shanghai, China</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICIT48603.2022.10002795</pub-id>
</citation>
</ref>
<ref id="B126">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>FPB-UNet++: semantic segmentation for remote sensing images of reservoir area <italic>via</italic> improved UNet++ with FPN</article-title>,&#x201d; in <conf-name>Proceedings of the ICIAI 2022</conf-name>, <conf-loc>Guangzhou, China</conf-loc>. <publisher-name>ACM</publisher-name>, <fpage>100</fpage>&#x2013;<lpage>104</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/12.2582338</pub-id>
</citation>
</ref>
<ref id="B127">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Y. H.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>D. F.</given-names>
</name>
<name>
<surname>Sha</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Mask DeepLab: end-to-end image segmentation for change detection in high-resolution remote sensing images</article-title>. <source>Int. J. Appl. Earth Obs. Geoinform.</source> <volume>104</volume>, <elocation-id>102582</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.jag.2021.102582</pub-id>
</citation>
</ref>
<ref id="B128">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>L. B.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>X. L.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A novel Transformer based semantic segmentation scheme for fine-resolution remote sensing images</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2022.3143368</pub-id>
</citation>
</ref>
<ref id="B129">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>X. L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Transformer meets convolution: a bilateral awareness network for semantic segmentation of very fine resolution urban scene images</article-title>. <source>Remote. Sens.</source> <volume>13</volume> (<issue>16</issue>), <fpage>3065.37</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13163065</pub-id>
</citation>
</ref>
<ref id="B130">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Z. M.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. S.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L. M.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>F. J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X. Y.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>Semantic segmentation of high-resolution remote sensing images based on a class feature attention mechanism fused with Deeplabv3+</article-title>. <source>Comput. Geosci.</source> <volume>158</volume>, <elocation-id>104969</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/J.CAGEO.2021.104969</pub-id>
</citation>
</ref>
<ref id="B131">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Zhai</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Multi-scale network for remote sensing segmentation</article-title>. <source>IET Image Process.</source> <volume>16</volume>, <fpage>1742</fpage>&#x2013;<lpage>1751</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1049/ipr2.12444</pub-id>
</citation>
</ref>
<ref id="B132">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Gross</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C. L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>B. H.</given-names>
</name>
</person-group> (<year>2022</year>b). <article-title>Fused adaptive receptive field mechanism and dynamic multiscale dilated convolution for side-scan sonar image segmentation</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>17</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2022.3201248</pub-id>
</citation>
</ref>
<ref id="B133">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>A. L.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>X. Y.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>Y. F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>LoveDA: a remote sensing land-cover dataset for domain adaptive semantic segmentation</article-title>. <source>arXiv</source> <volume>2110</volume>, <elocation-id>8733</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2110.08733</pub-id>
</citation>
</ref>
<ref id="B134">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Simultaneous road surface and centerline extraction from Large-scale remote sensing images using CNN-based segmentation and tracing</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>99</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2020.2991733</pub-id>
</citation>
</ref>
<ref id="B135">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weiss</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Jacob</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Duveiller</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Remote sensing for agricultural applications: a meta-review</article-title>. <source>Remote Sens. Environ.</source> <volume>236</volume>, <elocation-id>111402</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.rse.2019.111402</pub-id>
</citation>
</ref>
<ref id="B136">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weng</surname> <given-names>L. G.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Y. M.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y. H.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Y. Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Water areas segmentation from remote sensing images using a separable residual segnet network</article-title>. <source>ISPRS Int. J. Geo Inf.</source> <volume>9</volume>, <elocation-id>256</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/ijgi9040256</pub-id>
</citation>
</ref>
<ref id="B137">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Woo</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Park</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>J. Y.</given-names>
</name>
<name>
<surname>Kweon</surname> <given-names>I. S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>CBAM: convolutional block attention module</article-title>. <source>Proceedings of the ECCV</source> (<publisher-loc>Munich, Germany</publisher-loc>), <fpage>3</fpage>&#x2013;<lpage>19</lpage>. </citation>
</ref>
<ref id="B138">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wurm</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Stark</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>X. X.</given-names>
</name>
<name>
<surname>Weigand</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Hannes.</surname> <given-names>T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Semantic segmentation of slums in satellite images using transfer learning on fully convolutional neural networks</article-title>. <source>ISPRS J. Photogrammetry Remote Sens.</source> <volume>150</volume>, <fpage>59</fpage>&#x2013;<lpage>69</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2019.02.006</pub-id>
</citation>
</ref>
<ref id="B139">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>X. M.</given-names>
</name>
<name>
<surname>Ai</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>F. L.</given-names>
</name>
<name>
<surname>Wen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>X. M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Feature-selection high-resolution network with hypersphere embedding for semantic segmentation of VHR remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2022.3183144</pub-id>
</citation>
</ref>
<ref id="B140">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>W. C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T. X.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Z. F.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J. Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Efficient Transformer for remote sensing image segmentation</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>3585</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13183585</pub-id>
</citation>
</ref>
<ref id="B141">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Transformer-based model with dynamic attention pyramid head for semantic segmentation of VHR remote sensing imagery</article-title>. <source>Entropy</source> <volume>24</volume>, <elocation-id>1619</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/e24111619</pub-id>
</citation>
</ref>
<ref id="B142">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Sparse and complete latent organization for geospatial semantic segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2022</conf-name>, <publisher-name>IEEE</publisher-name>, <fpage>1809</fpage>&#x2013;<lpage>1818</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR52688.2022.00185</pub-id>
</citation>
</ref>
<ref id="B143">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Vosselman</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>G. S.</given-names>
</name>
<name>
<surname>Yilmaz</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>M. Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>UAVid: a semantic segmentation dataset for UAV imagery</article-title>. <source>ISPRS J. Photogramm.</source> <volume>165</volume>, <fpage>108</fpage>&#x2013;<lpage>119</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.05.009</pub-id>
</citation>
</ref>
<ref id="B144">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yue</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>TreeUNet: adaptive tree convolutional neural networks for subdecimeter aerial image segmentation</article-title>. <source>ISPRS J. Photogramm.</source> <volume>156</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2019.07.007</pub-id>
</citation>
</ref>
<ref id="B145">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>X. Y.</given-names>
</name>
<name>
<surname>Duan</surname> <given-names>Q. Y.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>X. B.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>DHT: deformable hybrid Transformer for aerial image segmentation</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2022.3222916</pub-id>
</citation>
</ref>
<ref id="B146">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>W. S.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C. J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Transformer and CNN hybrid deep neural network for semantic segmentation of very-high-resolution remote sensing imagery</article-title>. <source>IEEE Trans. Geosci. Remote Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>20</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2022.3144894</pub-id>
</citation>
</ref>
<ref id="B147">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>X. Y.</given-names>
</name>
<name>
<surname>Leng</surname> <given-names>C. C.</given-names>
</name>
<name>
<surname>Hong</surname> <given-names>Y. M.</given-names>
</name>
<name>
<surname>Pei</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Basu</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multimodal remote sensing image registration methods and advancements: a survey</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>5128</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13245128</pub-id>
</citation>
</ref>
<ref id="B148">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Cattani</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liu.</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Diffusion-based image inpainting forensics <italic>via</italic> weighted least squares filtering enhancement</article-title>. <source>Multim. Tools Appl.</source> <volume>80</volume>, <fpage>30725</fpage>&#x2013;<lpage>30739</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11042-021-10623-7</pub-id>
</citation>
</ref>
<ref id="B149">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J. T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A survey algorithm research of scene parsing based on deepLearning</article-title>. <source>J. Com. Res. Develop.</source> <volume>57</volume>, <fpage>859</fpage>&#x2013;<lpage>875</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.7544/issn1000-1239.2020.20190513</pub-id>
</citation>
</ref>
<ref id="B150">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Object-specific optimization of hierarchical multiscale segmentations for high-spatial resolution remote sensing images - science direct</article-title>. <source>ISPRS J. Photogramm.</source> <volume>159</volume>, <fpage>308</fpage>&#x2013;<lpage>321</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2019.11.009</pub-id>
</citation>
</ref>
<ref id="B151">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xiong</surname> <given-names>N. N.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2022</year>b). <article-title>An end-to-end deep learning model for robust smooth filtering identification</article-title>. <source>Future Gener. Comp. Sy</source> <volume>127</volume>, <fpage>263</fpage>&#x2013;<lpage>275</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.future.2021.09.004</pub-id>
</citation>
</ref>
<ref id="B152">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y. R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Memory-augmented Transformer for remote sensing image semantic segmentation</article-title>. <source>Remote. Sens.</source> <volume>13</volume>, <elocation-id>4518</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs13224518</pub-id>
</citation>
</ref>
<ref id="B153">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J. H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y. W.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semantic segmentation with attention mechanism for remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2021.3085889</pub-id>
</citation>
</ref>
<ref id="B154">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>D. P.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C. X.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>Z. W.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>F. Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Semantic segmentation of remote sensing image based on regional self-attention mechanism</article-title>. <source>IEEE Geosci. Remote. Sens. Lett.</source> <volume>19</volume>, <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/LGRS.2021.3071624</pub-id>
</citation>
</ref>
<ref id="B155">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>J. Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>B. Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>J. Y.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>R.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Multi-source collaborative enhanced for remote sensing images semantic segmentation</article-title>. <source>Neurocomputing</source> <volume>493</volume>, <fpage>76</fpage>&#x2013;<lpage>90</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.neucom.2022.04.045</pub-id>
</citation>
</ref>
<ref id="B156">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Rethinking semantic segmentation from a sequence-to-sequence perspective with Transformers</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2021</conf-name>, <conf-loc>Nashville, TN, USA</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>6881</fpage>&#x2013;<lpage>6890</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.00681</pub-id>
</citation>
</ref>
<ref id="B157">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>Y. L.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>M. Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>X. J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Semi-supervised adversarial semantic segmentation network using Transformer and multiscale convolution for high-resolution remote sensing imagery</article-title>. <source>Remote. Sens.</source> <volume>14</volume>, <elocation-id>1786</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/rs14081786</pub-id>
</citation>
</ref>
<ref id="B158">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhong</surname> <given-names>Y. F.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>A. L.</given-names>
</name>
</person-group> (<year>2020</year>a). &#x201c;<article-title>Foreground-aware relation network for geospatial object segmentation in high spatial resolution remote sensing imagery</article-title>,&#x201d; in <conf-name>Proceedings of the CVPR 2020</conf-name>, <conf-loc>Washington, Seattle</conf-loc>. <publisher-name>IEEE</publisher-name>, <fpage>4096</fpage>&#x2013;<lpage>4105</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00415</pub-id>
</citation>
</ref>
<ref id="B159">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>X. W.</given-names>
</name>
<name>
<surname>Huan</surname> <given-names>L. X.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>J. Y.</given-names>
</name>
</person-group> (<year>2020</year>b). <article-title>Parsing very high resolution urban scene images by learning deep ConvNets with edge-aware loss</article-title>. <source>ISPRS J. Photogramm.</source> <volume>170</volume>, <fpage>15</fpage>&#x2013;<lpage>28</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.09.019</pub-id>
</citation>
</ref>
<ref id="B160">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname> <given-names>H. F.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>H. M.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>D. N.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z. H.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>R. S.</given-names>
</name>
</person-group> (<year>2022</year>a). <article-title>Lake water body extraction of optical remote sensing images based on semantic segmentation</article-title>. <source>Appl. Intell.</source> <volume>52</volume>, <fpage>1</fpage>&#x2013;<lpage>16</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s10489-022-03345-2</pub-id>
</citation>
</ref>
<ref id="B161">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname> <given-names>H. F.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>H. M.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>R. S.</given-names>
</name>
</person-group> (<year>2022</year>b). <article-title>NT-Net: a semantic segmentation network for extracting lake water bodies from optical remote sensing images based on Transformer</article-title>. <source>IEEE Trans. Geosci. Remote. Sens.</source> <volume>60</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TGRS.2022.3197402</pub-id>
</citation>
</ref>
<ref id="B162">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Z. W.</given-names>
</name>
<name>
<surname>Siddiquee</surname> <given-names>M. M.</given-names>
</name>
<name>
<surname>Tajbakhsh</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>J. M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Unet++: a nested u-net architecture for medical image segmentation</article-title>,&#x201d; in <conf-name>Proceedings of the DLMIA 2018</conf-name>, <conf-loc>Granada, Spain</conf-loc>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-00889-5_1</pub-id>
</citation>
</ref>
<ref id="B163">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname> <given-names>X. X.</given-names>
</name>
<name>
<surname>Devis</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Mou</surname> <given-names>L. C.</given-names>
</name>
<name>
<surname>Xia</surname> <given-names>G. S.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L. P.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>F.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Deep learning in remote sensing: a comprehensive review and list of resources</article-title>. <source>IEEE Geosc. Rem. Sen. M.</source> <volume>5</volume>, <fpage>8</fpage>&#x2013;<lpage>36</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/MGRS.2017.2762307</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>