<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Ecol. Evol.</journal-id>
<journal-title>Frontiers in Ecology and Evolution</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Ecol. Evol.</abbrev-journal-title>
<issn pub-type="epub">2296-701X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fevo.2022.1083801</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Ecology and Evolution</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>GSC-MIM: Global semantic integrated self-distilled complementary masked image model for remote sensing images scene classification</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Xuying</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1835202/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhang</surname> <given-names>Yunsheng</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Zhaoyang</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1843077/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Luo</surname> <given-names>Qinyao</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Yang</surname> <given-names>Jingfan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1843461/overview"/>
</contrib>
</contrib-group>
<aff><institution>School of Geosciences and Info-Physics, Central South University, Changsha</institution>, <addr-line>Hunan</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yansheng Li, Wuhan University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Chengli Peng, Wuhan University, China; Qiqi Zhu, China University of Geosciences Wuhan, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Yunsheng Zhang &#x02709;<email>zhangys&#x00040;csu.edu.cn</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Environmental Informatics and Remote Sensing, a section of the journal Frontiers in Ecology and Evolution</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>12</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>1083801</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>06</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Wang, Zhang, Zhang, Luo and Yang.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Wang, Zhang, Zhang, Luo and Yang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Masked image modeling (MIM) is a learning method in which the unmasked components of the input are utilized to learn and predict the masked signal, enabling learning from large amounts of unannotated data. However, due to the scale diversity and complexity of features in remote sensing images (RSIs), existing MIMs face two challenges in the RSI scene classification task: (1) If the critical local patches of small-scale objects are randomly masked out, the model will be unable to learn its representation. (2) The reconstruction of MIM relies on the visible local contextual information surrounding the masked regions and overemphasizing this local information will potentially lead the model to disregard the global semantic information of the input RSI. Regarding the above considerations, we proposed a global semantic integrated self-distilled complementary masked image model (GSC-MIM) for RSI scene classification. To prevent information loss, we proposed an information-preserved complementary masking strategy (IPC-Masking), which generates two complementary masked views for the same image to resolve the problem of masking critical areas of small-scale objects. To incorporate global information into the MIM pre-training process, we proposed the global semantic distillation strategy (GSD). Specifically, we introduced an auxiliary network pipeline to extract the global semantic information from the full input RSI and transfer the knowledge to the MIM by self-distillation. The proposed GSC-MIM is validated on three publicly available datasets of AID, NWPU-RESISC45, and UC-Merced Land Use, and the results show that the proposed method&#x00027;s Top-1 accuracy surpasses the baseline approaches in three datasets by up to 4.01, 3.87, and 5.26%, respectively.</p>
</abstract>
<kwd-group>
<kwd>self-supervised learning (SSL)</kwd>
<kwd>masked image modeling (MIM)</kwd>
<kwd>self-distillation</kwd>
<kwd>remote sensing images (RSIs)</kwd>
<kwd>scene classification</kwd>
</kwd-group>
<contract-num rid="cn001">41871364</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="8"/>
<table-count count="3"/>
<equation-count count="5"/>
<ref-count count="31"/>
<page-count count="11"/>
<word-count count="5609"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The propose and development of self-supervised learning (SSL) has freed conventional supervised learning from the heavy reliance on large-scale high-quality annotated data, making it a reality to obtain well-performed interpretation models label-free (Wang Y. et al., <xref ref-type="bibr" rid="B24">2022</xref>). Compared to supervised learning (Lu et al., <xref ref-type="bibr" rid="B20">2022</xref>; Zhu et al., <xref ref-type="bibr" rid="B30">2022a</xref>,<xref ref-type="bibr" rid="B31">b</xref>) which learn the projection data to label, SSL aims to find correlations between data and drive models to learn features by manually constructing pre-tasks to constitute labels, yielding competitive results in RSI scene classification (Tao et al., <xref ref-type="bibr" rid="B22">2020</xref>; Wang X. et al., <xref ref-type="bibr" rid="B23">2022</xref>), object detection (Ding et al., <xref ref-type="bibr" rid="B9">2021</xref>), and semantic segmentation (Li et al., <xref ref-type="bibr" rid="B14">2022</xref>) tasks.</p>
<p>The dominant self-supervised learning approaches can be divided into two main categories: contrastive and generative (Liu X. et al., <xref ref-type="bibr" rid="B17">2021</xref>). The core of the contrastive learning method (Chen et al., <xref ref-type="bibr" rid="B5">2020</xref>; He et al., <xref ref-type="bibr" rid="B12">2020</xref>) is to pull closer different views of the same image (positive samples) and push apart views of different images (negative samples), after which potential invariant features in the images are prompted to be learned by the model through the determination of the positive and negative samples. Currently, the contrastive self-supervised learning-based approaches obtain promising results in various RSI interpretation tasks (Wang Y. et al., <xref ref-type="bibr" rid="B24">2022</xref>). Some researchers (Li et al., <xref ref-type="bibr" rid="B15">2021a</xref>; Akiva et al., <xref ref-type="bibr" rid="B1">2022</xref>) follow the basic idea of contrastive learning to construct positive and negative samples to accomplish label-free training of RSI interpretation models. In addition, since RSIs contain abundant information, some researchers include geographic features (Ayush et al., <xref ref-type="bibr" rid="B2">2021</xref>; Li et al., <xref ref-type="bibr" rid="B16">2021b</xref>), time series features (Manas et al., <xref ref-type="bibr" rid="B21">2021</xref>), and audio features (Heidler et al., <xref ref-type="bibr" rid="B13">2021</xref>) of RSIs in the contrastive learning process to encourage models to learn the invariance of RS-specific features. However, contrastive SSL has its own inherent limitations for RSI interpretation tasks. Specifically, RSI contains a complex variety of land objects and their corresponding labels are at the scene level in scene classification task. If the RSI containing the same type of object is selected as a negative sample, it will negatively influence the feature learning of such object in the pushing apart process, and vice versa (Zhang et al., <xref ref-type="bibr" rid="B28">2022</xref>). <xref ref-type="fig" rid="F1">Figure 1</xref> shows a brief introduction of the contrastive learning methods and masked image modeling method. Take the negative pair in <xref ref-type="fig" rid="F1">Figure 1</xref> as example, if we push this negative pair apart, it will inevitably push the feature of orange building apart, which is clearly not ideal. Therefore, we argue that the conventional contrastive SSL is suboptimal for the RSI interpretation tasks.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>A brief introduction of the contrastive learning methods and masked image modeling method.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0001.tif"/>
</fig>
<p>The fundamental concept of generative self-supervised approaches is to train models to reconstruct and restore manually corrupted images, which essentially enables the models to learn to present well representations of the original image (Liu X. et al., <xref ref-type="bibr" rid="B17">2021</xref>). Mask image modeling (MIM) as a classical generative self-supervised pre-training paradigm, which leverages a large amount of data to drive the training of vision transformer (ViT) (Dosovitskiy et al., <xref ref-type="bibr" rid="B10">2020</xref>), has achieved competitive results in many image interpretation tasks (Bao et al., <xref ref-type="bibr" rid="B3">2021</xref>; He et al., <xref ref-type="bibr" rid="B11">2022</xref>; Xie et al., <xref ref-type="bibr" rid="B26">2022</xref>). The core idea of MIM is to crop the input image into several local semantically meaningful visual tokens and randomly mask some of them, then train the vision transformer to reconstruct the masked parts based on the adjacent visible parts. Based on the above considerations, we assume that MIM-based generative self-supervised learning is more suitable for training remote sensing image interpretation models as it can flexibly acquire various features embedded inside remote sensing images and obtain well-formed representations of RSI without relying on data augmentation methods or negative samples.</p>
<p>However, when MIM is used for remote sensing image scene classification tasks, it usually encounters the following two issues:</p>
<list list-type="bullet">
<list-item><p>Since remote sensing images contain complex multi-scale land objects, if the critical local patch token containing a certain type of small-scale object is randomly masked out, the model will not be able to obtain its features, resulting in irreversible information loss.</p></list-item>
<list-item><p>The reconstruction mechanism of MIM is essentially a model inference based on local contextual information near the masked region, which causes the model to ignore the global semantic information of the whole input image, which is critical for RSI scene classification.</p></list-item>
</list>
<p>Considering the above limitations, we proposed a global semantic integrated self-distilled complementary masked image model (GSC-MIM) for RSI scene classification, which consist of two strategies: information preserved complementary masking strategy (IPC-Masking) and global semantic distillation strategy (GSD). Existing random masking strategies tend to lose dense and small-scale features in the RSIs. Two versions of the complementary visible-mask patch are generated from the same remote sensing image to maximize the retention of small-scale RS object information when reconstructing them with MIM. The reconstruction of masked patches in MIM relies on local adjacent visible patches and does not emphasize the global scene semantic information of the whole input image. To incorporate global information into the MIM pre-training process, we propose a global semantic distillation strategy (GSD). As an emerging method of SSL, self-distillation (Caron et al., <xref ref-type="bibr" rid="B4">2021</xref>; Cino et al., <xref ref-type="bibr" rid="B8">2022</xref>) learning leverages the network&#x00027;s past self to distillate needed knowledge to the present self to achieve discriminative self-learning. In view of this, we introduce an auxiliary network pipeline to extract global semantic information from the full input image and transfer the global knowledge to the MIM <italic>via</italic> self-distillation.</p>
<p>We use GSC-MIM to obtain the fundamental pre-trained model, and the features obtained by the model can be used for downstream scene classification tasks. We evaluated the performance of the model on three RSI scene classification public datasets, Aerial Image Dataset (AID) (Xia et al., <xref ref-type="bibr" rid="B25">2017</xref>), UC-Merced Land Use (UCM) (Yang and Newsam, <xref ref-type="bibr" rid="B27">2010</xref>), and NWPU-RESISC45 (NWPU45) (Cheng et al., <xref ref-type="bibr" rid="B7">2017</xref>). Experimental results show that our proposed GSC-MIM achieves better results compared to the classical contrastive self-supervised learning, generative self-supervised learning and self-distillation learning approaches. The main contributions of this paper are summarized as follows:</p>
<list list-type="order">
<list-item><p>We propose an information preserved complementary masking strategy (IPC-Masking), which aims to reduce the loss of small-scale features caused by MIM when used for RSI interpretation tasks. We verified that the information of the original input image could be maximally preserved by the simultaneous reconstruction of its two complementary masked-visible-region views.</p></list-item>
<list-item><p>We highlight the neglect of RSI global semantic information during the MIM training process. We propose a global semantic distillation strategy (GSD) to extract the global semantic information of the input images using additional network pipeline and distill the knowledge into the MIM network to compensate for the lack of global information.</p></list-item>
<list-item><p>Experiments on public datasets show that our proposed GSC-MIM achieves a maximum accuracy improvement of up to 5.26% on the RSI scene classification task under the equivalent conditions. Moreover, the network attention visualization results show that our model captures more highly detailed features of the land objects and locates the global scene information regions of the RSI more precisely.</p></list-item>
</list>
<p>The rest of this paper is organized as follows: Section 2 introduces the details of our proposed method. Section 3 discusses the experimental and visualization results organized on three public datasets and future work prospects. Section 4 concludes this paper.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2. Methodology</title>
<p>RSI has complex geographical features and a multi-scale spatial layout, if the critical patches containing certain types of features are randomly masked during the MIM training process, it will lead to information loss; Moreover, the reconstruction of local patch by MIM may lead the model to emphasize the local information and ignore the global information associated with the input image&#x00027;s category. Inspired by the above facts, the GSC-MIM is developed to preserve information of different scales&#x00027; land objects, and to integrate global meanings into the training process <italic>via</italic> self-distillation. The architecture of GSC-MIM is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The architecture consists of two identical structured ViT backbone networks, which we refer to as the teacher and student networks. The network includes two strategies, information-preserved complementary masking strategy (IPC-Masking) and global semantic distillation strategy (GSD), which will be described in the following sections.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The architecture of GSC-MIM.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0002.tif"/>
</fig>
<sec>
<title>2.1. Information-preserved complementary masking strategy (IPC-Masking)</title>
<p>The IPC-Masking is inspired by the conventional MIM&#x00027;s masking strategy. Firstly, for each input image <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, we crop it into <italic>N</italic> patches and feed into the linear projecting layer to obtain patch token sequence <inline-formula><mml:math id="M2"><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. After that, we randomly give each patch token a mask indicator <italic>m</italic> &#x02208; {0, 1}<sup><italic>N</italic></sup>, where <italic>N</italic> is the number of tokens. Specifically, wherever the mask indicator <italic>m</italic><sub><italic>i</italic></sub> is 1, the corresponding <italic>x</italic><sub><italic>i</italic></sub> is replaced by a mask <italic>e</italic><sub>[<italic>MASK</italic>]</sub>, which yields vision 1 of the masked image as:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>x</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>&#x0225C;</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>x</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02223;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>x</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:mi>e</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>MASK</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Vision 2 of the masked image is complementary to version 1, i.e., the visible-masked tokens of the two images are opposite. The illustration of IPC-Masking is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. In this work, we perform the IPC-Masking to the input image of the student network. Through the vision transformer backbone and decoder, the reconstructed patch tokens <inline-formula><mml:math id="M4"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> are generated for comparison.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The illustration of IPC-masking.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0003.tif"/>
</fig>
<p>In the teacher network, the two different augmented non-masked views of the same image serve as input to obtain their projections <inline-formula><mml:math id="M5"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>We then define the training objective of complementary MIM as:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>reconstructed</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x000B7;</mml:mo><mml:msup><mml:mrow><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[patch]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msup><mml:mo class="qopname">log</mml:mo><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[patch]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>We symmetrize the loss by averaging with the above CE term between two pairs of <inline-formula><mml:math id="M7"><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M8"><mml:msubsup><mml:mrow><mml:mi>I</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula>.</p>
</sec>
<sec>
<title>2.2. Global semantic distillation strategy (GSD)</title>
<p>As a branch of self-supervised learning, self-distillation dexterously transmits knowledge from the model&#x00027;s past itself and is cast as a discriminative objective. The GSD proposed in this paper utilized the framework of self-distillation to generate and integrate global semantic information into the student network.</p>
<p>We followed the self-distillation paradigm proposed in Caron et al. (<xref ref-type="bibr" rid="B4">2021</xref>). We adopt the same architecture for student and teacher networks, consisting of backbone <italic>f</italic> and projection head <italic>h</italic> : <italic>g</italic> &#x0003D; <italic>h</italic> &#x025E6; <italic>f</italic>. The parameters of the student network <italic>&#x003B8;</italic><sub><italic>s</italic></sub> are updated by back propagation according to the loss function, while the teacher&#x00027;s parameters <italic>&#x003B8;</italic><sub><italic>t</italic></sub> are exponentially moving averaged (EMA) of the updated <italic>&#x003B8;</italic><sub><italic>s</italic></sub>, which is described as follows:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02190;</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BB; following a cosine schedule from 0.996 to 1 during training.</p>
<p>For vanilla ViT (Dosovitskiy et al., <xref ref-type="bibr" rid="B10">2020</xref>), [<italic>CLS</italic>] token is a learnable embedding to the sequence of embedded patches whose state at the output of the Transformer encode and contains the predictive categorical distributions of the input image <italic>x</italic>. In GSC-MIM, we utilize [<italic>CLS</italic>] tokens as a proxy for global semantic information and to perform knowledge distillation. For a training set <inline-formula><mml:math id="M10"><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow></mml:math></inline-formula>, an image <inline-formula><mml:math id="M11"><mml:mi>x</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="-tex-caligraphic">I</mml:mi></mml:mrow></mml:math></inline-formula> is sampled uniformly and applied two random augmentations, yielding two distorted views <italic>t</italic><sub>1</sub> and <italic>t</italic><sub>2</sub>. We also applied complementary masking to the same image to get two corrupted views <italic>s</italic><sub>1</sub> and <italic>s</italic><sub>2</sub>. After feeding <italic>t</italic><sub>1</sub> and <italic>t</italic><sub>2</sub>, <italic>s</italic><sub>1</sub>, and <italic>s</italic><sub>2</sub> with learnable [<italic>CLS</italic>] tokens into the teacher and student network correspondingly, we get the global semantically meaningful tokens of <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M13"><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold-italic"><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:math></inline-formula>. Since the global semantics of the input image is modeled by the [CLS] token generated <italic>via</italic> the teacher network with unmasked input. The global context encoding ability of the target student masked image model (MIM) is improved by minimizing the cross-view cross-entropy between the teacher&#x00027;s global semantic meaningful [CLS] token and the student&#x00027;s [CLS] token, which we assumed to be global semantically deficient, formulated as:</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M14"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:msup><mml:mo class="qopname">log</mml:mo><mml:msubsup><mml:mrow><mml:mi>h</mml:mi></mml:mrow><mml:mrow><mml:mstyle mathvariant="bold"><mml:msub><mml:mrow><mml:mi>s</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mstyle></mml:mrow><mml:mrow><mml:mstyle class="mbox"><mml:mtext>[CLS]</mml:mtext></mml:mstyle></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>2.3. Loss function and architecture</title>
<p>With the information preserved complementary masking strategy and global semantic distillation, we designed the overall training objective as follows:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M15"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003BB;</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>reconstructed</mml:mtext></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003BB;<sub>1</sub>, &#x003BB;<sub>2</sub> represent the loss coefficients that balance the importance of [<italic>CLS</italic>] losses and reconstructed losses.</p>
<p>The backbone <italic>f</italic> adopted by GSC-MIM is vision transformer (ViT) encoder (Dosovitskiy et al., <xref ref-type="bibr" rid="B10">2020</xref>), shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. We also evaluated the different amounts of backbone parameters of ViT-S/16, ViT-B/16, and ViT-L/16. The projection head <italic>h</italic> of GSC-MIM is set as 3-layers MLP, which is proven optimal in both Caron et al. (<xref ref-type="bibr" rid="B4">2021</xref>) and Zhou et al. (<xref ref-type="bibr" rid="B29">2021</xref>). Moreover, to further borrow the capability of semantic abstraction obtained by self-distillation on [<italic>CLS</italic>] tokens, we share the projection head parameters for both patch tokens and [<italic>CLS</italic>] tokens. The scene classification head takes [<italic>CLS</italic>] token as input and is implemented by a MLP projection head with three hidden layers at pre-training time and by a single linear layer at fine-tuning time.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Vision transformer encoder.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0004.tif"/>
</fig>
</sec>
</sec>
<sec id="s3">
<title>3. Experimentation and results discussion</title>
<p>Three public remote sensing datasets are used for thorough tests to assess the efficacy of the proposed GSC-MIM approach, including parameter analysis, an experimental comparison of several algorithms, and visualization analysis.</p>
<sec>
<title>3.1. Experimental setup</title>
<p>(1) <bold>Datasets description</bold>. In this paper, the Aerial Image Dataset (AID) (Xia et al., <xref ref-type="bibr" rid="B25">2017</xref>), the NWPU-RESISC45 Dataset (NWPU45) (Cheng et al., <xref ref-type="bibr" rid="B7">2017</xref>), and the UC-Merced Land Use Dataset (UCM) (Yang and Newsam, <xref ref-type="bibr" rid="B27">2010</xref>) are selected to conduct the experiments. With high image size and spatial resolution diversity, these three datasets are challenging for RSI scene classification task. More detailed information about datasets is shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Datasets description.</p></caption>
<table frame="box" rules="all">
<thead><tr style="border-right: thin solid #000000;background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Datasets</bold></th>
<th valign="top" align="left"><bold>AID</bold></th>
<th valign="top" align="left"><bold>NWPU45</bold></th>
<th valign="top" align="center"><bold>UCM</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Classes</td>
<td valign="top" align="left">30</td>
<td valign="top" align="left">45</td>
<td valign="top" align="center">21</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Images per class</td>
<td valign="top" align="left">220&#x02013;420</td>
<td valign="top" align="left">700</td>
<td valign="top" align="center">100</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Total images</td>
<td valign="top" align="left">10,000</td>
<td valign="top" align="left">31,500</td>
<td valign="top" align="center">2,100</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Spatial resolution (m)</td>
<td valign="top" align="left">0.5&#x02013;8</td>
<td valign="top" align="left">0.2&#x02013;30</td>
<td valign="top" align="center">0.3</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Image size</td>
<td valign="top" align="left">600 &#x000D7; 600</td>
<td valign="top" align="left">256 &#x000D7; 256</td>
<td valign="top" align="center">256 &#x000D7; 256</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Data source</td>
<td valign="top" align="left">Google Earth</td>
<td valign="top" align="left">Google Earth</td>
<td valign="top" align="center">USGS</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">Published year</td>
<td valign="top" align="left">2017</td>
<td valign="top" align="left">2017</td>
<td valign="top" align="center">2010</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>(2) <bold>Hardware and software environment</bold>. All experiments are carried out on NVIDIA A100 Tensor Core GPU with a software environment of Python 3.7 with PyTorch performing on the Ubuntu system.</p>
<p>(3) <bold>Architecture parameter setup</bold>. For the backbone <italic>f</italic>, we inherit the ViT models&#x00027; parameters from Caron et al. (<xref ref-type="bibr" rid="B4">2021</xref>) and train the network without loading the pretrained weight. For ViTs, /16 denotes the patch size being 16. We also evaluate patch sizes 8 and 32 in the following experiment. For the projection head <italic>h</italic>, a 3&#x02212;<italic>layerMLP</italic> with <italic>l</italic>2&#x02212;<italic>normalized</italic> bottleneck is chosen to generate representations. Since we share <italic>h</italic> for both patch tokens and [<italic>CLS</italic>] tokens, they all get the output dimension of 8,192.</p>
<p>(4) <bold>Data processing and evaluation metrics</bold>. The proposed GSC-MIM uses multi-scale crop, random flip and random rotation as two different types of data augmentation methodologies to generate the distorted views of <italic>t</italic><sub>1</sub>, <italic>t</italic><sub>2</sub> and <italic>s</italic><sub>1</sub>, <italic>s</italic><sub>2</sub>. For the three datasets, 80% of each category is selected as the training set and the remaining 20% served as the testing set. After pretraining, we perform the linear evaluation using 1% of each category in the training set to elaborate on the effectiveness of the proposed GSC-MIM among the compared methods. The linear evaluation metric is the same as the works in Chen et al. (<xref ref-type="bibr" rid="B5">2020</xref>); Caron et al. (<xref ref-type="bibr" rid="B4">2021</xref>), and Zhou et al. (<xref ref-type="bibr" rid="B29">2021</xref>). Furthermore, the Top-1 Accuracy (Top-1 Acc) with visualized histogram is employed to illustrate the classification performance of the proposed method on the three datasets.</p>
<p>(5) <bold>Parameter optimization setup</bold>. We by default pre-train GSC-MIM on the above training datasets and set 16 as default patch size for IPC-Masking. The batch size is set to 64 for ViT-S and ViT-B, and 16 for ViT-L due to the limitation of GPU. For both teacher and student networks, the AdamW (Loshchilov and Hutter, <xref ref-type="bibr" rid="B19">2017</xref>) is employed as the optimizer for better convergence, the corresponding momentum is set to 0.996, and the weight decay is 0.004. For the student network, we employ random MIM, with prediction ratio <italic>r</italic> uniformly chosen from the range [0.1, 0.5] with a probability of 0.5 and set as 0 with a probability of 0.5. Moreover, we set &#x003BB;<sub>1</sub> and &#x003BB;<sub>2</sub> equal to 1 in <inline-formula><mml:math id="M16"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, i.e., we sum <inline-formula><mml:math id="M17"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>CLS</mml:mtext></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula><mml:math id="M18"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="-tex-caligraphic">L</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mtext>reconstructed</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> up without scaling. Finally, we train and evaluate the GSC-MIM for 300 epochs as default. During the first 10 epochs, the learning rate is linearly ramped up to its base value scaled with the entire batch size: <italic>lr</italic> &#x0003D; 5<italic>e</italic><sup>&#x02212;4</sup>&#x000D7;<italic>batch</italic>_<italic>size</italic>/256.</p>
</sec>
<sec>
<title>3.2. Results and discussion</title>
<sec>
<title>3.2.1. Performance results</title>
<p>Various state-of-the-art (SOTA) approaches are presented to compare with GSC-MIM on the three datasets to highlight the advantages of the proposed method. The results obtained from the linear evaluation are summarized in <xref ref-type="table" rid="T2">Table 2</xref>. The training and testing ratios for all listed methods remain the same for a fair comparison. Furthermore, all methods are divided into four categories to show the results effectively, including classical contrastive learning approaches (&#x02020;) (Chen et al., <xref ref-type="bibr" rid="B5">2020</xref>; He et al., <xref ref-type="bibr" rid="B12">2020</xref>), classical MIM-based methods (&#x025B3;) (He et al., <xref ref-type="bibr" rid="B11">2022</xref>; Xie et al., <xref ref-type="bibr" rid="B26">2022</xref>), self-distillation method (&#x02666;) (Caron et al., <xref ref-type="bibr" rid="B4">2021</xref>), and the proposed GSC-MIM (&#x022C6;). We also provide a bar chart for a more intuitive illustration, as shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. Several analysis can be drawn from these results.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Experimental results comparison with SOTA methods.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="border-right: thin solid #000000;background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method type</bold></th>
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Backbone</bold></th>
<th valign="top" align="center"><bold>Params. (M)</bold></th>
<th valign="top" align="center" colspan="3"><bold>Top-1 Acc (%)</bold></th>
</tr>
<tr style="border-right: thin solid #000000;background-color:#919498;color:#ffffff">
<th/>
<th/>
<th/>
<th/>
<th valign="top" align="center" colspan="3"><bold>Datasets</bold></th>
</tr>
<tr style="border-right: thin solid #000000;background-color:#919498;color:#ffffff">
<th/>
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>AID</bold></th>
<th valign="top" align="center"><bold>NWPU45</bold></th>
<th valign="top" align="center"><bold>UCM</bold></th>
</tr>
</thead>
<tbody>
<tr style="border-top: thin solid #000000;">
<td valign="top" align="left">&#x02020;</td>
<td valign="top" align="left">SimCLR</td>
<td valign="top" align="left">Resnet50</td>
<td valign="top" align="center">21</td>
<td valign="top" align="center">61.37</td>
<td valign="top" align="center">65.16</td>
<td valign="top" align="center">52.89</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">MoCo v3</td>
<td valign="top" align="left">ViT-small</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">63.67</td>
<td valign="top" align="center">67.11</td>
<td valign="top" align="center">50.38</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">&#x025B3;</td>
<td valign="top" align="left">MAE</td>
<td valign="top" align="left">ViT-base</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">65.33</td>
<td valign="top" align="center">71.56</td>
<td valign="top" align="center">50.63</td>
</tr>
<tr>
<td/>
<td valign="top" align="left">SimMIM</td>
<td valign="top" align="left">Swin-base</td>
<td valign="top" align="center">88</td>
<td valign="top" align="center"><bold>69.49</bold></td>
<td valign="top" align="center"><bold>73.01</bold></td>
<td valign="top" align="center">54.89</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">&#x02666;</td>
<td valign="top" align="left">DINO</td>
<td valign="top" align="left">ViT-small</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center">64.26</td>
<td valign="top" align="center">69.03</td>
<td valign="top" align="center">53.63</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left">&#x022C6;</td>
<td valign="top" align="left">GSC-MIM</td>
<td valign="top" align="left">ViT-small</td>
<td valign="top" align="center">22</td>
<td valign="top" align="center"><bold>65.38</bold></td>
<td valign="top" align="center"><bold>69.54</bold></td>
<td valign="top" align="center"><bold>54.14</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">GSC-MIM</td>
<td valign="top" align="left">ViT-base</td>
<td valign="top" align="center">86</td>
<td valign="top" align="center">68.98</td>
<td valign="top" align="center">72.47</td>
<td valign="top" align="center"><bold>55.89</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="left">GSC-MIM</td>
<td valign="top" align="left">ViT-large</td>
<td valign="top" align="center">304</td>
<td valign="top" align="center">72.03</td>
<td valign="top" align="center">76.04</td>
<td valign="top" align="center">None</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>&#x02020;Contrastive learning with ResNet/Vision Transformer backbone.</p>
<p>&#x025B3;MIM based methods.</p>
<p>&#x02666;Self-distillation based method.</p>
<p>&#x022C6;The proposed GSC-MIM.</p>
<p>The bold values represent the best-performing methods for similar amounts of network parameters.</p>
</table-wrap-foot>
</table-wrap>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Bar chart of experimental results comparison with SOTA methods.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0005.tif"/>
</fig>
<p>Firstly, the results of GSC-MIM (ViT-S) are significant compared to the conventional contrastive learning approaches with roughly the same amount of parameters such as SimCLR (ResNet50) (Chen et al., <xref ref-type="bibr" rid="B5">2020</xref>) and MoCo v3 (ViT-S) (Chen et al., <xref ref-type="bibr" rid="B6">2021</xref>). These observations suggest that the proposed MIM-based method benefits from the internal signal generated from the image itself compared to the handcraft supervision signal such as data distortion.</p>
<p>Secondly, we compared the classical MIM-based methods such as MAE (ViT-B) (He et al., <xref ref-type="bibr" rid="B11">2022</xref>) and SimMIM (Swin-B) (Xie et al., <xref ref-type="bibr" rid="B26">2022</xref>). The proposed GSC-MIM (ViT-B) outperformed MAE by 0.91&#x02013;5.26%. The improvement for Top-1 Acc indicates that the global semantic integrated strategy can help the model better acquire the input image&#x00027;s global long-dependence features and achieve higher classification results. Concerning no significant increments found between GSC-MIM and SimMIM, we suppose it is due to the advancement of Swin Transformer (Liu Z. et al., <xref ref-type="bibr" rid="B18">2021</xref>) as a backbone network.</p>
<p>Thirdly, another observation from the results is the accuracy increment compared to DINO(ViT-B) (Caron et al., <xref ref-type="bibr" rid="B4">2021</xref>), which is attributed to the information preserved complementary masking strategy of GSC-MIM that prevents the network from losing detailed information. We will provide visualization results in subsection 3.2.3. to further support our suggestion.</p>
<p>It also should be mentioned that under the scenario of limited training samples of UCM dataset, it is easy to overfit for the proposed GSC-MIM based on large amounts of parameters backbone such as ViT-L (Params. 304 M), and leads to unauthentic results.</p>
</sec>
<sec>
<title>3.2.2. Patch size as hyperparameter</title>
<p>Since small-scale features usually occupy relatively small areas in the RSI, the patch size theoretically affects the ability of the model to capture the small-scale features and the detailed information in the RSI. To further evaluate the impact of different patch sizes on the network effects, we conduct patch-size-specific experiments on the NWPU45 dataset. All experiments follow the default hyperparameters excluding patch size. The results are shown in <xref ref-type="table" rid="T3">Table 3</xref>. It is a rather evident tendency that classification accuracy is increasing with the decreasing of patch size. It is also worth noting that the training time of the network also increases substantially. This result is probably caused by two reasons:</p>
<list list-type="bullet">
<list-item><p>A small patch size will force the model to reconstruct more detailed information and thus learn fine-grained features of RSI.</p></list-item>
<list-item><p>Smaller patch size is more suitable for small-scale objects. When very small-scale objects, such as cars, are in remote sensing images, only a smaller patch size can retain adequate information.</p></list-item>
</list>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Experimental results comparison with patch size as hyperparameter.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="border-right: thin solid #000000;background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="left"><bold>Patch size</bold></th>
<th valign="top" align="center"><bold>Top-1 Acc (%)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">GSC-MIM<break/> (ViT-Base/p16)</td>
<td valign="top" align="center">32</td>
<td valign="top" align="center">66.91</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td/>
<td valign="top" align="center">16</td>
<td valign="top" align="center">72.47</td>
</tr>
<tr style="border-top: thin solid #000000;">
<td/>
<td valign="top" align="center">8</td>
<td valign="top" align="center">77.60</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We also provide attention map visualization diagrams to support our statements, shown as <xref ref-type="fig" rid="F6">Figure 6</xref>.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Attention map visualization diagrams with patch size of 8 <bold>(top)</bold>, 16 <bold>(middle)</bold>, and 32 <bold>(bottom)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0006.tif"/>
</fig>
</sec>
<sec>
<title>3.2.3. Visualization analysis</title>
<p>This section uses two visualization methods to interpret the network effects. We also perform the same visualization experiment for DINO (ViT-B) (Caron et al., <xref ref-type="bibr" rid="B4">2021</xref>) for a comparison. All the images are randomly selected from NWPU45 Dataset. As mentioned before, the mask of critical patches will cause information loss, so we proposed an information-preserved complementary masking strategy to mitigate this negative impact. To illustrate the effectiveness, we visualize the attention map of the last layer of each attention head in ViT. The results are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. Note that the colorful images are the visualization of [<italic>CLS</italic>] tokens. It can be seen that GSC-MIM shows visually more vital ability to detect multi-scale land objects and can separate different land objects or different parts of one land object apart, while DINO tends to capture the ambiguous semantic boundaries of land objects.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Attention map visualization for GSC-MIM <bold>(left)</bold> and DINO <bold>(right)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0007.tif"/>
</fig>
<p>Since scene labels are highly related to key areas and objects for RSI classification task, and accurately locating the attention regions is conducive to improving classification results, the proposed GSC-MIM integrates the input image&#x00027;s global semantic information into the network. To better reflect the key attention regions, we generated the energy map of the selected samples. As shown in <xref ref-type="fig" rid="F8">Figure 8</xref>, it is evident that the GSC-MIM-generated distribution of attention regions is accurate than DINO model.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Energy map visualization for GSC-MIM <bold>(left)</bold> and DINO <bold>(right)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-1083801-g0008.tif"/>
</fig>
</sec>
<sec>
<title>3.2.4. Discussion and future work</title>
<p>In this paper, we propose GSC-MIM, aiming to address two main problems encountered when applying MIM-based generative self-supervised learning to RSI scene classification tasks: the small-scale information absence problem and the global semantic information neglect problem, for which we propose the information-preserved complementary masking strategy (IPC-Masking) and semantic distillation strategy (GSD). Experimental results show that our model can obtain up to 5.26% accuracy improvement on the RSI scene classification task compared to the benchmark approaches.</p>
<p>We find that by introducing IPC-Masking, the network&#x00027;s ability to capture small-scale and fine-grained features is improved, and this ability is further enhanced as the patch size is reduced. Intuitively, the effectiveness of the network shows an inverse relationship with the patch size. However, when the patch size is excessively small (taking a single pixel as an example), the contained semantic information is lost; when the patch size is excessively large (assuming a whole input image as an example), the MIM mechanism loses its ability to work, so how to find the appropriate patch size for different RSI datasets is a problem worth further investigation.</p>
<p>Moreover, our experiments demonstrate that the model has accurate responses for the scene semantically meaningful regions when global information is introduced as a supervised signal. This proves the significance of global knowledge for using MIM in RSI scene classification tasks. Due to the specificity of the RSI imaging view, the critical information that determines the semantics of RSI scenes is often mixed with the background. When the global scene information of the field is not well-defined, it might not bring significant gain. Further clustering of the obtained global semantic futures to get more accurate scene information is a possible beneficial direction for future work.</p>
</sec>
</sec>
</sec>
<sec sec-type="conclusions" id="s4">
<title>4. Conclusion</title>
<p>This paper proposed a global semantic integrated self-distilled complementary masked image model, called GSC-MIM, for RSI scene classification. In GSC-MIM, the information preserved complementary masking strategy is proposed to prevent the information loss of local patches. The global semantic distillation strategy is employed to integrate the input image&#x00027;s global semantic information into the network and achieve label-free training. Experiments on the AID, NWPU-RESISC45, and UCM datasets indicate that the proposed GSC-MIM can better catch features of multi-scale land objects and achieves competitive classification accuracies.</p>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec sec-type="author-contributions" id="s6">
<title>Author contributions</title>
<p>XW: conceptualization, methodology, software, and writing&#x02014;original draft preparation. YZ: supervision. ZZ: software and validation. QL: data curation. JY: writing&#x02014;reviewing and editing. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This work was supported by the National Natural Science Foundation of China (Grants 41871364 and 42171376) and supported by the High Performance Computing Platform of Central South University.</p>
</sec>
<ack>
<p>We thank reviewers for valuable comments on the manuscript.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Akiva</surname> <given-names>P.</given-names></name> <name><surname>Purri</surname> <given-names>M.</given-names></name> <name><surname>Leotta</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Self-supervised material and texture representation learning for remote sensing tasks,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>8203</fpage>&#x02013;<lpage>8215</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.00803</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ayush</surname> <given-names>K.</given-names></name> <name><surname>Uzkent</surname> <given-names>B.</given-names></name> <name><surname>Meng</surname> <given-names>C.</given-names></name> <name><surname>Tanmay</surname> <given-names>K.</given-names></name> <name><surname>Burke</surname> <given-names>M.</given-names></name> <name><surname>Lobell</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Geography-aware self-supervised learning,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>10181</fpage>&#x02013;<lpage>10190</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01002</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>H.</given-names></name> <name><surname>Dong</surname> <given-names>L.</given-names></name> <name><surname>Wei</surname> <given-names>F.</given-names></name></person-group> (<year>2021</year>). <article-title>BEiT: BERT pre-training of image transformers</article-title>. <source>arXiv preprint arXiv:2106.08254</source>.</citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Caron</surname> <given-names>M.</given-names></name> <name><surname>Touvron</surname> <given-names>H.</given-names></name> <name><surname>Misra</surname> <given-names>I.</given-names></name> <name><surname>J&#x00027;egou</surname> <given-names>H.</given-names></name> <name><surname>Mairal</surname> <given-names>J.</given-names></name> <name><surname>Bojanowski</surname> <given-names>P.</given-names></name> <name><surname>Joulin</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Emerging properties in self-supervised vision transformers,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>9630</fpage>&#x02013;<lpage>9640</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00951</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Kornblith</surname> <given-names>S.</given-names></name> <name><surname>Norouzi</surname> <given-names>M.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;A simple framework for contrastive learning of visual representations,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Machine Learning</source>, Vol. <volume>119</volume>, <fpage>1597</fpage>&#x02013;<lpage>1607</lpage>.<pub-id pub-id-type="pmid">36188422</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;An empirical study of training self-supervised vision transformers,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>9640</fpage>&#x02013;<lpage>9649</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00950</pub-id><pub-id pub-id-type="pmid">36383492</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Remote sensing image scene classification: benchmark and state of the art</article-title>. <source>Proc. IEEE</source> <volume>105</volume>, <fpage>1865</fpage>&#x02013;<lpage>1883</lpage>. <pub-id pub-id-type="doi">10.1109/JPROC.2017.2675998</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Cino</surname> <given-names>L.</given-names></name> <name><surname>Mazzeo</surname> <given-names>P. L.</given-names></name> <name><surname>Distante</surname> <given-names>C.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Comparison of different supervised and self-supervised learning techniques in skin disease classification,&#x0201D;</article-title> in <source>IEEE International Conference on Image Information Processing</source> (<publisher-loc>Lecce</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>77</fpage>&#x02013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-031-06427-2_7</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ding</surname> <given-names>J.</given-names></name> <name><surname>Xie</surname> <given-names>E.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Jiang</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>P.</given-names></name> <name><surname>Xia</surname> <given-names>G.-S.</given-names></name></person-group> (<year>2021</year>). <article-title>Unsupervised pretraining for object detection by patch reidentification</article-title>. <source>arXiv preprint arXiv:2103.04814</source>. <pub-id pub-id-type="doi">10.1109/TPAMI.2022.3164911</pub-id><pub-id pub-id-type="pmid">35380956</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dosovitskiy</surname> <given-names>A.</given-names></name> <name><surname>Beyer</surname> <given-names>L.</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A.</given-names></name> <name><surname>Weissenborn</surname> <given-names>D.</given-names></name> <name><surname>Zhai</surname> <given-names>X.</given-names></name> <name><surname>Unterthiner</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. <source>arXiv preprint arXiv:2010.11929</source>.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Doll&#x000E1;r</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Masked autoencoders are scalable vision learners,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>16000</fpage>&#x02013;<lpage>16009</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.01553</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Fan</surname> <given-names>H.</given-names></name> <name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>S.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Momentum contrast for unsupervised visual representation learning,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>9729</fpage>&#x02013;<lpage>9738</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.00975</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Heidler</surname> <given-names>K.</given-names></name> <name><surname>Mou</surname> <given-names>L.</given-names></name> <name><surname>Hu</surname> <given-names>D.</given-names></name> <name><surname>Jin</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>G.</given-names></name> <name><surname>Gan</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Self-supervised audiovisual representation learning for remote sensing data</article-title>. <source>arXiv preprint arXiv:2108.00688</source>.</citation>
</ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>R.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Zhu</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Global and local contrastive self-supervised learning for semantic segmentation of HR remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>60</volume>, <fpage>1</fpage>&#x02013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2022.3147513</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>Z.</given-names></name></person-group> (<year>2021a</year>). <article-title>Semantic segmentation of remote sensing images with self-supervised multitask representation learning</article-title>. <source>IEEE J. Sel. Top. Appl. Earth Observ. Remote Sens</source>. <volume>14</volume>, <fpage>6438</fpage>&#x02013;<lpage>6450</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2021.3090418</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>Z.</given-names></name></person-group> (<year>2021b</year>). <article-title>Geographical knowledge-driven representation learning for remote sensing images</article-title>. <source>IEEE Trans. Geosci. Remote Sensors</source> <volume>60</volume>, <fpage>1</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2021.3115569</pub-id></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>F.</given-names></name> <name><surname>Hou</surname> <given-names>Z.</given-names></name> <name><surname>Mian</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Tang</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). <article-title>Self-supervised learning: Generative or contrastive</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <fpage>1</fpage>&#x02013;<lpage>20</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2021.3090866</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>H.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Swin transformer: hierarchical vision transformer using shifted windows,&#x0201D;</article-title> in <source>Proceedings of International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>10012</fpage>&#x02013;<lpage>10022</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00986</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Loshchilov</surname> <given-names>I.</given-names></name> <name><surname>Hutter</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>Decoupled weight decay regularization</article-title>. <source>arXiv preprint arXiv:1711.05101</source>.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>W.</given-names></name> <name><surname>Tao</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Qi</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name></person-group> (<year>2022</year>). <article-title>A unified deep learning framework for urban functional zone extraction based on multi-source heterogeneous data</article-title>. <source>Remote Sens. Environ</source>. <volume>270</volume>, <fpage>112830</fpage>. <pub-id pub-id-type="doi">10.1016/j.rse.2021.112830</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Manas</surname> <given-names>O.</given-names></name> <name><surname>Lacoste</surname> <given-names>A.</given-names></name> <name><surname>Gir&#x000F3;-i Nieto</surname> <given-names>X.</given-names></name> <name><surname>Vazquez</surname> <given-names>D.</given-names></name> <name><surname>Rodriguez</surname> <given-names>P.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Seasonal contrast: unsupervised pre-training from uncurated remote sensing data,&#x0201D;</article-title> in <source>Proceedings of International Conference on Computer Vision</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>9414</fpage>&#x02013;<lpage>9423</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00928</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tao</surname> <given-names>C.</given-names></name> <name><surname>Qi</surname> <given-names>J.</given-names></name> <name><surname>Lu</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Remote sensing image scene classification with self-supervised paradigm under limited labeled samples</article-title>. <source>IEEE Geosci. Remote Sens. Lett</source>. <volume>19</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2020.3038420</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Zhu</surname> <given-names>J.</given-names></name> <name><surname>Yan</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>LaST: label-free self-distillation contrastive learning with transformer architecture for remote sensing image scene classification</article-title>. <source>IEEE Geosci. Remote Sens. Lett</source>. <volume>19</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2022.3185088</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Albrecht</surname> <given-names>C. M.</given-names></name> <name><surname>Braham</surname> <given-names>N. A. A.</given-names></name> <name><surname>Mou</surname> <given-names>L.</given-names></name> <name><surname>Zhu</surname> <given-names>X. X.</given-names></name></person-group> (<year>2022</year>). <article-title>Self-supervised learning in remote sensing: a review</article-title>. <source>arXiv preprint arXiv:2206.13188</source>. <pub-id pub-id-type="doi">10.1109/MGRS.2022.3198244</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xia</surname> <given-names>G.-S.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Hu</surname> <given-names>F.</given-names></name> <name><surname>Shi</surname> <given-names>B.</given-names></name> <name><surname>Bai</surname> <given-names>X.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>AID: a benchmark data set for performance evaluation of aerial scene classification</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>55</volume>, <fpage>3965</fpage>&#x02013;<lpage>3981</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2017.2685945</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Lin</surname> <given-names>Y.</given-names></name> <name><surname>Bao</surname> <given-names>J.</given-names></name> <name><surname>Yao</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>&#x0201C;SimMIM: a simple framework for masked image modeling,&#x0201D;</article-title> in <source>Proceedings of IEEE Computer Society Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>New Orleans, LA</publisher-loc>), <fpage>9653</fpage>&#x02013;<lpage>9663</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR52688.2022.00943</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Newsam</surname> <given-names>S.</given-names></name></person-group> (<year>2010</year>). <article-title>&#x0201C;Bag-of-visual-words and spatial extensions for land-use classification,&#x0201D;</article-title> in <source>ACM SIGSPATIAL GIS</source> (<publisher-loc>New York, NY</publisher-loc>), <fpage>270</fpage>&#x02013;<lpage>279</lpage>. <pub-id pub-id-type="doi">10.1145/1869790.1869829</pub-id></citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Mei</surname> <given-names>X.</given-names></name> <name><surname>Tao</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>FALSE: false negative samples aware contrastive learning for semantic segmentation of high-resolution remote sensing image</article-title>. <source>IEEE Geosci. Remote Sens. Lett</source>. <volume>19</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2022.3222836</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Wei</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Shen</surname> <given-names>W.</given-names></name> <name><surname>Xie</surname> <given-names>C.</given-names></name> <name><surname>Yuille</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>iBOT: image bert pre-training with online tokenizer</article-title>. <source>arXiv preprint arXiv:2111.07832</source>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Q.</given-names></name> <name><surname>Lei</surname> <given-names>Y.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Guan</surname> <given-names>Q.</given-names></name> <name><surname>Zhong</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2022a</year>). <article-title>Knowledge-guided land pattern depiction for urban land use mapping: a case study of Chinese cities</article-title>. <source>Remote Sens. Environ</source>. <volume>272</volume>, <fpage>112916</fpage>. <pub-id pub-id-type="doi">10.1016/j.rse.2022.112916</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Q.</given-names></name> <name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Guan</surname> <given-names>Q.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Lin</surname> <given-names>W.</given-names></name></person-group> (<year>2022b</year>). <article-title>A weakly pseudo-supervised decorrelated subdomain adaptation framework for cross-domain land-use classification</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>60</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2022.3170335</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>
