<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Ecol. Evol.</journal-id>
<journal-title>Frontiers in Ecology and Evolution</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Ecol. Evol.</abbrev-journal-title>
<issn pub-type="epub">2296-701X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fevo.2022.840464</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Ecology and Evolution</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Lightweight Anchor-Free Subsidence Basin Detection Model With Adaptive Sample Assignment in Interferometric Synthetic Aperture Radar Interferogram</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Yu</surname> <given-names>Yaran</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1606281/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wang</surname> <given-names>Zhiyong</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1443643/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Zhenjin</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Ye</surname> <given-names>Kaile</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Hao</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Zihao</given-names></name>
</contrib>
</contrib-group>
<aff><institution>College of Geodesy and Geomatics, Shandong University of Science and Technology</institution>, <addr-line>Qingdao</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Yu Chen, China University of Mining and Technology, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Tao Li, Ministry of Natural Resources of the People&#x2019;s Republic of China, China; Hongyu Liang, Tongji University, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Zhiyong Wang, <email>skd994177@sdust.edu.cn</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Environmental Informatics and Remote Sensing, a section of the journal Frontiers in Ecology and Evolution</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>08</day>
<month>03</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>10</volume>
<elocation-id>840464</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>02</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Yu, Wang, Li, Ye, Li and Wang.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Yu, Wang, Li, Ye, Li and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>The excessive exploitation of coal resources has caused serious land subsidence, which seriously threatens the lives of the residents and the ecological environment in coal mining areas. Therefore, it is of great significance to precisely monitor and analyze the land subsidence in the mining area. To automatically detect the subsidence basins in the mining area from the interferometric synthetic aperture radar (InSAR) interferograms with wide swath, a lightweight model for detecting the subsidence basins with an anchor-free and adaptive sample assignment based on the YOLO V5 network, named Light YOLO-Basin model, is proposed in this paper. First, the depth and width scaling of the convolution layers and the depthwise separable convolution are used to make the model lightweight to reduce the memory consumption of the CSPDarknet53 backbone network. Furthermore, the anchor-free detection box encoding method is used to deal with the inapplicability of the anchor box parameters, and an optimal transport assignment (OTA) adaptive sample assignment method is introduced to solve the difficulty of optimizing the model caused by abandoning the anchor box. To verify the accuracy and reliability of the proposed model, we acquired 62 Sentinel-1A images over Jining and Huaibei coalfield (China) for the training model and experimental verification. In contrast with the original YOLO V5 model, the mean average precision (mAP) value of the Light YOLO-Basin model increases from 45.92 to 55.12%. The lightweight modules of the model sped up the calculation with the one billion floating-point operations (GFLOPs) from 32.81 to 10.07 and reduced the parameters from 207.10 to 40.39 MB. The Light YOLO-Basin model proposed in this paper can effectively recognize and detect the subsidence basins in the mining areas from the InSAR interferograms.</p>
</abstract>
<kwd-group>
<kwd>InSAR</kwd>
<kwd>subsidence basin detecting</kwd>
<kwd>YOLO V5</kwd>
<kwd>depthwise separable convolution</kwd>
<kwd>anchor-free</kwd>
<kwd>optimal transport assignment (OTA)</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="13"/>
<table-count count="6"/>
<equation-count count="13"/>
<ref-count count="65"/>
<page-count count="16"/>
<word-count count="11232"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>Large-scale land subsidence resulting from coal mining has caused a series of ecological and environmental problems, including destroying farmlands, damaging buildings, and even inducing geological disasters (<xref ref-type="bibr" rid="B50">Wang et al., 2021a</xref>; <xref ref-type="bibr" rid="B63">Yuan et al., 2021</xref>). It threatens the lives and property of the local residents and restricts the economic sustainable development in mining areas (<xref ref-type="bibr" rid="B15">Fan et al., 2018</xref>; <xref ref-type="bibr" rid="B49">Wang et al., 2020</xref>). Therefore, it is of great significance to continuously monitor and analyze the land subsidence in mining areas. Traditional geodetic surveying methods mainly include precise leveling, total station measurement, and global navigation satellite system (GNSS) (<xref ref-type="bibr" rid="B15">Fan et al., 2018</xref>). These methods have some limitations such as low spatial resolution, limited monitoring area, and low observation efficiency (<xref ref-type="bibr" rid="B4">Chen et al., 2015</xref>; <xref ref-type="bibr" rid="B9">Chen Y. et al., 2020</xref>; <xref ref-type="bibr" rid="B44">Shi M. Y. et al., 2021</xref>). Interferometric synthetic aperture radar (InSAR) technology has become a new means for monitoring the surface deformation in mining areas with the advantages of day/night data acquisition, all-weather imaging capability, and strong penetrability (<xref ref-type="bibr" rid="B35">Ng et al., 2017</xref>; <xref ref-type="bibr" rid="B5">Chen B. Q. et al., 2020</xref>; <xref ref-type="bibr" rid="B49">Wang et al., 2020</xref>).</p>
<p>Currently, the main research directions of InSAR technology for monitoring mine subsidence have evolved from acquiring single surface deformation information to three-dimensional deformation information or subsidence prediction based on deformation theory (<xref ref-type="bibr" rid="B61">Yang et al., 2017a</xref>,<xref ref-type="bibr" rid="B62">b</xref>, <xref ref-type="bibr" rid="B58">2018a</xref>,<xref ref-type="bibr" rid="B59">b</xref>; <xref ref-type="bibr" rid="B6">Chen et al., 2021</xref>; <xref ref-type="bibr" rid="B12">Dong et al., 2021</xref>; <xref ref-type="bibr" rid="B16">Fan et al., 2021</xref>). Most of these works are targeted at one or more subsidence basins. However, the imaging mode of mainstream synthetic aperture radar (SAR) satellites, such as ALOS-2 and RadarSat-2, has an image width of more than 80 km. The image width of the interferometric wide swath (IW) mode of the Sentinel-1A satellite even reaches 250 km (<xref ref-type="bibr" rid="B65">Zheng et al., 2018</xref>). It is time-consuming and labor-intensive to search for subsidence basins with a radius of only a few hundred meters in a wide range of images. At present, there have been few studies on the automatic detection of subsidence basins from the InSAR interferogram. <xref ref-type="bibr" rid="B24">Hu et al. (2013)</xref> proposed a differential interferometric synthetic aperture radar (D-InSAR) based illegal-mining detection system, aiming to increase both accuracy and efficiency of underground-mining detection. The detection results are highly dependent on the quality of the phase unwrapping and subjective processing experience. By using the methods of D-InSAR technology, geographic information system (GIS), and mining subsidence, <xref ref-type="bibr" rid="B55">Xia et al. (2018)</xref> proposed a novel theory of effectively recognizing the subsidence basins. This method can detect subsidence basins without manual intervention. However, like the method proposed by <xref ref-type="bibr" rid="B24">Hu et al. (2013)</xref>, the detection accuracy will also be affected by the quality of the phase unwrapping. <xref ref-type="bibr" rid="B60">Yang et al. (2018c)</xref> proposed a space-based method for recognizing the subsidence basins by relating the geometric parameters of subsidence basins to the InSAR derived line of sight deformation with the probability integral method (PIM). This method can determine the boundary of the subsidence basins, but it is not suitable for detecting subsidence basins in large-scale areas. <xref ref-type="bibr" rid="B13">Du et al. (2019)</xref> proposed a feature-point-based method for efficiently detecting subsidence basins. Du first used D-InSAR to monitor subsidence basins caused by mining and then used the PIM to determine the inflection and boundary points of the subsidence basins. <xref ref-type="bibr" rid="B51">Wang et al. (2021b)</xref> proposed a model for detecting the subsidence basin based on the histogram of oriented gradient features and support vector machine classifier. This method is limited by the feature detection operator and cannot effectively detect the subsidence basin with obscure edge features and too small scope. <xref ref-type="bibr" rid="B1">Bala et al. (2021)</xref> proposed a circlet transform method for detecting subsidence basins. This method reduces manual intervention, but the detection time is longer. Due to the long detection time, the above methods are difficult to carry out in a large-scale area. Furthermore, these methods conduct detection mostly based on the deformation gradient and the shape characteristics of the subsidence basin on the deformation map. The quality of the phase unwrapping is susceptible to atmospheric effects and noises, resulting in lower detection accuracy. Therefore, we propose a new method to detect subsidence basins from InSAR interferogram with large-scale areas.</p>
<p>After underground coal mining, a series of subsidence basins will appear in the mining area. Subsidence basins are scattered and large in number in the InSAR interferogram. It is difficult to identify these subsidence basins manually. The single subsidence basin of the mining area in the InSAR interferogram is approximately concentric circles or concentric ellipses, with a small scale and obvious edge features. Currently, the convolution neural network (CNN)-based objection detection method has been widely applied in many research fields (<xref ref-type="bibr" rid="B29">LeCun et al., 2015</xref>; <xref ref-type="bibr" rid="B43">Shi et al., 2020</xref>; <xref ref-type="bibr" rid="B41">Ren et al., 2021</xref>). For SAR images, it is mainly used to identify ships (<xref ref-type="bibr" rid="B3">Chang et al., 2019</xref>; <xref ref-type="bibr" rid="B26">Jiang et al., 2021</xref>; <xref ref-type="bibr" rid="B54">Wu Z. T. et al., 2021</xref>) and marine oil spills, etc. The CNN-based objection detection method can realize the automatic detection of subsidence basins. CNN-based objection detection frameworks primarily consist of three components, including backbone network, neck network, and detection head (<xref ref-type="bibr" rid="B8">Chen Q. S. et al., 2020</xref>; <xref ref-type="bibr" rid="B18">Fu K. et al., 2020</xref>). The backbone network mainly extracts the basic features of the input image, such as the ResNet (<xref ref-type="bibr" rid="B22">He et al., 2016</xref>; <xref ref-type="bibr" rid="B56">Xie et al., 2017</xref>) series and the DarkNet series. The main function of the neck network is to further strengthen the features extracted by the backbone network. For example, the feature pyramid network (FPN) combined features of different scales with lateral connections in a top-down manner to construct a series of scale-invariant feature maps, and multiple scale-dependent classifiers were trained on these feature pyramids (<xref ref-type="bibr" rid="B31">Lin et al., 2017a</xref>). The detection head network is responsible to predict and refine the bounding box, calculating the bounding box coordinates, confidence score, and classification score. According to the different head networks, object detection frameworks can be primarily divided into two categories. One is two-stage detectors that have high detection accuracy, mainly including R-CNN (<xref ref-type="bibr" rid="B21">Girshick et al., 2014</xref>), Fast-RCNN (<xref ref-type="bibr" rid="B20">Girshick, 2015</xref>), and Faster-RCNN (<xref ref-type="bibr" rid="B42">Ren et al., 2015</xref>), etc. Two-stage detectors first use a region proposal network to generate a sparse set of candidate object bounding boxes and then to extract features from each candidate bounding box for the following classification and bounding box regression tasks. The other is one-stage detectors that achieve high inference speed, mainly including the YOLO series (<xref ref-type="bibr" rid="B40">Redmon et al., 2016</xref>; <xref ref-type="bibr" rid="B38">Redmon and Farhadi, 2017</xref>, <xref ref-type="bibr" rid="B39">2018</xref>; <xref ref-type="bibr" rid="B2">Bochkovskiy et al., 2020</xref>), SSD (<xref ref-type="bibr" rid="B33">Liu et al., 2016</xref>), and RetinaNet (<xref ref-type="bibr" rid="B32">Lin et al., 2017b</xref>), etc. One-stage detectors generate prediction boxes, confidence scores, and object classes concurrently.</p>
<p>At present, many CNN-based object detection methods have been proposed, most of which are designed to detect objects in natural images. However, there are two problems for recognizing the subsidence basins when directly using these methods. First, due to the various object categories and shapes in natural images, big networks (such as DarkNet53, Resnet101) are often used as the backbone. However, the shape of the subsidence basin is relatively simple and the detection category is single, hence there is no need for a heavy network to detect subsidence basin. In addition, with the continuous development of SAR satellites, the image width is also increasing, then it will also increase the burden on the computer hardware when using heavy networks. Second, the anchor boxes obtained by clustering object structure in natural images are not suitable for the detection of subsidence basins.</p>
<p>The YOLO V5 model, a one-stage detector, has the advantages of high accuracy and fast speed. It has been widely used in various domains such as face recognition, text detection, and logo detection (<xref ref-type="bibr" rid="B52">Wu W. et al., 2021</xref>; <xref ref-type="bibr" rid="B57">Yan et al., 2021</xref>; <xref ref-type="bibr" rid="B64">Zhao et al., 2021</xref>). In order to automatically recognize and detect subsidence basins in large-scale mining areas with high accuracy, we proposed a lightweight detection model with adaptive sample assignment. The proposed model is based on the path aggregation network (PANet) of YOLO V5.</p>
<p>The main sections of this paper are organized as follows. In section &#x201C;Data and Materials,&#x201D; the model proposed in this paper for detecting the subsidence basins from the InSAR interferogram is described. It mainly includes the depthwise separable convolution, anchor-free, and OTA adaptive sample assignment. The experimental results and quantitative evaluation are presented in Section &#x201C;Method.&#x201D; Section &#x201C;Results and Analysis&#x201D; shows the discussions and the analysis of each module in the proposed model and the comparison results with the original YOLOV5 model. Finally, some valuable conclusions of this study are drawn in Section &#x201C;Discussion.&#x201D;</p>
</sec>
<sec id="S2">
<title>Data and Materials</title>
<sec id="S2.SS1">
<title>Study Area</title>
<p>We selected the Jining and Huaibei mining areas in China as the study areas. The two mining areas are rich in coal resources and have a long mining history. There are many mines in the two areas with a complex distribution. The Jining mining area (115&#x00B0;50&#x2032;&#x2013;117&#x00B0;48&#x2032;E, 34&#x00B0;58&#x2032;&#x2013;35&#x00B0;59&#x2032;N) is located in the southwest of Shandong Province, China, with a cumulative proven coal reserve of nearly 15.1 billion tons. The Huaibei mining area (116&#x00B0;23&#x2032;&#x2013;117&#x00B0;12&#x2032;E, 33&#x00B0;16&#x2032;&#x2013;34&#x00B0;14&#x2032;N) is located in the north of the Huaihe River of Anhui Province, China, with a cumulative proven coal reserve of 13 billion tons. The locations of the study areas are shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Location of the study area. <bold>(A)</bold> The geographic location of the Jining mining area; <bold>(B)</bold> geographic location of the Huaibei mining area.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g001.tif"/>
</fig>
</sec>
<sec id="S2.SS2">
<title>Experimental Data</title>
<p>We used 24 Sentinel-1A images acquired from November 2017 to March 2020 over the Jining mining area and used 34 Sentinel-1A images acquired from November 2017 to March 2020 over the Huaibei mining area. The specific information of the partial interferometric pairs is shown in <xref ref-type="table" rid="T1">Table 1</xref>. The experimental Sentinel-1A data are interferometric wide swath imaging mode (Level 1 single-look complex images products) and in VV polarization with the 38.9&#x00B0; incidence angle. The revisit period of Sentinel-1A is 12 days and has a pixel size of about 2.33 &#x00D7; 13.91 m. Moreover, the shuttle radar topography mission digital elevation model (SRTM DEM) released by the National Aeronautics and Space Administration (NASA) was applied to remove both the flattening and terrain phases in D-InSAR data processing.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Sentinel-1A interferometric pairs for constructing the training/testing datasets and verifying the performance of the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td/>
<td valign="top" align="center">Mining area</td>
<td valign="top" align="center">Master image</td>
<td valign="top" align="center">Slave image</td>
<td valign="top" align="center">Path</td>
<td valign="top" align="center">Frame</td>
<td valign="top" align="center">Temporal baseline/d</td>
<td valign="top" align="center">Perpendicular baseline/m</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Datasets</td>
<td valign="top" align="center">Jining</td>
<td valign="top" align="center">28/11/2017<xref ref-type="table-fn" rid="t1fns1">&#x002A;</xref></td>
<td valign="top" align="center">22/12/2017</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">111</td>
<td valign="top" align="center">24</td>
<td valign="top" align="center">101.52</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Jining</td>
<td valign="top" align="center">11/11/2018</td>
<td valign="top" align="center">05/12/2018</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">111</td>
<td/>
<td valign="top" align="center">22.76</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Jining</td>
<td valign="top" align="center">30/11/2019</td>
<td valign="top" align="center">24/12/2019</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">111</td>
<td/>
<td valign="top" align="center">56.14</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Huaibei</td>
<td valign="top" align="center">10/12/2017</td>
<td valign="top" align="center">03/01/2018</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">106</td>
<td/>
<td valign="top" align="center">81.53</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Huaibei</td>
<td valign="top" align="center">11/11/2018</td>
<td valign="top" align="center">05/12/2018</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">106</td>
<td/>
<td valign="top" align="center">21.79</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Huaibei</td>
<td valign="top" align="center">18/11/2019</td>
<td valign="top" align="center">12/12/2019</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">106</td>
<td/>
<td valign="top" align="center">115.70</td>
</tr>
<tr>
<td valign="top" align="left">Verifying data</td>
<td valign="top" align="center">Jining</td>
<td valign="top" align="center">30/12/2020</td>
<td valign="top" align="center">11/01/2021</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">111</td>
<td valign="top" align="center">12</td>
<td valign="top" align="center">25.10</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Huaibei</td>
<td valign="top" align="center">11/01/2021</td>
<td valign="top" align="center">23/01/2021</td>
<td valign="top" align="center">142</td>
<td valign="top" align="center">106</td>
<td/>
<td valign="top" align="center">&#x2212;22.02</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="t1fns1"><p><italic>&#x002A;Date/month/year.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>In order to obtain the differential interferograms of the two mining areas, we used the two-pass D-InSAR technology (<xref ref-type="bibr" rid="B36">Ou et al., 2018</xref>; <xref ref-type="bibr" rid="B11">Dai et al., 2020</xref>) to process these Sentinel-1A data. The procedures of D-InSAR technology mainly include interferogram generation, SAR simulation based on digital elevation model (DEM), differential processing between the real interferogram and the simulated interferogram, phase unwrapping, transformation from phase to deformation, and geocoding (<xref ref-type="bibr" rid="B25">Ilieva et al., 2019</xref>; <xref ref-type="bibr" rid="B7">Chen D. H. et al., 2020</xref>). <xref ref-type="fig" rid="F2">Figure 2</xref> is a flow chart of the two-pass D-InSAR data processing. Interferograms with serious decoherence and large noise influence were excluded. We set the temporal baseline threshold of the interferogram to 36 days and the spatial baseline threshold to 200 m. The longest spatial baseline of the interferogram in this paper is 155.71 m. We obtained 62 interferograms in the Jining and Huaibei mining areas. The ratio of multi-looking is 1:5. The pixel size in the range direction is 18.54 m, and the pixel size in the azimuth direction is 13.89 m for the interferogram. The interferograms are too large to be used by the deep neural architecture, which generally accepts an image with a size of 416 &#x00D7; 416 as an input. Therefore, we segment the interferograms into smaller sub-images with a size of 416 &#x00D7; 416 according to the standard of the YOLO V5 model. The image data annotation software called &#x201C;LabelImg&#x201D; was used to draw the outer rectangular boxes of the subsidence basins on each sub-image, realizing the manual annotation of the border and labels of the ground truth box. The ground truth boxes were annotated according to the features of the subsidence basin. The subsidence basin on the InSAR interferogram is a series of approximately circular or elliptical interferometric fringes with a small scale and obvious edge features (<xref ref-type="bibr" rid="B51">Wang et al., 2021b</xref>). We have obtained a total of 1,160 sample images. In this study, 812 sample images were randomly selected from 1,160 images as the training samples; the remaining 30% were selected as the testing samples, which had a total of 348 images. The partial examples of the sample datasets are shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. In order to alleviate the over-fitting phenomenon during the training model incurred by limited sample datasets, rotation, translation, and flipping were used for data augmentation.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>The data processing flow of the two-pass differential interferometric synthetic aperture radar (D-InSAR). The blue part is the data processing implemented in the paper.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Some examples of sample datasets. <bold>(A)</bold> Partial training data; <bold>(B)</bold> partial testing data. The red boxes represent the manual annotation of the ground truth box in panel <bold>(A)</bold>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g003.tif"/>
</fig>
<p>We selected two Sentinel-1A images acquired from December 2020 to January 2021 over the Jining mining area and two Sentinel-1A images acquired from January 2021 over the Huaibei mining area, constituting a total of two interferometric pairs to test the performance of the proposed model for detecting the subsidence basins from the interferograms. The parameters of the two interferometric pairs are listed in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
</sec>
</sec>
<sec id="S3" sec-type="method">
<title>Method</title>
<p>To automatically detect the subsidence basins from the InSAR interferograms, a lightweight detection model with an adaptive sample assignment based on the YOLO V5 model was proposed. We first use the channel number scaling and depthwise separable convolution (<xref ref-type="bibr" rid="B23">Howard et al., 2017</xref>) to make the CSPDarknet53 network lightweight, and then model the prediction boxes as a width and height fit problem based on the center point like the anchor-free strategy in FCOS (<xref ref-type="bibr" rid="B48">Tian et al., 2019</xref>). In addition, we also introduced an OTA (<xref ref-type="bibr" rid="B19">Ge et al., 2021</xref>) that assigns positive and negative samples in an adaptive manner. The proposed model for detecting the subsidence basin was named the Light YOLO-Basin model. The network architecture of the Light YOLO-Basin model is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Architecture of the Light YOLO-Basin network. SPP is spatial pyramid pooling. Upsample uses bilinear interpolation. Obj and distance intersection over union loss (DIoU) represent confidence loss and bounding box regression loss, respectively. In addition, convolution modules that can be lightweight are in orange.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g004.tif"/>
</fig>
<sec id="S3.SS1">
<title>YOLO V5 Network</title>
<p>The YOLO V5 model, as the basic framework for detecting the subsidence basins in this paper, mainly consists of three components: backbone network, neck network, and detection head. The backbone network is designed to extract the features of the subsidence basin, mainly composed of the CSPDarknet53 network and spatial pyramid pooling (SPP) (<xref ref-type="bibr" rid="B37">Purkait et al., 2017</xref>). The neck network adopts the path aggregation network instead of feature pyramid networks in YOLO V5. The detection head, as the final detection part of the model, is used to output the detection results of the subsidence basin. It utilizes the high-level semantic information outputted from the neck network to classify the category and regress the location of the objects.</p>
<p>The loss function of the YOLO V5 network mainly consists of three parts: bounding box regression loss, classification loss, and confidence loss (<xref ref-type="bibr" rid="B45">Shi P. F. et al., 2021</xref>). Since this paper only has the category of the subsidence basins, we only used confidence loss and bounding box regression loss, as shown in Formula (1).</p>
<disp-formula id="S3.E1"><label>(1)</label><mml:math id="M1"><mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>L</italic><sub><italic>obj</italic></sub> and <italic>L</italic><sub><italic>DIoU</italic></sub> mean confidence loss and bounding box regression loss, respectively. &#x03BB; is the balancing factor and the value is 5 (<xref ref-type="bibr" rid="B57">Yan et al., 2021</xref>).</p>
<p>The Light YOLO-Basin model used the Focal Loss (<xref ref-type="bibr" rid="B32">Lin et al., 2017b</xref>) function confidence loss to alleviate the problem caused by the imbalanced number of hard and easy samples, as shown in Formula (2), in which the positive sample <italic>p</italic> with the high confidence is the easy sample and vice versa. Focal loss reduces the weight of easy samples so that the model can focus more on hard samples, ensuring that the contributions of all samples to model parameter updating are relatively balanced.</p>
<disp-formula id="S3.E2"><label>(2)</label><mml:math id="M2"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi mathvariant="normal">&#x03B3;</mml:mi></mml:msup><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="left"><mml:mrow><mml:mrow><mml:mrow><mml:mo>-</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi mathvariant="normal">&#x03B1;</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mi mathvariant="normal">&#x03B3;</mml:mi></mml:msup><mml:mi>l</mml:mi><mml:mi>o</mml:mi><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mi>p</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mi/></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>y</italic> = 1 means positive samples and <italic>y</italic> = 0 means negative samples; the parameter &#x03B1; is used to balance the weight of the positive and negative samples during the model training; the parameter &#x03B3; is used to balance the weight of the easy samples in the model; the parameter <italic>p</italic> &#x03F5; [0,1] is the model estimated probability for the confidence loss. The bounding box regression loss of the original YOLO V5 model adopts generalized intersection over union loss (GIoU Loss), but it only focuses on the overlapping areas and other non-overlapping areas, which ignores the impact of the bounding box on the IoU. The Light YOLO-Basin model used distance IoU loss (DIoU Loss) (<xref ref-type="bibr" rid="B34">Luo et al., 2020</xref>) as the bounding box regression loss. DIoU Loss is defined as Formula (3). It not only considers the distance of the center and the overlapping area between the ground truth box and prediction box, but also minimizes the center point distance.</p>
<disp-formula id="S3.E3"><label>(3)</label><mml:math id="M3"><mml:mtable displaystyle="true" rowspacing="0pt"><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:mrow><mml:mtext>I</mml:mtext><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo>&#x2229;</mml:mo><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd columnalign="center"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac><mml:mrow><mml:msup><mml:mi mathvariant="normal">&#x03C1;</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>b</mml:mi><mml:mrow><mml:mi>g</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:msup><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mfrac></mml:mstyle></mml:mrow></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where A represents the area of the prediction box and B represents the area of the ground truth box; the parameter <italic>b</italic> and <italic>b<sup>gt</sup></italic> mean the center points of the prediction box A and ground truth box B, respectively; &#x03C1;<sup>2</sup> is the Euclidean Distance between <italic>b</italic> and <italic>b<sup>gt</sup></italic>; the parameter c represents the diagonal length of the smallest closed shape that includes the ground truth box A and the prediction box B.</p>
<p>The detection head module of the YOLO V5 model directly uses a single convolutional layer to calculate the classification loss, confidence loss, and bounding box regression loss. However, there is no classification loss in our study, the structure of the detection head requires to be changed. Some researches demonstrated that there is a conflict between classification and regression tasks (<xref ref-type="bibr" rid="B46">Song et al., 2020</xref>; <xref ref-type="bibr" rid="B53">Wu et al., 2020</xref>), so referencing the ideas of FCOS (<xref ref-type="bibr" rid="B48">Tian et al., 2019</xref>) and the literature (<xref ref-type="bibr" rid="B46">Song et al., 2020</xref>), the Light YOLO-Basin model used double-head as the detection head module. The Light YOLO-Basin model architecture after decoupling is shown in <xref ref-type="fig" rid="F4">Figure 4</xref>. The double-head splits the output features of the subsidence basin into regression and confidence branches. The regression branch provides prediction box coordinates. Meanwhile, the confidence branch calculates the probability of positive samples.</p>
</sec>
<sec id="S3.SS2">
<title>The Lightweight of the CSPDarknet53 Network</title>
<p>At present, the YOLO V5 model has achieved great success in natural image datasets such as PASCAL VOC, ImageNet, and MS COCO. However, compared with objects in natural images, subsidence basins on the interferograms have more obvious edges and texture features and simpler shapes. Intuitively, the detection of the subsidence basins on the interferogram from D-InSAR is simpler than the multi-category detection on natural images. This also leads to the conjecture of whether the heavy CSPDarknet53 network is necessary for detecting the subsidence basins. In order to verify that, we used the three versions of the YOLO V5 model (Large, Middle, and Small) to detect the subsidence basins from the interferograms. The variation of detection accuracy is shown in <xref ref-type="fig" rid="F5">Figure 5</xref>. Experimental results show that as the number of parameters and computations of the model reduces, the detection accuracy of the model shows a rising trend instead of decreasing. The results demonstrate that the detection of subsidence basins does not require a heavy network. In addition, since the swath of SAR images is large wide, for example, the swath of Sentinel-1 is about 250 km, a lightweight model is used to detect subsidence basins from the interferograms to lighten the burden on the computer hardware. Therefore, we introduced depthwise separable convolutions to make the CSPDarknet53 network lightweight.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Comparison results between the mean average precision (mAP) and parameters/inference memory of the three versions of YOLO V5 (Large, Middle, and Small) model. The abscissa in panel <bold>(A)</bold> represents the parameter amount of the model; the abscissa in panel <bold>(B)</bold> means the inference memory of the model for processing a 416 &#x00D7; 416 image. The numerical values of the mAP of the three versions of the YOLO V5 model.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g005.tif"/>
</fig>
<p>Compared with standard convolution, depthwise separable convolution generally sacrifices a small amount of detection accuracy to save the computations and parameters of the model. Depthwise separable convolution is a form of factorized convolutions that factorizes a standard convolution into a depthwise convolution and a pointwise convolution. Firstly, the depthwise convolution applies a single convolution to each channel feature map to extract feature information and keeps the number of feature maps unchanged. Secondly, the pointwise convolution applies multiple 1 &#x00D7; 1 convolutions to combine the feature maps obtained by the depthwise convolution. The process of performing depthwise separable convolution on a feature map with the size of W &#x00D7; H &#x00D7; C (W and H are the spatial width and height of the feature map, respectively, and C is the number of feature channels) is shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. The computation of the standard convolution is 3 &#x00D7; 3 &#x00D7; C &#x00D7; N &#x00D7; W &#x00D7; H, while the computation of the depthwise separable convolution is 3 &#x00D7; 3 &#x00D7; C &#x00D7; W &#x00D7; H+1 &#x00D7; 1 &#x00D7; C &#x00D7; N &#x00D7; W &#x00D7; H = (3 &#x00D7; 3+N) &#x00D7; (C &#x00D7; W &#x00D7; H), which is approximately 1/N + 1/(3 &#x00D7; 3) of the standard convolution. The parameter of the standard convolution is 3 &#x00D7; 3 &#x00D7; C &#x00D7; N, while the parameter of the depthwise separable convolution is 3 &#x00D7; 3 &#x00D7; C + 1 &#x00D7; 1 &#x00D7; C &#x00D7; N, which is about 1/N + 1/(3 &#x00D7; 3) of the standard convolution. The depthwise separable convolution has the effect of drastically reducing the computation and model size. In this study, the depthwise separable convolution is introduced into the CSPDarknet53 network, which is combined with the scaling of the depth (the number of convolutional layers) and the width (the number of channels of the convolution kernel) to realize the lightweight of the model, as shown in the orange part of <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>The structure diagram of depthwise separable convolution. W and H are the spatial width and height of the feature map, respectively. C is the number of feature channels; N is the number of feature channels after depthwise separable convolution.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g006.tif"/>
</fig>
</sec>
<sec id="S3.SS3">
<title>Anchor-Free and Adaptive Sample Assignment</title>
<p>YOLO V5, an anchor-based method, predicts boxes by fitting the deviation of the anchor boxes, where the anchor boxes have a preset width and height. Although anchor-based methods improve the detection accuracy to a certain extent, there are still some drawbacks (<xref ref-type="bibr" rid="B17">Fu J. M. et al., 2020</xref>). First, a large number of anchor boxes are used in the model, resulting in excessive redundant computation and slowing down the detection speed of the model. Second, only a tiny fraction of the anchor boxes is labeled as positive samples, resulting in a huge imbalance between positive and negative samples. Third, fixed anchor boxes cannot be applied to various data, which increases the difficulty of model optimization. Anchor boxes generally require to be reset according to different data. Currently, the anchor boxes obtained by clustering objects in natural images are quite different from the features of the subsidence basins in the InSAR interferograms. In addition, since the training data cannot represent all the subsidence basins, the anchor-boxes obtained by clustering the characteristics of the subsidence basins cannot guarantee their versatility. Therefore, we used an anchor-free method to detect the subsidence basins from interferograms.</p>
<p>The location of the prediction box in the original YOLO V5 model is determined by the center point and the width and height of the anchor box, as shown in <xref ref-type="fig" rid="F7">Figure 7A</xref>. The coordinate value of the center point is represented by the offset from the top left corner of the cell, and the width and height are the scaling ratios of the corresponding anchor box. However, if the anchor box is abandoned, the width and height cannot be expressed. Hence, the Light YOLO-Basin model directly fits the width and height instead of the ratio of the anchor box and keeps the computation of the center point unchanged to achieve anchor-free (<xref ref-type="bibr" rid="B28">Law and Deng, 2018</xref>; <xref ref-type="bibr" rid="B14">Duan et al., 2019</xref>), as shown in <xref ref-type="fig" rid="F7">Figure 7B</xref>.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Location encoding of bounding box in YOLO V5 and Light YOLO-Basin. <bold>(A)</bold> Location encoding of YOLO V5 bounding box. The red box denotes the anchor box and the black box denotes the prediction box. <italic>P<sub>w</sub></italic> and <italic>P<sub>h</sub></italic> mean the width and height of the anchor box, respectively; <italic>t<sub>w</sub></italic> and <italic>t<sub>h</sub></italic> are the predicted ratio of the width and height between the prediction box and anchor box. <bold>(B)</bold> Anchor-free location encoding of the bounding box. The black box denotes the prediction box; <italic>t<sub>w</sub></italic> and <italic>t<sub>h</sub></italic> directly predicted the width and height. <italic>b<sub>x</sub></italic> and <italic>b<sub>y</sub></italic>, respectively represent the coordinates of the center point <italic>x</italic> and <italic>y</italic> of the prediction box; <italic>b<sub>w</sub></italic> and <italic>b</italic><sub><italic>h</italic></sub> represent the width and height of the prediction box, respectively. <italic>t<sub>x</sub></italic> and <italic>t<sub>y</sub></italic> are the offsets between the center point of the prediction box and the top left corner of the cell computed by the model; &#x03C3; is the Sigmoid function; <italic>C<sub>x</sub></italic> and <italic>C<sub>y</sub></italic> are the coordinates of the top left corner of the cell.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g007.tif"/>
</fig>
<p>The lack of the constraint of anchor boxes in the anchor-free method increases the degree of freedom of the model and also increases the difficulty of optimizing the prediction box. Adopting IoU at a certain threshold as positive and negative assignment criterion, which is commonly used in anchor-based detection methods, is usually not suitable for the anchor-free method. For example, the two-stage Faster R-CNN network labels the prediction box with an IoU value greater than 0.7 as a positive sample and less than 0.3 as a negative sample; the one-stage YOLO V3 model labels the prediction box with the highest IoU value as a positive sample, as shown in <xref ref-type="fig" rid="F8">Figure 8B</xref>. Moreover, based on the anchor-free method, using a certain IoU threshold to assign prediction boxes will cause a large number of useful prediction boxes to be discarded. To increase the number of positive samples, the YOLO V5 model expands the selection range of positive samples to the surrounding pixels, reducing the difficulty of model optimization, as shown in <xref ref-type="fig" rid="F8">Figure 8C</xref>. However, the YOLO V5 model still selects a fixed number of positive samples with a certain IoU threshold, and it cannot adaptively determine the number and assignment of hard and easy positive samples. Therefore, we used the OTA (<xref ref-type="bibr" rid="B19">Ge et al., 2021</xref>) to model the positive and negative samples assignment as an optimal transmission problem, which selects and balances positive and negative samples in an adaptive manner.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>Sample assignment of YOLO V3, YOLO V5, and optimal transport assignment (OTA). Panel <bold>(A)</bold> is the image with bounding box; panel <bold>(B)</bold> is the sample assignment of YOLO V3; panel <bold>(C)</bold> is the sample assignment of YOLO V5; panel <bold>(D)</bold> is the adaptive sample assignment of OTA. The value is the transporting cost value in panel <bold>(D)</bold>. The blue grid and the green box denote the positive sample and the ground truth box, respectively, and the red point is the center point of the ground truth.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g008.tif"/>
</fig>
<p>The OTA method treats the ground truth box as the supplier in the optimal transport theory, and the prediction box in the model training as the demander. The unit transportation cost between each demander and supplier is defined as the weighted summation of losses between the ground truth box and prediction box. If a demander receives enough goods from the supplier, this demander becomes one positive sample. The model needs the best positive and negative samples assignment solution to minimize the global transportation cost. Concretely, assuming that there are <italic>m</italic> ground truth boxes and <italic>n</italic> prediction boxes for image <italic>I</italic>, the ground truth box, namely supplier, holds <italic>k</italic> units of goods, while the prediction box, namely demander, needs d units of goods. c<italic><sub><italic>i,j</italic></sub></italic> represents the transporting cost of each unit of good from the <italic>i</italic>-th supplier to the <italic>j</italic>-th demander. The positive and negative sample assignment problem can be defined as finding an optimal transmission strategy &#x03C0; = {&#x03C0;<sub><italic>i</italic>,<italic>j</italic></sub>|<italic>i</italic> = 1, 2, &#x2026;, <italic>m</italic>, <italic>j</italic> = 1, 2, &#x2026;, <italic>n</italic>} to minimize the transportation cost, as shown in Formula (4). The formula behind s.t. is a condition that needs to be satisfied during optimization.</p>
<disp-formula id="S3.Ex1"><mml:math id="M4"><mml:mrow><mml:munder><mml:mo movablelimits="false">min</mml:mo><mml:mi mathvariant="normal">&#x03C0;</mml:mi></mml:munder><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S3.E4"><label>(4)</label><mml:math id="M5"><mml:mrow><mml:mi>s</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mrow><mml:mrow><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:msub><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msub><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S3.Ex2"><mml:math id="M6"><mml:mrow><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>m</mml:mi></mml:munderover><mml:msub><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:munderover><mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo><mml:mi>j</mml:mi><mml:mi>n</mml:mi></mml:munderover><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S3.Ex3"><mml:math id="M7"><mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x03C0;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>The key of the OTA is how to define the transportation cost. The original OTA method defines transportation cost between the ground truth box and the prediction box as the weighted summation of their regression loss and classification loss. Since there is only the category of subsidence basin in our study, the transporting cost was defined as the weighted summation of confidence loss and bounding box regression loss. The computation equation is as follow:</p>
<disp-formula id="S3.E5"><label>(5)</label><mml:math id="M8"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>G</mml:mi><mml:mi>i</mml:mi><mml:mi mathvariant="bold">obj</mml:mi></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03B1;</mml:mi><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>P</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>G</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>L</italic><sub><italic>obj</italic></sub> and <italic>L</italic><sub><italic>reg</italic></sub> represent the binary cross-entropy and bounding box regression loss; <inline-formula><mml:math id="INEQ14"><mml:msubsup><mml:mi>P</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="INEQ15"><mml:msubsup><mml:mi>G</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denote the confidence score of the model prediction box and ground truth box, respectively; <inline-formula><mml:math id="INEQ16"><mml:msubsup><mml:mi>P</mml:mi><mml:mi>j</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="INEQ17"><mml:msubsup><mml:mi>G</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, respectively, represent the coordinates of the prediction box and the ground truth box. &#x03B1; is the balanced coefficient with the value of 3. The visualization result of transporting cost value is shown in <xref ref-type="fig" rid="F8">Figure 8D</xref>.</p>
<p>The original OTA method uses Sinkhorn-Knopp Iteration (<xref ref-type="bibr" rid="B10">Cuturi, 2013</xref>) to solve the optimal transmission problem, but the multiple iteration optimization of the Sinkhorn-Knopp greatly reduces the training efficiency of the model. Hence, the Light YOLO-Basin adopted an approximate method to solve the optimal transmission problem. The OTA solution mainly consists of two parts: (1) <italic>k</italic> value, which is the number of positive samples of each ground truth box; (2) the assignment of the <italic>k</italic> value, that is, how the positive samples are assigned to the prediction box. Since the cost matrix has been determined, the assignment solution of <italic>k</italic> can be simplified to assign <italic>k</italic> of each ground truth box to the prediction box with the lowest cost in turn, as shown in the blue grid in <xref ref-type="fig" rid="F8">Figure 8D</xref>. The determination of the value of <italic>k</italic> for each ground truth box is simplified to a statistical problem: first, sorting all the prediction boxes in the model according to the IoU value, and then adding up the cost value of the prediction box with the top Z (default 10) IoU value to round to get the estimated value of k, as shown in Formula (6).</p>
<disp-formula id="S3.E6"><label>(6)</label><mml:math id="M9"><mml:mrow><mml:mrow><mml:msub><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>o</mml:mi><mml:mi>p</mml:mi><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mmultiscripts><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:none/><mml:mprescripts/><mml:none/><mml:mo>&#x002A;</mml:mo></mml:mmultiscripts></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mrow><mml:mo rspace="5.8pt">,</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x22EF;</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <italic>k<sub>i</sub></italic> means the <italic>k</italic> value of the i-th ground truth box; &#x03BB; is set to 0.1 derived from experiments.</p>
</sec>
</sec>
<sec id="S4" sec-type="results">
<title>Results and Analysis</title>
<sec id="S4.SS1">
<title>Experimental Setting</title>
<sec id="S4.SS1.SSS1">
<title>Accuracy Evaluation Index</title>
<p>The average precision (AP) and mAP (<xref ref-type="bibr" rid="B30">Li et al., 2017</xref>; <xref ref-type="bibr" rid="B47">Sun et al., 2021</xref>) were used to evaluate the performance of the Light YOLO-Basin model proposed in this study. The detection results mainly include four categories: true-positive (TP) and false-positive (FP)are the numbers of positive samples that are correctly predicted and incorrectly predicted, respectively; true-negative (TN) and false-negative (FN) mean the number of negative samples that are correctly predicted and incorrectly predicted, respectively. Precision refers to the proportion of all detected samples that are correct, and recall refers to the proportion of the objects recognized by the model among all the objects that require to be recognized. The precision (P) and recall (R) are defined in Formula (7). The P-R curve takes the precision as the ordinate and the recall as the abscissa.</p>
<disp-formula id="S4.Ex4"><mml:math id="M10"><mml:mrow><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<disp-formula id="S4.E7"><label>(7)</label><mml:math id="M11"><mml:mrow><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>The AP measures the detection performance of a single category, as shown in Formula (8). The evaluation indicators of the Light YOLO-Basin model proposed in this paper are AP<sub>50</sub> : AP<sub>95</sub>. AP<sub>50</sub> : AP<sub>95</sub> is the value of AP at different IoU thresholds that the IoU ranges from 0.5 to 0.95 and the step size is 0.05. The mAP measures the detection performance of all the categories. Since there is only the category of subsidence basin in this paper, mAP is defined as the average AP at different IoU thresholds, as shown in Formula (9).</p>
<disp-formula id="S4.E8"><label>(8)</label><mml:math id="M12"><mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x222B;</mml:mo><mml:mn>0</mml:mn><mml:mn>1</mml:mn></mml:msubsup><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>R</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo mathvariant="italic" rspace="0pt">d</mml:mo><mml:mi>R</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:mrow></mml:math></disp-formula>
<disp-formula id="S4.E9"><label>(9)</label><mml:math id="M13"><mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mn>9</mml:mn></mml:msubsup><mml:mrow><mml:mi>A</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mn>50</mml:mn><mml:mo>+</mml:mo><mml:mi>i</mml:mi><mml:mmultiscripts><mml:mn>5</mml:mn><mml:mprescripts/><mml:none/><mml:mo>&#x002A;</mml:mo></mml:mmultiscripts></mml:mrow></mml:msub></mml:mrow></mml:mrow><mml:mn>10</mml:mn></mml:mfrac></mml:mrow></mml:math></disp-formula>
</sec>
<sec id="S4.SS1.SSS2">
<title>Training Settings</title>
<p>The experiment was conducted using the Microsoft Windows, 64-bit operating system. The central processing unit (CPU) is Intel Core i5-8300H. The graphics processing unit (GPU) is NVIDIA GeForce GTX 1050 Ti (4GB video memory). The deep learning framework is Facebook PyTorch 1.8. In this study, all subsidence basin detection models in this paper were trained by the adaptive moment estimation (Adam) optimization method (<xref ref-type="bibr" rid="B27">Kingma and Ba, 2014</xref>). The initial learning rate is set to 0.001 and decayed according to the formula 0.<inline-formula><mml:math id="INEQ18"><mml:mrow><mml:mn>001</mml:mn><mml:mmultiscripts><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mo>/</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mn>0.9</mml:mn></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mprescripts/><mml:none/><mml:mo>&#x002A;</mml:mo></mml:mmultiscripts></mml:mrow></mml:math></inline-formula>, where <italic>iter</italic> is the current number of iterations and the maximum number of iterations (<italic>max</italic><sub><italic>iter</italic></sub>) is set to 60,000.</p>
</sec>
</sec>
<sec id="S4.SS2" sec-type="results">
<title>Results</title>
<p>According to the standard of the YOLO V5 model, we divided the Light YOLO-Basin into three versions: Large (L), Middle (M), and Small (S). The three versions are distinguished by three scaling ratios of the depth (the number of convolutional layers) and the width (the number of channels of the convolution kernel), which are (1.0, 1.0), (0.67, 0.75), and (0.33, 0.50), respectively. The model that does not introduce depthwise separable convolution in this paper is called YOLO-Basin. Examples of the qualitative comparison results of the Light YOLO-Basin model and the YOLO V5 model are shown in <xref ref-type="fig" rid="F9">Figure 9</xref>. It can be seen that the Light YOLO-Basin model can detect the subsidence basin misdetected by the YOLO V5 model. Importantly, in order to verify the performance of the Light YOLO-Basin model in the actual scene, we selected two representative InSAR interferograms with the size of 7,636 &#x00D7; 8,205 and 8,127 &#x00D7; 10,338, respectively. We first cut the whole images regularly to obtain a large number of sub-images with a size of 416 &#x00D7; 416, then used the Light YOLO-Basin model to detect the subsidence basins for each sub-image, and finally stitched the detection results of each sub-image. <xref ref-type="fig" rid="F10">Figure 10</xref> exhibits part of the detection results of subsidence basins using the Light YOLO-Basin model. It can be seen that most subsidence basins have been correctly detected. 42 and 40 subsidence basins were detected in Jining and Huaibei mining areas, respectively. Regardless of the time consumption on image segmenting and stitching, the Light YOLO-Basin model only consumed 16.28 s to detect the whole image. We made statistics on the deformation values of the detected subsidence basins from the Jining and Huaibei mining areas, among which, the deformation value of the subsidence basin with the smallest deformation is 1.5 cm.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption><p>Results of qualitative comparison experiment. Large, Middle, and Small mean the three models produced by scaling, respectively. <bold>(A)</bold> Cropped original image; <bold>(B)</bold> detection results of subsidence basins using YOLO V5 model; <bold>(C)</bold> detection results of subsidence basins using YOLO-Basin model; <bold>(D)</bold> detection results of subsidence basins using Light YOLO-Basin. <bold>(E)</bold> Ground truth box.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g009.tif"/>
</fig>
<fig id="F10" position="float">
<label>FIGURE 10</label>
<caption><p>Detection results of partial subsidence basins in the Interferometric synthetic aperture radar (InSAR) interferogram with wide swath. We select two differential interferograms obtained from the experimental data in <xref ref-type="table" rid="T1">Table 1</xref> to test the performance of the Light YOLO-Basin model. Panels <bold>(A,B)</bold> are representative areas selected from the detection results of the Jining mining area; panels <bold>(C,D)</bold> are representative areas selected from the detection results of the Huaibei mining area.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g010.tif"/>
</fig>
</sec>
<sec id="S4.SS3">
<title>Quantitative Evaluation</title>
<p>The detection accuracy of the Light YOLO-Basin model and the YOLO V5 model are shown in <xref ref-type="table" rid="T2">Table 2</xref>. The experimental results demonstrate: (1) with the same experimental method, the smaller the scaling ratio of the model is, the higher the detection accuracy is. The value of mAP increased from 45.92% of the YOLO V5-L model to 47.65% of the YOLO V5-S model. This further verified the hypothesis proposed above that the detection of subsidence basins does not require a heavy network. (2) Benefiting from the introduction of the anchor-free detection box encoding method and OTA, the mAP value of the YOLO-Basin-S model has greatly improved compared to the YOLO V5-L model which was 6% higher than that of the original YOLO V5-L model. By comparing the detection accuracy in <xref ref-type="table" rid="T2">Table 2</xref>, it can be found that the improvement of the YOLO-Basin model detection accuracy is mainly manifested by the strict evaluation indicators such as AP<sub>70</sub> and AP<sub>80</sub>. The AP<sub>70</sub> value increased from 65.70 to 71.24%, and the AP<sub>80</sub> value increased by 15.20%. (3) The detection accuracy of the Light YOLO-Basin model is further improved, benefiting from the introduction of the depthwise separable convolution. Similarly, the improvement of the Light YOLO-Basin model detection accuracy is also mainly manifested by the evaluation indicators such as AP<sub>70</sub> and AP<sub>80</sub>. The AP<sub>70</sub> value improved from 71.24 to 75.20%, and the AP<sub>80</sub> value increased by 7.85% compared to the non-lightweight model. In summary, on the one hand, the experimental results prove the effectiveness of anchor-free and OTA methods in detecting the subsidence basins. On the other hand, depthwise separable convolution can improve the detection accuracy of the Light YOLO-Basin model with less model parameters.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Quantitative experiment comparison between the Light YOLO-Basin model and the YOLO V5 model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">Backbone</td>
<td valign="top" align="center">mAP (%)</td>
<td valign="top" align="center">AP<sub>50</sub> (%)</td>
<td valign="top" align="center">AP<sub>60</sub> (%)</td>
<td valign="top" align="center">AP<sub>70</sub> (%)</td>
<td valign="top" align="center">AP<sub>80</sub> (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">YOLO V5-L</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">45.92</td>
<td valign="top" align="center">88.33</td>
<td valign="top" align="center">80.76</td>
<td valign="top" align="center">61.91</td>
<td valign="top" align="center">21.94</td>
</tr>
<tr>
<td valign="top" align="left">YOLO V5-M</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">46.82</td>
<td valign="top" align="center">86.57</td>
<td valign="top" align="center">78.43</td>
<td valign="top" align="center">63.15</td>
<td valign="top" align="center">28.31</td>
</tr>
<tr>
<td valign="top" align="left">YOLO V5-S</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">47.65</td>
<td valign="top" align="center">89.78</td>
<td valign="top" align="center">84.36</td>
<td valign="top" align="center">65.70</td>
<td valign="top" align="center">20.12</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-Basin-L</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">49.96</td>
<td valign="top" align="center">87.45</td>
<td valign="top" align="center">82.79</td>
<td valign="top" align="center">69.21</td>
<td valign="top" align="center">30.89</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-Basin-M</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">51.77</td>
<td valign="top" align="center">89.08</td>
<td valign="top" align="center">84.23</td>
<td valign="top" align="center">69.18</td>
<td valign="top" align="center">36.21</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-Basin-S</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">51.92</td>
<td valign="top" align="center">90.03</td>
<td valign="top" align="center">85.12</td>
<td valign="top" align="center">71.24</td>
<td valign="top" align="center">35.32</td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-L</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center">53.46</td>
<td valign="top" align="center">88.95</td>
<td valign="top" align="center">84.02</td>
<td valign="top" align="center">71.68</td>
<td valign="top" align="center">40.53</td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-M</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center">54.61</td>
<td valign="top" align="center">90.37</td>
<td valign="top" align="center"><bold>86.48</bold></td>
<td valign="top" align="center"><bold>75.29</bold></td>
<td valign="top" align="center">41.96</td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-S</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center"><bold>55.12</bold></td>
<td valign="top" align="center"><bold>90.64</bold></td>
<td valign="top" align="center">86.21</td>
<td valign="top" align="center">75.20</td>
<td valign="top" align="center"><bold>43.17</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>Light CSPDarknet-53 represents the CSPDarknet-53 backbone network after introducing the depthwise separable convolution. L, M, and S represent the three versions of YOLO V5, YOLO-Basin, and Light YOLO-Basin model Large, Middle, and Small, respectively. The bold values is the maximum value of each column.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="S5" sec-type="discussion">
<title>Discussion</title>
<sec id="S5.SS1">
<title>Efficiency Experiment of the Lightweight Module</title>
<p>The lightweight module of the Light YOLO-Bain model mainly includes scaling and depthwise separable convolution. The lightweight detection model needs to pay attention to two aspects: (1) whether the lightweight module can improve model computing efficiency and reduce memory utilization; (2) whether the lightweight module affects detection accuracy. We used network parameters, GFLOPs, and inference memory as the evaluation indicators of model efficiency and mAP as the evaluation indicator of model accuracy. The results are shown in <xref ref-type="table" rid="T3">Table 3</xref>. The image size is set to 416 &#x00D7; 416 and the batch size is set to 1 when training the model.</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>The experiment for test the accuracy and efficiency of lightweight module.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">Backbone</td>
<td valign="top" align="center">mAP (%)</td>
<td valign="top" align="center">Parameters (MB)</td>
<td valign="top" align="center">GFLOPs</td>
<td valign="top" align="center">Inference memory (MB)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">YOLO-Basin-L</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">49.96</td>
<td valign="top" align="center">207.10</td>
<td valign="top" align="center">32.81</td>
<td valign="top" align="center">434.96</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-Basin-M</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">51.77</td>
<td valign="top" align="center">108.99</td>
<td valign="top" align="center">19.25</td>
<td valign="top" align="center">285.76</td>
</tr>
<tr>
<td valign="top" align="left">YOLO-Basin-S</td>
<td valign="top" align="center">CSPDarknet-53</td>
<td valign="top" align="center">51.92</td>
<td valign="top" align="center">55.08</td>
<td valign="top" align="center">11.99</td>
<td valign="top" align="center"><bold>172.22</bold></td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-L</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center">53.46</td>
<td valign="top" align="center">92.09</td>
<td valign="top" align="center">16.63</td>
<td valign="top" align="center">559.73</td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-M</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center">54.61</td>
<td valign="top" align="center">60.12</td>
<td valign="top" align="center">12.54</td>
<td valign="top" align="center">352.60</td>
</tr>
<tr>
<td valign="top" align="left">Light YOLO-Basin-S</td>
<td valign="top" align="center">Light CSPDarknet-53</td>
<td valign="top" align="center"><bold>55.12</bold></td>
<td valign="top" align="center"><bold>40.39</bold></td>
<td valign="top" align="center"><bold>10.07</bold></td>
<td valign="top" align="center">198.95</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The bold values is the maximum value of each column.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>Firstly, by analyzing the values of the evaluation indicators of model efficiency in <xref ref-type="table" rid="T3">Table 3</xref>, introducing depthwise separable convolution and scaling can exponentially decrease the number of model parameters and speed up model training. The improvement of model efficiency by depthwise separable convolution is mainly reflected in the number of model parameters and detection speed. Note that the smaller the model is, the smaller the effect of the depthwise separable convolution is. In addition, since the depthwise separable convolution factorizes a standard convolution into two parts, the computation memory utilization of the model increases. <xref ref-type="table" rid="T3">Table 3</xref> indicates the use of scaling and depthwise separable convolution has a better lightweight effect on the model. Secondly, observing the model accuracy evaluation indicators data in <xref ref-type="table" rid="T3">Table 3</xref>, compared with the detection method of natural images, the lightweight of the model can improve the accuracy of the detection of subsidence basins rather than reducing the accuracy. Compared with the YOLO-Basin-S model, the Light YOLO-Basin-S model increases the mAP value from 51.92 to 55.12%.</p>
</sec>
<sec id="S5.SS2">
<title>Ablation Study for Detection Head and Loss Function</title>
<p>To address the problem caused by the anchor boxes, the Light YOLO-Basin model introduces the anchor-free method and the OTA method. In addition, we also changed the neck network and loss function of the YOLO V5 model, as seen in Section &#x201C;YOLO V5 Network.&#x201D; We used mAP and AP<sub>50</sub> : AP<sub>95</sub> as the evaluation indicators of model accuracy in <xref ref-type="table" rid="T4">Table 4</xref>, showing the accuracy changes of the modules added to the YOLO V5-S model.</p>
<table-wrap position="float" id="T4">
<label>TABLE 4</label>
<caption><p>Roadmap of the Light YOLO-Basin model in terms of mAP and average precision (AP) (%).</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td/>
<td valign="top" align="center">mAP (%)</td>
<td valign="top" align="center">AP<sub>50</sub> (%)</td>
<td valign="top" align="center">AP<sub>60</sub> (%)</td>
<td valign="top" align="center">AP<sub>70</sub> (%)</td>
<td valign="top" align="center">AP<sub>80</sub> (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">YOLOV5-S</td>
<td valign="top" align="center">47.65</td>
<td valign="top" align="center">89.78</td>
<td valign="top" align="center">84.36</td>
<td valign="top" align="center">65.70</td>
<td valign="top" align="center">20.12</td>
</tr>
<tr>
<td valign="top" align="left">+Depthwise</td>
<td valign="top" align="center">46.98</td>
<td valign="top" align="center">86.77</td>
<td valign="top" align="center">79.19</td>
<td valign="top" align="center">66.01</td>
<td valign="top" align="center">24.34</td>
</tr>
<tr>
<td valign="top" align="left">+Anchor-free and OTA</td>
<td valign="top" align="center">54.84</td>
<td valign="top" align="center"><bold>91.52</bold></td>
<td valign="top" align="center"><bold>86.82</bold></td>
<td valign="top" align="center"><bold>75.85</bold></td>
<td valign="top" align="center">40.88</td>
</tr>
<tr>
<td valign="top" align="left">+Double-head</td>
<td valign="top" align="center">54.62</td>
<td valign="top" align="center">88.62</td>
<td valign="top" align="center">86.17</td>
<td valign="top" align="center">75.48</td>
<td valign="top" align="center">42.75</td>
</tr>
<tr>
<td valign="top" align="left">+Focal Loss</td>
<td valign="top" align="center"><bold>55.12</bold></td>
<td valign="top" align="center">90.64</td>
<td valign="top" align="center">86.21</td>
<td valign="top" align="center">75.20</td>
<td valign="top" align="center"><bold>43.17</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The bold values is the maximum value of each column.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>It visually shows the change of model accuracy with added different modules in <xref ref-type="table" rid="T4">Table 4</xref>. By analyzing the changes in model accuracy, we can draw the following three conclusions: (1) The anchor-free detection box encoding method and OTA have the greatest effect on improving the detection accuracy of the model, greatly increasing the value of AP<sub>70</sub> and AP<sub>80</sub>. The accuracy evaluation indicators mAP has also been significantly improved, increasing from 47.65 to 55.12%. (2) The introduction of depthwise separable convolution did not improve the detection accuracy of the YOLO V5 model. However, it can improve the accuracy of the Light YOLO-Basin model, perhaps benefiting from the combined effects of depthwise separable convolution and OTA. (3) Compared with the OTA, the double-head and Focal loss only less improve the accuracy.</p>
<p>We also analyzed the influence of different IoU loss functions on the Light YOLO-Bain model accuracy, as shown in <xref ref-type="table" rid="T5">Table 5</xref>. It can be seen from <xref ref-type="table" rid="T5">Table 5</xref> that the detection accuracy is the highest when we used the DIoU loss function. Hence, we used DIoU loss as the bounding box regression loss function.</p>
<table-wrap position="float" id="T5">
<label>TABLE 5</label>
<caption><p>Ablation study for loss function.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Loss</td>
<td valign="top" align="center">mAP (%)</td>
<td valign="top" align="center">AP<sub>50</sub> (%)</td>
<td valign="top" align="center">AP<sub>60</sub> (%)</td>
<td valign="top" align="center">AP<sub>70</sub> (%)</td>
<td valign="top" align="center">AP<sub>80</sub> (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">IoU loss</td>
<td valign="top" align="center">54.91</td>
<td valign="top" align="center">90.28</td>
<td valign="top" align="center">87.40</td>
<td valign="top" align="center">74.70</td>
<td valign="top" align="center">41.41</td>
</tr>
<tr>
<td valign="top" align="left">GIoU loss</td>
<td valign="top" align="center">54.08</td>
<td valign="top" align="center">89.67</td>
<td valign="top" align="center">85.73</td>
<td valign="top" align="center">74.89</td>
<td valign="top" align="center">42.01</td>
</tr>
<tr>
<td valign="top" align="left">DIoU loss</td>
<td valign="top" align="center"><bold>55.12</bold></td>
<td valign="top" align="center">90.64</td>
<td valign="top" align="center">86.21</td>
<td valign="top" align="center"><bold>75.20</bold></td>
<td valign="top" align="center"><bold>43.17</bold></td>
</tr>
<tr>
<td valign="top" align="left">CIoU loss</td>
<td valign="top" align="center">54.89</td>
<td valign="top" align="center"><bold>91.27</bold></td>
<td valign="top" align="center"><bold>87.86</bold></td>
<td valign="top" align="center">75.05</td>
<td valign="top" align="center">40.98</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The bold values is the maximum value of each column.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S5.SS3">
<title>Ablation Study for Optimal Transport Assignment</title>
<p>To avoid the inefficient iterative computation of Sinkhorn-Knopp Iteration, we used statistical methods to estimate the <italic>k</italic> value corresponding to the ground truth box in the Light YOLO-Bain model, each of which represents the number of corresponding positive samples. The <italic>k</italic> value is obtained by adding up the cost value of the prediction box of the top Z in the IoU value and rounding it. Hence, the number of Z determines the size of the <italic>k</italic> value to a certain extent. According to Formula (6), the larger the value of Z is, the larger the corresponding <italic>k</italic> value is. That is, each ground truth box is assigned more positive samples. However, too many positive samples will divide the poorly optimized prediction boxes into positive samples, resulting in incorrect detection of the Light YOLO-Bain model. Too few positive samples will cause the imbalance of positive and negative samples, increasing the difficulty of model optimization. Therefore, it is important to choose a suitable Z. <xref ref-type="table" rid="T6">Table 6</xref> shows the effect of different values of Z on the accuracy of the Light YOLO-Basin model. It can be seen that the detection accuracy of the model is highest when the value of Z is 10.</p>
<table-wrap position="float" id="T6">
<label>TABLE 6</label>
<caption><p>Ablation study of Z in optimal transport assignment (OTA).</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Z</td>
<td valign="top" align="center">mAP (%)</td>
<td valign="top" align="center">AP<sub>50</sub> (%)</td>
<td valign="top" align="center">AP<sub>60</sub> (%)</td>
<td valign="top" align="center">AP<sub>70</sub> (%)</td>
<td valign="top" align="center">AP<sub>80</sub> (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">54.12</td>
<td valign="top" align="center">90.33</td>
<td valign="top" align="center"><bold>86.81</bold></td>
<td valign="top" align="center">74.62</td>
<td valign="top" align="center">38.63</td>
</tr>
<tr>
<td valign="top" align="left">10</td>
<td valign="top" align="center"><bold>55.12</bold></td>
<td valign="top" align="center">90.64</td>
<td valign="top" align="center">86.21</td>
<td valign="top" align="center">75.20</td>
<td valign="top" align="center"><bold>43.17</bold></td>
</tr>
<tr>
<td valign="top" align="left">15</td>
<td valign="top" align="center">54.47</td>
<td valign="top" align="center"><bold>90.70</bold></td>
<td valign="top" align="center">88.43</td>
<td valign="top" align="center"><bold>75.91</bold></td>
<td valign="top" align="center">37.43</td>
</tr>
<tr>
<td valign="top" align="left">20</td>
<td valign="top" align="center">54.15</td>
<td valign="top" align="center">90.01</td>
<td valign="top" align="center">85.15</td>
<td valign="top" align="center">74.52</td>
<td valign="top" align="center">42.31</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The bold values is the maximum value of each column.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S5.SS4">
<title>Robustness of Light YOLO-Basin Model</title>
<p>To test the robustness of the model, we, respectively, tested the effects of DEM, decorrelation noise, and the number of interferogram multi looking on detection. We first conducted a set of comparative experiments using DEM with a spatial resolution of 30 and 90 m. The detection results of subsidence basins using different levels of DEM are shown in <xref ref-type="fig" rid="F11">Figure 11</xref>. Since the DEM does not affect the features of the subsidence basins, it has little influence on the detection results where 14 and 13 subsidence basins were detected using DEM with a spatial resolution of 30 and 90 m, respectively. Note that the spatial resolution of DEM used in this paper is 30 m. Then, to test the effect of decorrelation noise on detection, the Gaussian complex noise was used to simulate the decorrelation noise. <xref ref-type="fig" rid="F12">Figure 12</xref> gives the influence of noise on the detection results. It can be seen that noise has a greater influence on the detection results. Hence, we used the Goldstein filtering method to remove image noise in this paper. Finally, we conducted a set of comparative experiments with different numbers of multi-looking. The comparison of the detection results is shown in <xref ref-type="fig" rid="F13">Figure 13</xref>. It can be seen that the number of interferogram multi-looking changes the aspect ratio of the subsidence basin but has little effect on the detection result of the subsidence basin. We adopted a more appropriate ratio of multi looking is 1: 5.</p>
<fig id="F11" position="float">
<label>FIGURE 11</label>
<caption><p>The detection results using different spatial resolutions of DEM. Panel <bold>(A)</bold> is the result using DEM data with a spatial resolution of 30 m; panel <bold>(B)</bold> is the result using DEM data with a spatial resolution of 90 m.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g011.tif"/>
</fig>
<fig id="F12" position="float">
<label>FIGURE 12</label>
<caption><p>The influence of noise on the detection results. Panel <bold>(A)</bold> is the detection result of this paper to remove noise; panels <bold>(B,C)</bold> are the detection results under Gaussian complex noise with the variance of 0.001 and 0.01, respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g012.tif"/>
</fig>
<fig id="F13" position="float">
<label>FIGURE 13</label>
<caption><p>Detection results of different multi looking numbers of the interferogram. Panels <bold>(A&#x2013;D)</bold> are the detection results with multi looking ratios of 1:5, 1:4, 1:3, and 1:2, respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fevo-10-840464-g013.tif"/>
</fig>
</sec>
</sec>
<sec id="S6" sec-type="conclusion">
<title>Conclusion</title>
<p>Based on the YOLO V5 network architecture, the Light YOLO-Basin model for automatically detecting subsidence basins from interferograms was proposed in this paper. The Light YOLO-Basin model uses depthwise separable convolution as the lightweight module and introduces the anchor-free detection box encoding method and OTA to solve the problem caused by fixed anchor boxes. Through experiment, the valuable conclusions can be obtained as follows:</p>
<list list-type="simple">
<list-item>
<label>1.</label>
<p>Depthwise separable convolution generally sacrifices a small amount of accuracy to improve the detection efficiency. The depthwise separable convolution of the Light YOLO-Basin model can improve the detection speed and reduce the model parameters from 207.10 to 40.39 MB. More importantly, the detection accuracy of the Light YOLO-Basin model has also been significantly improved. The value of mAP is increased by 3.2% compared with the non-lightweight model, which verifies the assumption that it does not require a heavy network when detecting the subsidence basins from interferograms.</p>
</list-item>
<list-item>
<label>2.</label>
<p>It can effectively detect the subsidence basins through combined anchor-free and OTA adaptive sample assignment methods. The ablation experiments in this study indicate that anchor-free and OTA methods in the Light YOLO-Basin model increase the value of mAP from 46.98 to 54.84%, and the value of AP<sub>50</sub>, AP<sub>60</sub>, and AP<sub>70</sub> increase by 4.75, 7.63, and 9.84%, respectively.</p>
</list-item>
<list-item>
<label>3.</label>
<p>We introduce the Focal Loss function in the Light YOLO-Basin model when computing the confidence loss to balance the weight of the hard and easy samples during model training, increasing the value of mAP from 54.62 to 55.12% and AP<sub>50</sub> from 88.62 to 90.64%.</p>
</list-item>
</list>
<p>The Light YOLO-Basin model proposed in this paper has good performance to detect subsidence basins from InSAR interferograms with wide swaths. This study also has some limitations. When making sample datasets, the labeling of training samples has a greater impact on the detection accuracy of the model. For subsidence basins with poor visual interpretation, the Light YOLO-Basin model also has false detection or missing detection in the poor interferogram. In addition, the reason why the depthwise separable convolution improves the detection accuracy of the subsidence basin may be related to the shape of the subsidence basin on the InSAR interferogram. We will solve these above problems in future works. We will propose a better detection model by analyzing the difference in morphological characteristics between the subsidence basin on the InSAR interferogram and the object in the natural image.</p>
</sec>
<sec id="S7" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="S8">
<title>Author Contributions</title>
<p>YY collected and analyzed the data, and wrote the manuscript. ZYW proposed the method, designed its structure, and revised the manuscript. ZL and KY helped in collecting and analyzing the data. HL and ZHW critically revised the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S9" sec-type="funding-information">
<title>Funding</title>
<p>This research was funded by the Major Science and Technology Innovation Projects of Shandong Province (No. 2019JZZY020103). This research was supported by funding from the National Natural Science Foundation of China (No. 41876202).</p>
</sec>
<ack>
<p>We thank the European Space Agency (ESA) for providing the Sentinel-1A remote sensing data. We also thank NASA for providing the SRTM DEM data. We want to thank the reviewers for their valuable suggestions and comments which improved the quality of the manuscript.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bala</surname> <given-names>J.</given-names></name> <name><surname>Dwornik</surname> <given-names>M.</given-names></name> <name><surname>Franczyk</surname> <given-names>A.</given-names></name></person-group> (<year>2021</year>). <article-title>Automatic subsidence troughs detection in SAR interferograms using circlet transform.</article-title> <source><italic>Sensors</italic></source> <volume>21</volume>:<issue>1706</issue>. <pub-id pub-id-type="doi">10.3390/s21051706</pub-id> <pub-id pub-id-type="pmid">33801252</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bochkovskiy</surname> <given-names>A.</given-names></name> <name><surname>Wang</surname> <given-names>C. Y.</given-names></name> <name><surname>Liao</surname> <given-names>H. Y. M.</given-names></name></person-group> (<year>2020</year>). <article-title>Yolov4: optimal speed and accuracy of object detection.</article-title> <source><italic>arXiv</italic></source> <comment>[Preprint] arXiv: 2004.10934</comment>,</citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chang</surname> <given-names>Y. L.</given-names></name> <name><surname>Anagaw</surname> <given-names>A.</given-names></name> <name><surname>Chang</surname> <given-names>L. N.</given-names></name> <name><surname>Wang</surname> <given-names>Y. C.</given-names></name> <name><surname>Hsiao</surname> <given-names>C. Y.</given-names></name> <name><surname>Lee</surname> <given-names>W. H.</given-names></name></person-group> (<year>2019</year>). <article-title>Ship detection based on YOLOv2 for SAR imagery.</article-title> <source><italic>Remote Sens.</italic></source> <volume>11</volume>:<issue>786</issue>. <pub-id pub-id-type="doi">10.3390/rs11070786</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>B. Q.</given-names></name> <name><surname>Deng</surname> <given-names>K. Z.</given-names></name> <name><surname>Fan</surname> <given-names>H. D.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name></person-group> (<year>2015</year>). <article-title>Combining SAR interferometric phase and intensity information for monitoring of large gradient deformation in coal mining area.</article-title> <source><italic>Eur. J. Remote Sens.</italic></source> <volume>48</volume> <fpage>701</fpage>&#x2013;<lpage>717</lpage>. <pub-id pub-id-type="doi">10.5721/EuJRS20154839</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>B. Q.</given-names></name> <name><surname>Li</surname> <given-names>Z. H.</given-names></name> <name><surname>Yu</surname> <given-names>C.</given-names></name> <name><surname>Fairbairn</surname> <given-names>D.</given-names></name> <name><surname>Kang</surname> <given-names>J. R.</given-names></name> <name><surname>Hu</surname> <given-names>J. S.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Three-dimensional time-varying large surface displacements in coal exploiting areas revealed through integration of SAR pixel offset measurements and mining subsidence model.</article-title> <source><italic>Remote Sens. Environ.</italic></source> <volume>240</volume>:<issue>111663</issue>. <pub-id pub-id-type="doi">10.1016/j.rse.2020.111663</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>B. Q.</given-names></name> <name><surname>Mei</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Z. H.</given-names></name> <name><surname>Wang</surname> <given-names>Z. S.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Retrieving three-dimensional large surface displacements in coal mining areas by combining SAR pixel offset measurements with an improved mining subsidence model.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>2541</issue>. <pub-id pub-id-type="doi">10.3390/rs13132541</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>D. H.</given-names></name> <name><surname>Chen</surname> <given-names>H. E.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name> <name><surname>Cao</surname> <given-names>C.</given-names></name> <name><surname>Zhu</surname> <given-names>K. X.</given-names></name> <name><surname>Yuan</surname> <given-names>X. Q.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Characteristics of the residual surface deformation of multiple abandoned mined-out areas based on a field investigation and SBAS-InSAR: a case study in Jilin, China.</article-title> <source><italic>Remote Sens.</italic></source> <volume>12</volume> <issue>3752</issue>. <pub-id pub-id-type="doi">10.3390/rs12223752</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Q. S.</given-names></name> <name><surname>Fan</surname> <given-names>C.</given-names></name> <name><surname>Jin</surname> <given-names>W. Z.</given-names></name> <name><surname>Zou</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>F. Y.</given-names></name> <name><surname>Li</surname> <given-names>X. P.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>EPGNet: enhanced point cloud generation for 3D object detection.</article-title> <source><italic>Sensors</italic></source> <volume>20</volume>:<issue>6927</issue>. <pub-id pub-id-type="doi">10.3390/s20236927</pub-id> <pub-id pub-id-type="pmid">33291527</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Tong</surname> <given-names>Y. X.</given-names></name> <name><surname>Tan</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>Coal mining deformation monitoring using SBAS-InSAR and offset tracking: a case study of yu county, china.</article-title> <source><italic>IEEE J. Sel. Topics Appl. Earth Observ.</italic></source> <volume>13</volume> <fpage>6077</fpage>&#x2013;<lpage>6087</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2020.3028083</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cuturi</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Sinkhorn distances: lightspeed computation of optimal transport.</article-title> <source><italic>Adv. Neural Inf. Process. Syst.</italic></source> <volume>26</volume> <fpage>2292</fpage>&#x2013;<lpage>2300</lpage>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dai</surname> <given-names>Y. W.</given-names></name> <name><surname>Ng</surname> <given-names>A. H. M.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>L. Y.</given-names></name> <name><surname>Tao</surname> <given-names>T. Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Modeling-assisted InSAR phase-unwrapping method for mapping mine subsidence.</article-title> <source><italic>IEEE Geosci. Remote Sens. Lett.</italic></source> <volume>18</volume> <fpage>1059</fpage>&#x2013;<lpage>1063</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2020.2991687</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>L. K.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Tang</surname> <given-names>Y. X.</given-names></name> <name><surname>Tang</surname> <given-names>F. Q.</given-names></name> <name><surname>Duan</surname> <given-names>W.</given-names></name></person-group> (<year>2021</year>). <article-title>Time series InSAR three-dimensional displacement inversion model of coal mining areas based on symmetrical features of mining subsidence.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume> <fpage>2143</fpage>&#x2013;<lpage>2159</lpage>. <pub-id pub-id-type="doi">10.3390/rs13112143</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>Y. J.</given-names></name> <name><surname>Zheng</surname> <given-names>M. N.</given-names></name> <name><surname>Zhou</surname> <given-names>D. W.</given-names></name> <name><surname>Xia</surname> <given-names>Y. P.</given-names></name></person-group> (<year>2019</year>). <article-title>Goaf locating based on InSAR and probability integration method.</article-title> <source><italic>Remote Sens.</italic></source> <volume>11</volume>:<issue>8</issue>. <pub-id pub-id-type="doi">10.3390/rs11070812</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Duan</surname> <given-names>K. W.</given-names></name> <name><surname>Bai</surname> <given-names>S.</given-names></name> <name><surname>Xie</surname> <given-names>L. X.</given-names></name> <name><surname>Qi</surname> <given-names>H. G.</given-names></name> <name><surname>Huang</surname> <given-names>Q. M.</given-names></name> <name><surname>Tian</surname> <given-names>Q.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Centernet: keypoint triplets for object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2019 IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Seoul</publisher-loc>, <fpage>6569</fpage>&#x2013;<lpage>6578</lpage>.</citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>H. D.</given-names></name> <name><surname>Lu</surname> <given-names>L.</given-names></name> <name><surname>Yao</surname> <given-names>Y. H.</given-names></name></person-group> (<year>2018</year>). <article-title>Method combining probability integration model and a small baseline subset for time series monitoring of mining subsidence.</article-title> <source><italic>Remote Sens.</italic></source> <volume>10</volume>:<issue>1444</issue>. <pub-id pub-id-type="doi">10.3390/rs10091444</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>H. D.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Wen</surname> <given-names>B. F.</given-names></name> <name><surname>Du</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>A new model for three-dimensional deformation extraction with single-track InSAR based on mining subsidence characteristics.</article-title> <source><italic>Int. J. Appl. Earth Obs.</italic></source> <volume>94</volume>:<issue>102223</issue>. <pub-id pub-id-type="doi">10.1016/j.jag.2020.102223</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>J. M.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>Z. R.</given-names></name> <name><surname>Fu</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>An anchor-free method based on feature balancing and refinement network for multiscale ship detection in SAR images.</article-title> <source><italic>IEEE Trans. Geosci. Remote Sens.</italic></source> <volume>59</volume> <fpage>1331</fpage>&#x2013;<lpage>1344</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2020.3005151</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname> <given-names>K.</given-names></name> <name><surname>Chang</surname> <given-names>Z. H.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>G. L.</given-names></name> <name><surname>Zhang</surname> <given-names>K. S.</given-names></name> <name><surname>Sun</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Rotation-aware and multi-scale convolutional neural network for object detection in remote sensing images.</article-title> <source><italic>ISPRS J. Photogramm.</italic></source> <volume>161</volume> <fpage>294</fpage>&#x2013;<lpage>308</lpage>. <pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.01.025</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>S. T.</given-names></name> <name><surname>Li</surname> <given-names>Z. M.</given-names></name> <name><surname>Yoshie</surname> <given-names>O.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2021</year>). &#x201C;<article-title>OTA: optimal transport assignment for object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2021 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Nashville, TN</publisher-loc>, <fpage>303</fpage>&#x2013;<lpage>312</lpage>.</citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Girshick</surname> <given-names>R.</given-names></name></person-group> (<year>2015</year>). &#x201C;<article-title>Fast R-CNN</article-title>,&#x201D; in <source><italic>Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Santiago</publisher-loc>, <fpage>1440</fpage>&#x2013;<lpage>1448</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.169</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Donahue</surname> <given-names>J.</given-names></name> <name><surname>Darrell</surname> <given-names>T.</given-names></name> <name><surname>Malik</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). &#x201C;<article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>,&#x201D; in <source><italic>Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Columbus, OH</publisher-loc>, <fpage>580</fpage>&#x2013;<lpage>587</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2014.81</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; in <source><italic>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognitionn (CVPR)</italic></source>, <publisher-loc>Las Vegas, NV</publisher-loc>, <fpage>770</fpage>&#x2013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Howard</surname> <given-names>A. G.</given-names></name> <name><surname>Zhu</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Kalenichenko</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Weyand</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2017</year>). <article-title>Mobilenets: efficient convolutional neural networks for mobile vision applications.</article-title> <source><italic>arXiv</italic></source> <comment>[Preprint] arXiv: 1704.04861</comment>,</citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>Z.</given-names></name> <name><surname>Ge</surname> <given-names>L. L.</given-names></name> <name><surname>Li</surname> <given-names>X. J.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2013</year>). <article-title>An underground-mining detection system based on DInSAR.</article-title> <source><italic>IEEE Trans. Geosci. Remote Sens.</italic></source> <volume>51</volume> <fpage>615</fpage>&#x2013;<lpage>625</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2012.2202243</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ilieva</surname> <given-names>M.</given-names></name> <name><surname>Polanin</surname> <given-names>P.</given-names></name> <name><surname>Borkowski</surname> <given-names>A.</given-names></name> <name><surname>Gruchlik</surname> <given-names>P.</given-names></name> <name><surname>Smolak</surname> <given-names>K.</given-names></name> <name><surname>Kowalski</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2019</year>). <article-title>Mining deformation life cycle in the light of InSAR and deformation models.</article-title> <source><italic>Remote Sens.</italic></source> <volume>11</volume>:<issue>745</issue>. <pub-id pub-id-type="doi">10.3390/rs11070745</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>J. H.</given-names></name> <name><surname>Fu</surname> <given-names>X. J.</given-names></name> <name><surname>Qin</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>X. Y.</given-names></name> <name><surname>Ma</surname> <given-names>Z. F.</given-names></name></person-group> (<year>2021</year>). <article-title>High-speed lightweight ship detection algorithm based on YOLO-V4 for three-channels RGB SAR image.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>1909</issue>. <pub-id pub-id-type="doi">10.3390/rs13101909</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization.</article-title> <source><italic>arXiv</italic></source> <comment>[Preprint] arXiv: 1412.6980</comment>,</citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Law</surname> <given-names>H.</given-names></name> <name><surname>Deng</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Cornernet: detecting objects as paired keypoints</article-title>,&#x201D; in <source><italic>Proceedings of the 2018 European Conference on Computer Vision (ECCV)</italic></source>, <publisher-loc>Munich</publisher-loc>, <fpage>734</fpage>&#x2013;<lpage>750</lpage>. <pub-id pub-id-type="doi">10.1007/s11263-019-01204-1</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>LeCun</surname> <given-names>Y.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name> <name><surname>Hinton</surname> <given-names>G.</given-names></name></person-group> (<year>2015</year>). <article-title>Deep learning.</article-title> <source><italic>Nature</italic></source> <volume>2015</volume> <fpage>436</fpage>&#x2013;<lpage>444</lpage>. <pub-id pub-id-type="doi">10.1038/nature14539</pub-id> <pub-id pub-id-type="pmid">26017442</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>K.</given-names></name> <name><surname>Cheng</surname> <given-names>G.</given-names></name> <name><surname>Bu</surname> <given-names>S. H.</given-names></name> <name><surname>You</surname> <given-names>X.</given-names></name></person-group> (<year>2017</year>). <article-title>Rotation-insensitive and context-augmented object detection in remote sensing images.</article-title> <source><italic>IEEE Trans. Geosci. Remote Sens.</italic></source> <volume>56</volume> <fpage>2337</fpage>&#x2013;<lpage>2348</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2017.2778300</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>T. Y.</given-names></name> <name><surname>Doll&#x00E1;r</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>He</surname> <given-names>K. M.</given-names></name> <name><surname>Hariharan</surname> <given-names>B.</given-names></name> <name><surname>Belongie</surname> <given-names>S.</given-names></name></person-group> (<year>2017a</year>). &#x201C;<article-title>Feature pyramid networks for object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Honolulu, HI</publisher-loc>, <fpage>2117</fpage>&#x2013;<lpage>2125</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.106</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>T. Y.</given-names></name> <name><surname>Goyal</surname> <given-names>P.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>He</surname> <given-names>K. M.</given-names></name> <name><surname>Dollar</surname> <given-names>P.</given-names></name></person-group> (<year>2017b</year>). &#x201C;<article-title>Focal loss for dense object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Venice</publisher-loc>, <fpage>2980</fpage>&#x2013;<lpage>2988</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2018.2858826</pub-id> <pub-id pub-id-type="pmid">30040631</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Anguelov</surname> <given-names>D.</given-names></name> <name><surname>Erhan</surname> <given-names>D.</given-names></name> <name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Reed</surname> <given-names>S.</given-names></name> <name><surname>Fu</surname> <given-names>C. Y.</given-names></name><etal/></person-group> (<year>2016</year>). &#x201C;<article-title>SSD: single shot multibox detector</article-title>,&#x201D; in <source><italic>Proceedings of the 2016 European Conference on Computer Vision (ECCV)</italic></source>, <publisher-loc>Amsterdam</publisher-loc>, <fpage>21</fpage>&#x2013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-46448-0_2</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>Z.</given-names></name> <name><surname>Yu</surname> <given-names>H. L.</given-names></name> <name><surname>Zhang</surname> <given-names>Y. Z.</given-names></name></person-group> (<year>2020</year>). <article-title>Pine cone detection using boundary equilibrium generative adversarial networks and improved YOLOv3 model.</article-title> <source><italic>Sensors</italic></source> <volume>20</volume>:<issue>4430</issue>. <pub-id pub-id-type="doi">10.3390/s20164430</pub-id> <pub-id pub-id-type="pmid">32784403</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ng</surname> <given-names>A. H. M.</given-names></name> <name><surname>Ge</surname> <given-names>L. L.</given-names></name> <name><surname>Du</surname> <given-names>Z. Y.</given-names></name> <name><surname>Wang</surname> <given-names>S. R.</given-names></name> <name><surname>Ma</surname> <given-names>C.</given-names></name></person-group> (<year>2017</year>). <article-title>Satellite radar interferometry for monitoring subsidence induced by longwall mining activity using Radarsat-2, Sentinel-1 and ALOS-2 data.</article-title> <source><italic>Int. J. Appl. Earth Obs.</italic></source> <volume>61</volume> <fpage>92</fpage>&#x2013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1016/j.jag.2017.05.009</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ou</surname> <given-names>D. P.</given-names></name> <name><surname>Tan</surname> <given-names>K.</given-names></name> <name><surname>Du</surname> <given-names>Q.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Ding</surname> <given-names>J. W.</given-names></name></person-group> (<year>2018</year>). <article-title>Decision fusion of D-InSAR and pixel offset tracking for coal mining deformation monitoring.</article-title> <source><italic>Remote Sens.</italic></source> <volume>10</volume>:<issue>1055</issue>. <pub-id pub-id-type="doi">10.3390/rs10071055</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Purkait</surname> <given-names>P.</given-names></name> <name><surname>Zhao</surname> <given-names>C.</given-names></name> <name><surname>Zach</surname> <given-names>C.</given-names></name></person-group> (<year>2017</year>). <article-title>SPP-Net: deep absolute pose regression with synthetic views.</article-title> <source><italic>arXiv</italic></source> <comment>[Preprint]. arXiv: 1712.03452</comment>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Redmon</surname> <given-names>J.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>YOLO9000: better, faster, stronger</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Honolulu, HI</publisher-loc>, <fpage>7263</fpage>&#x2013;<lpage>7271</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.690</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Redmon</surname> <given-names>J.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Yolov3: an incremental improvement.</article-title> <source><italic>arXiv</italic></source> <comment>[Preprint] arXiv: 1804.02767</comment>,</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Redmon</surname> <given-names>J.</given-names></name> <name><surname>Divvala</surname> <given-names>S.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Farhadi</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>You only look once: unified, real-time object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Las Vegas, NV</publisher-loc>, <fpage>779</fpage>&#x2013;<lpage>788</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.91</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>C. J.</given-names></name> <name><surname>Jung</surname> <given-names>H. J.</given-names></name> <name><surname>Lee</surname> <given-names>S. H.</given-names></name> <name><surname>Jeong</surname> <given-names>D. W.</given-names></name></person-group> (<year>2021</year>). <article-title>Coastal waste detection based on deep convolutional neural networks.</article-title> <source><italic>Sensors</italic></source> <volume>21</volume>:<issue>7269</issue>. <pub-id pub-id-type="doi">10.3390/s21217269</pub-id> <pub-id pub-id-type="pmid">34770576</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>S. Q.</given-names></name> <name><surname>He</surname> <given-names>K. M.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Faster R-CNN: towards real-time object detection with region proposal networks.</article-title> <source><italic>IEEE Trans. Pattern Anal. Mach. Intell.</italic></source> <volume>39</volume>:<issue>6</issue>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id> <pub-id pub-id-type="pmid">27295650</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Global context-augmented objection detection in VHR optical remote sensing images.</article-title> <source><italic>IEEE Trans. Geosci. Remote Sens.</italic></source> <volume>59</volume> <fpage>10604</fpage>&#x2013;<lpage>10617</lpage>. <pub-id pub-id-type="doi">10.1109/tgrs.2020.3043252</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>M. Y.</given-names></name> <name><surname>Yang</surname> <given-names>H. L.</given-names></name> <name><surname>Wang</surname> <given-names>B. C.</given-names></name> <name><surname>Peng</surname> <given-names>J. H.</given-names></name> <name><surname>Zhang</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Improving boundary constraint of probability integral method in SBAS-InSAR for deformation monitoring in mining areas.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:1497. <pub-id pub-id-type="doi">10.3390/rs13081497</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>P. F.</given-names></name> <name><surname>Jiang</surname> <given-names>Q. G.</given-names></name> <name><surname>Shi</surname> <given-names>C.</given-names></name> <name><surname>Xi</surname> <given-names>J.</given-names></name> <name><surname>Tao</surname> <given-names>G. F.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2021</year>). <article-title>Oil well detection via large-scale and high-resolution remote sensing images based on improved YOLO v4.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>3243</issue>. <pub-id pub-id-type="doi">10.3390/rs13163243</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>G. L.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X. G.</given-names></name></person-group> (<year>2020</year>). &#x201C;<article-title>Revisiting the sibling head in object detector</article-title>,&#x201D; in <source><italic>Proceedings of the 2020 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Seattle, WA</publisher-loc>, <fpage>11563</fpage>&#x2013;<lpage>11572</lpage>.</citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>P. J.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>Y. F.</given-names></name> <name><surname>Fu</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>PBNet: Part-based convolutional neural network for complex composite object detection in remote sensing imagery.</article-title> <source><italic>ISPRS J. Photogramm.</italic></source> <volume>173</volume> <fpage>50</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1016/j.isprsjprs.2020.12.015</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>Z.</given-names></name> <name><surname>Shen</surname> <given-names>C. H.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>He</surname> <given-names>T.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Fcos: Fully convolutional one-stage object detection</article-title>,&#x201D; in <source><italic>Proceedings of the in 2019 IEEE International Conference on Computer Vision (ICCV)</italic></source>, <publisher-loc>Seoul</publisher-loc>, <fpage>9627</fpage>&#x2013;<lpage>9636</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00972</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>L. Y.</given-names></name> <name><surname>Deng</surname> <given-names>K. Z.</given-names></name> <name><surname>Zheng</surname> <given-names>M. N.</given-names></name></person-group> (<year>2020</year>). <article-title>Research on ground deformation monitoring method in mining areas using the probability integral model fusion D-InSAR, sub-band InSAR and offset-tracking.</article-title> <source><italic>Int. J. Appl. Earth Obs.</italic></source> <volume>85</volume>:<issue>101981</issue>. <pub-id pub-id-type="doi">10.1016/j.jag.2019.101981</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z. Y.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Yu</surname> <given-names>Y. R.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Z. J.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name></person-group> (<year>2021a</year>). <article-title>A novel phase unwrapping method used for monitoring the land subsidence in coal mining area based on U-Net convolutional neural network.</article-title> <source><italic>Front. Earth Sci.</italic></source> <volume>9</volume>:<issue>761653</issue>. <pub-id pub-id-type="doi">10.3389/feart.2021.761653</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z. Y.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name></person-group> (<year>2021b</year>). <article-title>A method of detecting the subsidence basin from InSAR interferogram in mining area based on HOG features.</article-title> <source><italic>J. China Univ. Mining Technol.</italic></source> <volume>50</volume> <fpage>404</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.13247/j.cnki.jcumt.001264</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2021</year>). <article-title>Application of local fully convolutional neural network combined with YOLO v5 algorithm in small target detection of remote sensing image.</article-title> <source><italic>PLoS One</italic></source> <volume>16</volume>:<issue>e0259283</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0259283</pub-id> <pub-id pub-id-type="pmid">34714878</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Y. P.</given-names></name> <name><surname>Yuan</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>Z. C.</given-names></name> <name><surname>Wang</surname> <given-names>L. J.</given-names></name> <name><surname>Li</surname> <given-names>H. Z.</given-names></name><etal/></person-group> (<year>2020</year>). &#x201C;<article-title>Rethinking classification and localization for object detection</article-title>,&#x201D; in <source><italic>Proceedings of the 2020 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Seattle, WA</publisher-loc>, <fpage>10186</fpage>&#x2013;<lpage>10195</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01020</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z. T.</given-names></name> <name><surname>Hou</surname> <given-names>B.</given-names></name> <name><surname>Ren</surname> <given-names>B.</given-names></name> <name><surname>Ren</surname> <given-names>Z. L.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Jiao</surname> <given-names>L. C.</given-names></name></person-group> (<year>2021</year>). <article-title>A deep detection network based on interaction of instance segmentation and object detection for SAR images.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>2582</issue>. <pub-id pub-id-type="doi">10.3390/rs13132582</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xia</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Du</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>Integration of D-InSAR and GIS technology for identifying illegal underground mining in Yangquan District, Shanxi Province, China.</article-title> <source><italic>Environ. Earth Sci.</italic></source> <volume>77</volume> <fpage>1</fpage>&#x2013;<lpage>19</lpage>.</citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xie</surname> <given-names>S. N.</given-names></name> <name><surname>Girshick</surname> <given-names>R.</given-names></name> <name><surname>Doll&#x00E1;r</surname> <given-names>P.</given-names></name> <name><surname>Tu</surname> <given-names>Z. W.</given-names></name> <name><surname>He</surname> <given-names>K. M.</given-names></name></person-group> (<year>2017</year>). &#x201C;<article-title>Aggregated residual transformations for deep neural networks</article-title>,&#x201D; in <source><italic>Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</italic></source>, <publisher-loc>Honolulu, HI</publisher-loc>, <fpage>1492</fpage>&#x2013;<lpage>1500</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.634</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>B.</given-names></name> <name><surname>Fan</surname> <given-names>P.</given-names></name> <name><surname>Lei</surname> <given-names>X. Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z. J.</given-names></name> <name><surname>Yang</surname> <given-names>F. Z.</given-names></name></person-group> (<year>2021</year>). <article-title>A real-time apple targets detection method for picking robot based on improved YOLOv5.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>1619</issue>. <pub-id pub-id-type="doi">10.3390/rs13091619</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. F.</given-names></name> <name><surname>Li</surname> <given-names>Z. W.</given-names></name> <name><surname>Zhu</surname> <given-names>J. J.</given-names></name> <name><surname>Feng</surname> <given-names>G. C.</given-names></name> <name><surname>Wang</surname> <given-names>Q. J.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2018a</year>). <article-title>Deriving time-series three-dimensional displacements of mining areas from a single-geometry InSAR dataset.</article-title> <source><italic>J. Geodesy</italic></source> <volume>92</volume> <fpage>529</fpage>&#x2013;<lpage>544</lpage>. <pub-id pub-id-type="doi">10.1007/s00190-017-1079-x</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. F.</given-names></name> <name><surname>Li</surname> <given-names>Z. W.</given-names></name> <name><surname>Zhu</surname> <given-names>J. J.</given-names></name> <name><surname>Preusse</surname> <given-names>A.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Feng</surname> <given-names>G. C.</given-names></name><etal/></person-group> (<year>2018b</year>). <article-title>An InSAR-based temporal probability integral method and its application for predicting mining-induced dynamic deformations and assessing progressive damage to surface buildings.</article-title> <source><italic>IEEE J. Sel. Topics Appl. Earth Observ.</italic></source> <volume>11</volume> <fpage>472</fpage>&#x2013;<lpage>484</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2018.2789341</pub-id></citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. F.</given-names></name> <name><surname>Li</surname> <given-names>Z. W.</given-names></name> <name><surname>Zhu</surname> <given-names>J. J.</given-names></name> <name><surname>Feng</surname> <given-names>G. C.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>L. X.</given-names></name><etal/></person-group> (<year>2018c</year>). <article-title>Locating and defining underground goaf caused by coal mining from space-borne SAR interferometry.</article-title> <source><italic>ISPRS J. Photogramm.</italic></source> <volume>135</volume> <fpage>112</fpage>&#x2013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1016/j.isprsjprs</pub-id></citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. F.</given-names></name> <name><surname>Li</surname> <given-names>Z. W.</given-names></name> <name><surname>Zhu</surname> <given-names>J. J.</given-names></name> <name><surname>Preusse</surname> <given-names>A.</given-names></name> <name><surname>Yi</surname> <given-names>H. W.</given-names></name> <name><surname>Wang</surname> <given-names>Y. J.</given-names></name><etal/></person-group> (<year>2017a</year>). <article-title>An extension of the InSAR-based probability integral method and its application for predicting 3-D mining-induced displacements under different extraction conditions.</article-title> <source><italic>IEEE Trans. Geosci. Remote Sens.</italic></source> <volume>55</volume> <fpage>3835</fpage>&#x2013;<lpage>3845</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2017.2682192</pub-id></citation></ref>
<ref id="B62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Z. F.</given-names></name> <name><surname>Li</surname> <given-names>Z. W.</given-names></name> <name><surname>Zhu</surname> <given-names>J. J.</given-names></name> <name><surname>Yi</surname> <given-names>H. W.</given-names></name> <name><surname>Hu</surname> <given-names>J.</given-names></name> <name><surname>Feng</surname> <given-names>G. C.</given-names></name></person-group> (<year>2017b</year>). <article-title>Deriving dynamic subsidence of coal mining areas using InSAR and logistic model.</article-title> <source><italic>Remote Sens.</italic></source> <volume>9</volume> <fpage>125</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.3390/rs9020125</pub-id></citation></ref>
<ref id="B63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>M. Z.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Lv</surname> <given-names>P. Y.</given-names></name> <name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Zheng</surname> <given-names>W. B.</given-names></name></person-group> (<year>2021</year>). <article-title>Subsidence monitoring base on SBAS-InSAR and slope stability analysis method for damage analysis in mountainous mining subsidence regions.</article-title> <source><italic>Remote Sens.</italic></source> <volume>13</volume>:<issue>3107</issue>. <pub-id pub-id-type="doi">10.3390/rs13163107</pub-id></citation></ref>
<ref id="B64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>J. Q.</given-names></name> <name><surname>Zhang</surname> <given-names>X. H.</given-names></name> <name><surname>Yan</surname> <given-names>J. W.</given-names></name> <name><surname>Qiu</surname> <given-names>X. L.</given-names></name> <name><surname>Yao</surname> <given-names>X.</given-names></name> <name><surname>Tian</surname> <given-names>Y. C.</given-names></name><etal/></person-group> (<year>2021</year>). <article-title>A wheat spike detection method in UAV images based on improved YOLOv5.</article-title> <source><italic>Remote Sens.</italic></source> <volume>133095</volume>. <pub-id pub-id-type="doi">10.3390/rs13163095</pub-id></citation></ref>
<ref id="B65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>M. N.</given-names></name> <name><surname>Deng</surname> <given-names>K. Z.</given-names></name> <name><surname>Fan</surname> <given-names>H. D.</given-names></name> <name><surname>Du</surname> <given-names>S.</given-names></name></person-group> (<year>2018</year>). <article-title>Monitoring and analysis of surface deformation in mining area based on InSAR and GRACE.</article-title> <source><italic>Remote Sens.</italic></source> <volume>10</volume> <fpage>1392</fpage>&#x2013;<lpage>1411</lpage>. <pub-id pub-id-type="doi">10.3390/rs10091392</pub-id></citation></ref>
</ref-list>
</back>
</article>
