<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurorobot.</journal-id>
<journal-title>Frontiers in Neurorobotics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurorobot.</abbrev-journal-title>
<issn pub-type="epub">1662-5218</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnbot.2023.1220166</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Context-aware lightweight remote-sensing image super-resolution network</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Peng</surname> <given-names>Guangwen</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2308955/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Xie</surname> <given-names>Minghong</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2308171/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Fang</surname> <given-names>Liuyang</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Faculty of Information Engineering and Automation, Kunming University of Science and Technology</institution>, <addr-line>Kunming</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Yunnan Key Laboratory of Digital Communications</institution>, <addr-line>Kunming</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Xin Jin, Yunnan University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yu Liu, Hefei University of Technology, China; Zhiqin Zhu, Chongqing University of Posts and Telecommunications, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Minghong Xie <email>minghongxie&#x00040;163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>23</day>
<month>06</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1220166</elocation-id>
<history>
<date date-type="received">
<day>10</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Peng, Xie and Fang.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Peng, Xie and Fang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>In recent years, remote-sensing image super-resolution (RSISR) methods based on convolutional neural networks (CNNs) have achieved significant progress. However, the limited receptive field of the convolutional kernel in CNNs hinders the network&#x00027;s ability to effectively capture long-range features in images, thus limiting further improvements in model performance. Additionally, the deployment of existing RSISR models to terminal devices is challenging due to their high computational complexity and large number of parameters. To address these issues, we propose a Context-Aware Lightweight Super-Resolution Network (CALSRN) for remote-sensing images. The proposed network primarily consists of Context-Aware Transformer Blocks (CATBs), which incorporate a Local Context Extraction Branch (LCEB) and a Global Context Extraction Branch (GCEB) to explore both local and global image features. Furthermore, a Dynamic Weight Generation Branch (DWGB) is designed to generate aggregation weights for global and local features, enabling dynamic adjustment of the aggregation process. Specifically, the GCEB employs a Swin Transformer-based structure to obtain global information, while the LCEB utilizes a CNN-based cross-attention mechanism to extract local information. Ultimately, global and local features are aggregated using the weights acquired from the DWGB, capturing the global and local dependencies of the image and enhancing the quality of super-resolution reconstruction. The experimental results demonstrate that the proposed method is capable of reconstructing high-quality images with fewer parameters and less computational complexity compared with existing methods.</p>
</abstract>
<kwd-group>
<kwd>convolutional neural network</kwd>
<kwd>transformer</kwd>
<kwd>remote-sensing image super-resolution</kwd>
<kwd>lightweight network</kwd>
<kwd>context-aware</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="7"/>
<equation-count count="11"/>
<ref-count count="54"/>
<page-count count="15"/>
<word-count count="7690"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The aim of single image super-resolution (SISR) is to reconstruct a high-resolution image from its associated low-resolution version. As a low-level visual task within the realm of computer vision, SISR algorithms serve to recover lost texture details in low-resolution images, thereby providing enhanced clarity for higher-level visual tasks, such as person re-identification (Li et al., <xref ref-type="bibr" rid="B23">2022b</xref>, <xref ref-type="bibr" rid="B22">2023a</xref>; Li S. et al., <xref ref-type="bibr" rid="B26">2022</xref>; Zhang et al., <xref ref-type="bibr" rid="B50">2022</xref>), medical imaging (Georgescu et al., <xref ref-type="bibr" rid="B10">2023</xref>), image dehazing/defogging (Zheng et al., <xref ref-type="bibr" rid="B51">2020</xref>; Zhu et al., <xref ref-type="bibr" rid="B54">2021c</xref>), low-resolution image fusion (Li et al., <xref ref-type="bibr" rid="B24">2016</xref>, <xref ref-type="bibr" rid="B19">2021</xref>; Xiao et al., <xref ref-type="bibr" rid="B46">2022</xref>), and remote sensing (Chen L. et al., <xref ref-type="bibr" rid="B4">2021</xref>; Jia et al., <xref ref-type="bibr" rid="B13">2023</xref>). In the field of remote sensing, high-resolution remote-sensing images can be used to obtain more detailed information about the detected area. The most direct method to obtain high-resolution remote-sensing images is to improve the precision of CMOS or charge-coupled device sensors. However, this approach entails substantial costs (Xu et al., <xref ref-type="bibr" rid="B47">2021</xref>). On the contrary, SISR technology can economically and conveniently improve the resolution of remote-sensing images.</p>
<p>In recent years, the advancement of deep learning has led to the proposal of numerous convolutional neural network (CNN)-based SISR methods, which have demonstrated remarkable performance. Dong et al. (<xref ref-type="bibr" rid="B6">2014</xref>) were the pioneers in applying CNNs to super-resolution (SR) tasks, introducing the Super-Resolution Convolutional Neural Network (SRCNN). SRCNN employs bicubic interpolation to enlarge the input low-resolution image to the target size and utilizes a three-layer convolutional network for nonlinear mapping to obtain a high-resolution image. This approach outperforms traditional super-resolution reconstruction methods. However, SRCNN suffers from high computational complexity and slow inference speed. To overcome these limitations, Dong et al. (<xref ref-type="bibr" rid="B7">2016</xref>) proposed the Fast Super-Resolution Convolutional Neural Network (FSRCNN), building upon SRCNN to directly extract features from low-resolution images, thus accelerating computation. Kim et al. introduced a Deeply-Recursive Convolutional Network (DRCN; Kim et al., <xref ref-type="bibr" rid="B15">2016b</xref>) and a Very Deep Convolutional Network for Image Super-Resolution (VDSR; Kim et al., <xref ref-type="bibr" rid="B14">2016a</xref>), both of which employ deeper convolutional layers and have achieved impressive results in SR tasks. This supports the notion that deeper CNNs can enhance model performance. Lim et al. (<xref ref-type="bibr" rid="B28">2017</xref>) proposed Enhanced Deep Residual Networks for Single Image Super-Resolution (EDSR), incorporating a deeper network and residual structure, further emphasizing that deeper networks yield superior super-resolution performance. Although high-resolution images can be obtained using the above SISR approaches, computational costs and memory consumption need to be considered when deploying the models on mobile devices, especially in the field of remote sensing (Qi et al., <xref ref-type="bibr" rid="B35">2022</xref>; Liu Y. et al., <xref ref-type="bibr" rid="B29">2023</xref>; Liu Z. et al., <xref ref-type="bibr" rid="B30">2023</xref>; Wang et al., <xref ref-type="bibr" rid="B43">2023</xref>). Wang et al. (<xref ref-type="bibr" rid="B45">2022</xref>) proposed a lightweight feature enhancement network (FeNet) for remote-sensing image super-resolution, which aims to achieve high-quality image reconstruction by effectively extracting and enhancing image features. FeNet can maintain high reconstruction quality while reducing computational complexity and memory consumption. Nonetheless, due to the limited receptive field of the convolution kernel, CNN-based super-resolution models can only acquire local image information during convolution operations, which restricts their performance. Consequently, super-resolution networks need to extract both global and local information from images to achieve further improvements in performance.</p>
<p>Transformer (Vaswani et al., <xref ref-type="bibr" rid="B42">2017</xref>) differs significantly from CNNs and is capable of capturing global information in images through its self-attention mechanism. Consequently, Liang et al. (<xref ref-type="bibr" rid="B27">2021</xref>) designed an image restoration network called SwinIR, which combines CNNs and Transformers. This network effectively models long-range dependencies in images, facilitating the restoration of global image information. However, SwinIR only relies on CNNs to extract shallow features, neglecting to fully exploit the CNN&#x00027;s potential to capture local information in intermediate layers. This results in the model&#x00027;s limited ability to acquire local information. Tu et al. (<xref ref-type="bibr" rid="B41">2022</xref>) proposed a generative adversarial network (GAN) called SWCGAN, which aims to address the limitations of convolutional layers in modeling long-range dependencies and uses a combination of Swin Transformer and convolutional layers to generate high-resolution remote-sensing images. To further investigate the aggregation of local and global information, Chen et al. (<xref ref-type="bibr" rid="B5">2022</xref>) and Gao et al. (<xref ref-type="bibr" rid="B9">2022b</xref>) proposed the image super-resolution networks HAT and LBNet, respectively. HAT employs a hybrid attention mechanism, combining channel attention and self-attention to activate more pixels, thereby enhancing the quality of super-resolution reconstruction images. Nevertheless, the hybrid attention mechanism leads to a substantial increase in the model&#x00027;s number of parameters and computational complexity. LBNet fuses symmetric CNNs with recursive Transformers to offer a high-performance, efficient solution for SISR tasks. However, LBNet directly cascades the CNN and recursive Transformer, overlooking the dynamic interaction between global and local information during the feature extraction process. Thus, further research is warranted to effectively harness the local feature extraction capabilities of CNNs and the global feature extraction capacities of Transformers to improve the performance of SISR models.</p>
<p>To address the above issues, we propose a context-aware lightweight super-resolution network (CALSRN) for remote-sensing images. This novel network is capable of extracting both local and global features from images and dynamically adjusting their fusion weights, thereby better representing image information and enhancing reconstruction quality. Furthermore, the proposed model has only about 320 K parameters, making it lighter than existing state-of-the-art lightweight super-resolution reconstruction networks while maintaining superior performance, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. This lower number of parameters results in lower computational complexity of the model. Overall, the proposed network achieves a good balance between performance and model complexity. In summary, our main contributions are as follows.</p>
<list list-type="order">
<list-item><p>We propose a lightweight remote-sensing image SR network consisting of &#x0007E;320 K parameters. In comparison to other state-of-the-art lightweight SR networks, the proposed network demonstrates the ability to reconstruct higher-quality images with a reduced number of parameters and lower computational complexity, thereby facilitating easier deployment on terminal devices.</p></list-item>
<list-item><p>We introduce a context-aware Transformer block (CATB) that is designed to not only capture local details but also concentrate on extracting global features. Simultaneously, dynamic adjustment branches are incorporated to adaptively learn the fusion weights between local and global features, resulting in a more effective feature representation and an enhanced quality of SR reconstruction images.</p></list-item>
<list-item><p>The experimental results demonstrate that the super-resolution reconstruction images generated by the proposed method exhibit substantial structural and textural details. Compared with other lightweight SISR networks, the proposed method achieves the optimum in terms of visual quality and performance evaluation.</p></list-item>
</list>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Comparison of SISR models (&#x000D7;4) in terms of accuracy, network parameters, and Multi-Adds from the RS-T1 dataset. The area of each circle denotes the number of Multi-Adds. The proposed model achieves comparable performance with fewer parameters and lower Multi-Adds. The star symbol represents the model proposed by us.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0001.tif"/>
</fig>
</sec>
<sec id="s2">
<title>2. Related works</title>
<sec>
<title>2.1. CNN-based SISR</title>
<p>In recent years, deep learning has been widely used in SISR in view of its excellent performance in image processing (Zhu et al., <xref ref-type="bibr" rid="B52">2021a</xref>; Li et al., <xref ref-type="bibr" rid="B21">2022a</xref>, <xref ref-type="bibr" rid="B25">2023b</xref>; Tang et al., <xref ref-type="bibr" rid="B38">2023</xref>) and recognition (Li et al., <xref ref-type="bibr" rid="B20">2020</xref>; Zhu et al., <xref ref-type="bibr" rid="B53">2021b</xref>; Yan et al., <xref ref-type="bibr" rid="B48">2022</xref>). Dong et al. (<xref ref-type="bibr" rid="B6">2014</xref>) first applied CNN to image super-resolution and proposed SRCNN, which outperforms the conventional SISR methods. To solve the problem of slow inference speed of SRCNN, Dong et al. (<xref ref-type="bibr" rid="B7">2016</xref>) improved SRCNN and proposed FSRCNN. To overcome the limitation of the limited receptive field of the convolutional kernel, Lei et al. (<xref ref-type="bibr" rid="B18">2017</xref>) proposed the local global combined network (LGCNet), which extracts local and global features in low-resolution images by local and global networks, respectively, and combines these two features together for super-resolution image reconstruction. Lai et al. (<xref ref-type="bibr" rid="B16">2017</xref>) Laplacian pyramid super-resolution networks (LapSRN), which combined Laplacian pyramid with deep learning to achieve multi-level super-resolution reconstruction. Since simply stacking convolutional layers may lead to gradient explosion, Kim et al. (<xref ref-type="bibr" rid="B14">2016a</xref>) proposed VDSR, which alleviates the gradient explosion problem by residual learning. From the perspective of reducing the model parameters and computational complexity, DRRN (Tai et al., <xref ref-type="bibr" rid="B36">2017a</xref>), MemNet (Tai et al., <xref ref-type="bibr" rid="B37">2017b</xref>), and LESRCNN (Tian et al., <xref ref-type="bibr" rid="B39">2020</xref>) use a recursive approach to increase the sharing of model parameters. These methods have achieved good performance for remote-sensing image super-resolution. However, their inference speed is limited by the fact that recurrent networks require deeper CNNs for information compensation. Hui et al. (<xref ref-type="bibr" rid="B12">2018</xref>) proposed information distillation network (IDN), which extracts detail and structural information through a knowledge distillation strategy to obtain better performance while reducing the model parameters. Lan et al. (<xref ref-type="bibr" rid="B17">2021</xref>) proposed a lightweight SISR model named MADNet, which effectively combines multi-scale residuals and attention mechanisms to enhance image feature representation. To effectively extract and fuse features from different levels, Lan et al. (<xref ref-type="bibr" rid="B17">2021</xref>) proposed a feature distillation and interaction weighting strategy to improve the super-resolution image quality. However, the number of parameters of the above models is still large. In order to address the limitations of memory consumption and computational burden in remote-sensing image super-resolution applications, Wang et al. (<xref ref-type="bibr" rid="B45">2022</xref>) proposed a lightweight feature enhancement network (FeNet) for accurate remote-sensing image super-resolution reconstruction. FeNet uses lightweight lattice blocks (LLB) and feature enhancement blocks (FEB) to extract and fuse features with different texture richness. FeNet has a smaller number of model parameters and faster inference speed, but its ability to capture global information is limited due to the constraints of convolutional kernel receptive field. In general, lightweight CNN-based SISR networks have difficulty in capturing global information of images, while CNN-based models with a larger number of parameters are challenging to deploy directly on terminal devices. Consequently, we design a lightweight SISR network capable of capturing both local and global image features.</p>
</sec>
<sec>
<title>2.2. Transformer-based SISR</title>
<p>In recent years, Transformer (Vaswani et al., <xref ref-type="bibr" rid="B42">2017</xref>) has been applied to low-level computer vision tasks with good results due to its global feature capture capability. Chen H. et al. (<xref ref-type="bibr" rid="B3">2021</xref>) proposed a pre-trained image processing Transformer for image recovery. Liang et al. (<xref ref-type="bibr" rid="B27">2021</xref>) proposed SwinIR network by migrating the Swin Transformer (Liu et al., <xref ref-type="bibr" rid="B31">2021</xref>) directly to the image recovery task with good results. However, the dual layer structures in the Swin Transformer block all use multi-head self-attention, which makes the SwinIR too complex. Lu et al. (<xref ref-type="bibr" rid="B32">2021</xref>) proposed an effective Transformer for SISR, which reduces GPU memory consumption through lightweight Transformers and feature separation strategies. Chen et al. (<xref ref-type="bibr" rid="B5">2022</xref>) proposed a SISR Transformer named HAT. HAT employs a hybrid attention combining channel attention and self-attention, while introducing an overlapping cross-attention module to better aggregate information across windows, achieving good super-resolution performance. However, the number of parameters in HAT is too large. Chen et al. (<xref ref-type="bibr" rid="B5">2022</xref>) proposed a lightweight super-resolution network LBNet, which combines CNN and Transformer. In LBNet, the symmetric CNN structure facilitates local feature extraction, and the recursive Transformer learns the long-term dependency relationship of images. However, LBNet only cascades CNN and Transformer, and the local and global features they extract are not well fused. In general, the aforementioned models do not adequately consider the effective aggregation of features extracted by CNN and Transformer, making it challenging to achieve an optimal balance between model size and performance. To strike a compromise between accuracy, complexity, and model size, the network&#x00027;s feature representation must be enhanced within a limited number of parameters. Consequently, we design the CATB, which can adaptively learn the fusion weights between local and global features, thereby improving the network&#x00027;s feature representation.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3. Methods</title>
<sec>
<title>3.1. Overview</title>
<p>The proposed context-aware lightweight super-resolution network (CALSRN) consists of three main parts: a shallow feature extraction module, a deep feature extraction module, and a reconstruction layer, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The shallow feature extraction module adopts a 3 &#x000D7; 3 convolution layer and a <italic>PReLU</italic> activation function to extract shallow feature, which contain more fine-grained information. Deep features are extracted through <italic>N</italic> cascaded CATBs, and features from different levels of CATBs are concatenated to obtain SR reconstruction images through reconstruction layer. CATB is a feature extraction block designed based on CNN and Transformer, which is composed of local context extraction branch (LCEB), global context extraction branch (GCEB), and dynamic weight generation branch (DWGB).</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Overall architecture of the proposed CALSRN. <bold>(A)</bold> Context-aware transformer block. <bold>(B)</bold> Local context extraction branch. <bold>(C)</bold> Global context extraction branch.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.2. Network structure</title>
<p>Given a degraded low-resolution image <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, its shallow features <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are extracted by the shallow feature extraction module, which can be formulated as:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M3"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>C</italic><sub><italic>in</italic></sub> and <italic>C</italic> denote the number of channels for low-resolution images and its shallow features, respectively. <italic>conv</italic><sub>3&#x000D7;3</sub> represents 3 &#x000D7; 3 convolution. <italic>PReLU</italic> is <italic>PReLU</italic> activation function.</p>
<p><italic><bold>F</bold></italic><sub>0</sub> is input to the deep feature extraction module to extract deep features. The deep feature extraction module consists of <italic>N</italic> CATBs. Assuming <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is the output of the <italic>n</italic>-th CATB, it can be expressed as:</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M5"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x000B7;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M6"><mml:msubsup><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mi>T</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the <italic>n</italic>-th CATB. The outputs of all CATBs are concatenated, and channel downscaling is performed using 1 &#x000D7; 1 convolution, and features of each level are fused by 3 &#x000D7; 3 convolution to obtain deep features. Then, the residual structure is used to sum the deep features and shallow features to obtain the feature <inline-formula><mml:math id="M7"><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M8"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where [&#x000B7;, &#x000B7;] denotes concatenation operation. <italic><bold>F</bold></italic><sub><italic>add</italic></sub> is fed into the reconstruction layer for super-resolution reconstruction.</p>
<p>The reconstruction layer consists of 3 &#x000D7; 3 convolution and Pixel Shuffle upsampling operation. The reconstructed result of <italic><bold>F</bold></italic><sub><italic>add</italic></sub> by the reconstruction layer are summed with the up-sampling result of the low-resolution image to obtain the super-resolution reconstruction image <italic><bold>I</bold></italic><sub><italic>SR</italic></sub>.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M9"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>d</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mi>H</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>H</italic><sub><italic>Pi</italic></sub> and <italic>H</italic><sub><italic>Bi</italic></sub> denote Pixel Shuffle upsampling operation and bilinear upsampling operation, respectively.</p>
<p>The reconstruction loss is used to constrain the proposed network. Assuming that the total number of training samples is <italic>B</italic>, the reconstruction loss can be expressed as:</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M10"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x02016;</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x02016;</mml:mo></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M11"><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="M12"><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>I</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are the <italic>i</italic>-th reconstructed super-resolution image and its corresponding labeled high-resolution image, respectively.</p>
</sec>
<sec>
<title>3.3. Context-aware transformer block</title>
<p>CATB consists of GCEB, LCEB, and DWGB. GCEB is designed based on the Swin Transformer (Liu et al., <xref ref-type="bibr" rid="B31">2021</xref>) to extract global information. LCEB uses a CNN-based cross-attention mechanism for extracting local information. DWGB adaptively generates fusion weights for weighted fusion of global and local features.</p>
<p>The structure of GCEB is shown in <xref ref-type="fig" rid="F2">Figure 2C</xref>. GCEB can be divided into two layers, which employ S-MSA (Liu et al., <xref ref-type="bibr" rid="B31">2021</xref>) and overlapping cross-attention (Lu et al., <xref ref-type="bibr" rid="B32">2021</xref>) mechanisms to achieve information interaction between windows, respectively. Among them, the first layer uses the S-MSA mechanism, and the window size determines the range of self-attention. A larger window size is beneficial for obtaining more relevant information, but expanding the window size will increase the number of parameters and model complexity. To reduce the computational complexity of the model, we introduce an overlapping cross-attention (OCA) mechanism in the second layer of GCEB. OCA enhances the expression of window self-attention by establishing cross-window connections, which is less computationally demanding than S-MSA. S-MSA is primarily used to capture spatial relationships within the input features. By applying multi-head self-attention on sliding windows, the network can focus on relevant spatial contexts and enhance its perception of local spatial details. On the other hand, OCA effectively aggregates cross-window information while reducing computational complexity, thereby enhancing the interaction between neighboring window features.</p>
<p>Let <inline-formula><mml:math id="M13"><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> denote the input of the CATB. <inline-formula><mml:math id="M14"><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> can be obtained after <italic><bold>F</bold></italic><sub><italic>n</italic>&#x02212;1</sub> is processed by the first layer of GCEB.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M15"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mo>-</mml:mo><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x02003;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>S</mml:mi><mml:mo>-</mml:mo><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>S</italic>&#x02212;<italic>MSA</italic> indicates sliding window multi-headed self-attentive operation. <italic>LN</italic> and <italic>MLP</italic> denote layer normalization and multi-layer perceptron, respectively. The output features of the second layer of GCEB are expressed as:</p>
<disp-formula id="E8"><label>(7)</label><mml:math id="M17"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mi>L</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;&#x02003;&#x000A0;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>O</mml:mi><mml:mi>C</mml:mi><mml:mi>A</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:msubsup><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic><bold>F</bold></italic><sub><italic>GCE</italic></sub> is the global feature extracted by GCEB. <italic>OCA</italic> indicates overlapping cross-attention operation.</p>
<p>As shown in <xref ref-type="fig" rid="F2">Figure 2B</xref>, in LCEB, the input feature <italic><bold>F</bold></italic><sub><italic>n</italic>&#x02212;1</sub> passes through the LN layer, 1 &#x000D7; 1 convolution for further feature extraction. The extracted features are divided into two parts along the channel: <inline-formula><mml:math id="M19"><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula> and <inline-formula><mml:math id="M20"><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:math></inline-formula>, which can be represented as:</p>
<disp-formula id="E10"><label>(8)</label><mml:math id="M21"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>i</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>L</mml:mi><mml:mi>N</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>split</italic> is the feature separation operation along the channel. <italic><bold>X</bold></italic> and <italic><bold>Y</bold></italic> go through 1 &#x000D7; 1 convolution to reshape their channel dimensions to <italic>C</italic>, respectively, and then they are passed through the PReLU activation function to obtain <inline-formula><mml:math id="M22"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M23"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:math></inline-formula>. Convolution operations are performed on <inline-formula><mml:math id="M24"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:math></inline-formula> and <inline-formula><mml:math id="M25"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle></mml:mrow><mml:mo>&#x0007E;</mml:mo></mml:mover></mml:math></inline-formula> using convolution kernels of different sizes to obtain features <italic><bold>X</bold></italic><sub>1</sub> and <italic><bold>Y</bold></italic><sub>1</sub> with different receptive fields. In order to integrate the features <italic><bold>X</bold></italic><sub>1</sub> and <italic><bold>Y</bold></italic><sub>1</sub>, we introduce a cross-attention mechanism. The features obtained by cross-attention fusion are the local features extracted by the network, which can be expressed as:</p>
<disp-formula id="E11"><label>(9)</label><mml:math id="M26"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02299;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x000D7;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>Y</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02299;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>X</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003C3; denotes Sigmoid activation function. &#x02299; denotes element-by-element multiplication.</p>
<p>In order to adaptively adjust the fusion weights of global and local features, we introduce a dynamic weight generation branch (DWGB), as shown in <xref ref-type="fig" rid="F2">Figure 2A</xref>. DWGB can adaptively learn the weighted fusion coefficients of global features and local features. The input of DWGB is <italic><bold>F</bold></italic><sub><italic>n</italic>&#x02212;1</sub> and the output is a two-dimensional vector [&#x003B1;, &#x003B2;].</p>
<disp-formula id="E12"><label>(10)</label><mml:math id="M27"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>C</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003B3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>G</mml:mi><mml:mi>A</mml:mi><mml:mi>P</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B1; and &#x003B2; are the fusion weights of local features and global features. &#x003B3; denotes ReLU activation function. <italic>FC</italic> is fully connected layer. <italic>GAP</italic> denotes global average pooling.</p>
<p>Finally, the output of CTAB is obtained by weighted fusion of local features and global features.</p>
<disp-formula id="E13"><label>(11)</label><mml:math id="M28"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mi>&#x003B2;</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold-italic'><mml:mi>F</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>G</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic><bold>F</bold></italic><sub><italic>n</italic></sub> denotes the output of CATB.</p>
</sec>
</sec>
<sec id="s4">
<title>4. Experimental results and analysis</title>
<sec>
<title>4.1. Experimental setup</title>
<p>The DIV2K dataset (Timofte et al., <xref ref-type="bibr" rid="B40">2017</xref>) was used to train the proposed network. This dataset consists of 800 training images, 100 validation images, and 100 test images, each with 2 K resolution. To comprehensively evaluate the model performance, we used two remote-sensing image datasets RS-T1 and RS-T2 (Wang et al., <xref ref-type="bibr" rid="B45">2022</xref>), as well as five super-resolution benchmark test sets: Set5 (Bevilacqua et al., <xref ref-type="bibr" rid="B2">2012</xref>), Set14 (Zeyde et al., <xref ref-type="bibr" rid="B49">2012</xref>), BSD100 (Huang et al., <xref ref-type="bibr" rid="B11">2015</xref>), Urban100 (Martin et al., <xref ref-type="bibr" rid="B33">2001</xref>), and Manga109 (Matsui et al., <xref ref-type="bibr" rid="B34">2017</xref>) to test the models. PSNR and SSIM (Wang et al., <xref ref-type="bibr" rid="B44">2004</xref>) were used as evaluation metrics to measure the quality of the reconstruction images. PSNR and SSIM are calculated on the Y channel after converting the reconstructed image from RGB space to YCbCR space. In addition, we used the number of parameters (Params) and the number of multiplication and addition operations (Muti-Adds) to evaluate the size and complexity of the model.</p>
<p>In the experiments, low-resolution images were generated from high-resolution images by bicubic downsampling with scale factors of 2&#x000D7;, 3&#x000D7;, and 4&#x000D7;. Moreover, we performed data expansion using random rotations of 90, 180, 270<sup><italic>o</italic></sup> and horizontal flips. The Adam optimizer was used to optimize the proposed network, where &#x003B2;<sub>1</sub> &#x0003D; 0.9, &#x003B2;<sub>1</sub> &#x0003D; 0.999, &#x003F5; &#x0003D; 10<sup>&#x02212;8</sup>. The size of mini-batch was set to 16. The initial learning rate was set to 5 &#x000D7; 10<sup>&#x02212;4</sup> and the learning rate was halved every 200 epochs. The total training epoch was 1, 000. The 2&#x000D7; super-resolution model was trained from scratch and it was used as a pre-training model for the 3&#x000D7; and 4&#x000D7; super-resolution models. The number of CATBs in the proposed network was 4, and 50 feature channels were used in the middle layer to ensure the lightweight of the model. All experiments were performed under Pytorch 1.12.1 framework using two NVIDIA GTX3090 GPUs (24 G).</p>
</sec>
<sec>
<title>4.2. Experiments on remote-sensing image datasets</title>
<p>To validate the effectiveness of the proposed model in this paper, we compare the proposed method with state-of-the-arts methods [SRCNN (Dong et al., <xref ref-type="bibr" rid="B6">2014</xref>), VDSR (Kim et al., <xref ref-type="bibr" rid="B14">2016a</xref>), LGCNet (Lei et al., <xref ref-type="bibr" rid="B18">2017</xref>), LapSRN (Lai et al., <xref ref-type="bibr" rid="B16">2017</xref>), CARN-M (Ahn et al., <xref ref-type="bibr" rid="B1">2018</xref>), IDN (Hui et al., <xref ref-type="bibr" rid="B12">2018</xref>), LESRCNN (Tian et al., <xref ref-type="bibr" rid="B39">2020</xref>), FeNet (Wang et al., <xref ref-type="bibr" rid="B45">2022</xref>), LBNet (Gao et al., <xref ref-type="bibr" rid="B9">2022b</xref>)] on the remote-sensing image datasets RS-T1 and RS-T2 (Wang et al., <xref ref-type="bibr" rid="B45">2022</xref>). Both RS-T1 and RS-T2 consist of 120 images covering 21 complex ground truth remote-sensing scenarios. For a fair comparison, all comparison methods are tested on the RS-T1 and RS-T2 datasets using models trained on the DIV2K dataset. <xref ref-type="table" rid="T1">Tables 1</xref>&#x02013;<xref ref-type="table" rid="T3">3</xref> demonstrate the results of the quantitative evaluation of the compared methods on the RS-T1 and RS-T2 datasets. According to <xref ref-type="table" rid="T1">Tables 1</xref>&#x02013;<xref ref-type="table" rid="T3">3</xref> that the PSNR/SSIM values of the 2&#x000D7;, 3&#x000D7;, 4&#x000D7; super-resolution reconstruction results of the proposed method on RS-T1 and RS-T2 datasets are optimal. Moreover, the Multi-Adds value of the proposed method is the best, and the number of parameters is about 30 K less than that of FeNet, which is the current optimal lightweight super-resolution reconstruction model for remote-sensing images. It confirms that the proposed method can achieve good performance with a small number of parameters.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Quantitative comparison of 2&#x000D7; super-resolution results obtained by different methods on RS-T1 and RS-T2 datasets.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-adds (G)</bold></th>
<th valign="top" align="center" colspan="2"><bold>PSNR/SSIM</bold></th>
</tr>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>RS-T1</bold></th>
<th valign="top" align="center"><bold>RS-T2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SRCNN</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">52.7</td>
<td valign="top" align="center">35.18/0.9243</td>
<td valign="top" align="center">32.87/0.9209</td>
</tr> <tr>
<td valign="top" align="left">VDSR</td>
<td valign="top" align="center">666</td>
<td valign="top" align="center">612.6</td>
<td valign="top" align="center">35.85/0.9312</td>
<td valign="top" align="center">33.86/0.9312</td>
</tr> <tr>
<td valign="top" align="left">LGCNet</td>
<td valign="top" align="center">193</td>
<td valign="top" align="center">178.1</td>
<td valign="top" align="center">35.65/0.9298</td>
<td valign="top" align="center">33.47/0.9281</td>
</tr> <tr>
<td valign="top" align="left">LapSRN</td>
<td valign="top" align="center">251</td>
<td valign="top" align="center">29.9</td>
<td valign="top" align="center">35.69/0.9304</td>
<td valign="top" align="center">33.57/0.9286</td>
</tr> <tr>
<td valign="top" align="left">CARN-M</td>
<td valign="top" align="center">412</td>
<td valign="top" align="center">91.2</td>
<td valign="top" align="center">35.77/0.9314</td>
<td valign="top" align="center">33.84/0.9315</td>
</tr> <tr>
<td valign="top" align="left">IDN</td>
<td valign="top" align="center">553</td>
<td valign="top" align="center">124.6</td>
<td valign="top" align="center">36.13/0.9339</td>
<td valign="top" align="center">34.07/0.9329</td>
</tr> <tr>
<td valign="top" align="left">LESRCNN</td>
<td valign="top" align="center">626</td>
<td valign="top" align="center">281.5</td>
<td valign="top" align="center">36.04/0.9328</td>
<td valign="top" align="center">34.00/0.9320</td>
</tr> <tr>
<td valign="top" align="left">FeNet</td>
<td valign="top" align="center">351</td>
<td valign="top" align="center">77.9</td>
<td valign="top" align="center">36.23/0.9341</td>
<td valign="top" align="center">34.22/0.9337</td>
</tr> <tr>
<td valign="top" align="left">LBNet</td>
<td valign="top" align="center">731</td>
<td valign="top" align="center">153.2</td>
<td valign="top" align="center">36.28/0.9345</td>
<td valign="top" align="center">34.30/0.9339</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">319</td>
<td valign="top" align="center">20.4</td>
<td valign="top" align="center"><bold>36.34</bold>/<bold>0.9356</bold></td>
<td valign="top" align="center"><bold>34.37</bold>/<bold>0.9349</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Multi-Adds is computed according to a 1, 280 &#x000D7; 720 image.</p>
<p>The bold values represent the best performance.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Quantitative comparison of 3&#x000D7; super-resolution results obtained by different methods on RS-T1 and RS-T2 datasets.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center" colspan="2"><bold>PSNR/SSIM</bold></th>
</tr>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>RS-T1</bold></th>
<th valign="top" align="center"><bold>RS-T2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SRCNN</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">52.7</td>
<td valign="top" align="center">30.95/0.8228</td>
<td valign="top" align="center">28.59/0.8180</td>
</tr> <tr>
<td valign="top" align="left">VDSR</td>
<td valign="top" align="center">666</td>
<td valign="top" align="center">612.6</td>
<td valign="top" align="center">31.55/0.9352</td>
<td valign="top" align="center">29.40/0.8391</td>
</tr> <tr>
<td valign="top" align="left">LGCNet</td>
<td valign="top" align="center">193</td>
<td valign="top" align="center">79.0</td>
<td valign="top" align="center">31.30/0.8314</td>
<td valign="top" align="center">29.03/0.8312</td>
</tr> <tr>
<td valign="top" align="left">LapSRN</td>
<td valign="top" align="center">290</td>
<td valign="top" align="center">115.2</td>
<td valign="top" align="center">31.47/0.8338</td>
<td valign="top" align="center">29.22/0.8352</td>
</tr> <tr>
<td valign="top" align="left">CARN-M</td>
<td valign="top" align="center">412</td>
<td valign="top" align="center">46.1</td>
<td valign="top" align="center">31.72/0.8426</td>
<td valign="top" align="center">29.62/0.8452</td>
</tr> <tr>
<td valign="top" align="left">IDN</td>
<td valign="top" align="center">553</td>
<td valign="top" align="center">56.3</td>
<td valign="top" align="center">31.73/0.8430</td>
<td valign="top" align="center">29.59/0.8450</td>
</tr> <tr>
<td valign="top" align="left">LESRCNN</td>
<td valign="top" align="center">810</td>
<td valign="top" align="center">238.9</td>
<td valign="top" align="center">31.68/0.8398</td>
<td valign="top" align="center">29.65/0.8444</td>
</tr> <tr>
<td valign="top" align="left">FeNet</td>
<td valign="top" align="center">357</td>
<td valign="top" align="center">35.2</td>
<td valign="top" align="center">31.89/0.8432</td>
<td valign="top" align="center">29.80/0.8481</td>
</tr> <tr>
<td valign="top" align="left">LBNet</td>
<td valign="top" align="center">736</td>
<td valign="top" align="center">51.5</td>
<td valign="top" align="center">31.96/0.8485</td>
<td valign="top" align="center">29.91/0.8516</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">326</td>
<td valign="top" align="center">20.8</td>
<td valign="top" align="center"><bold>32.05</bold>/<bold>0.8505</bold></td>
<td valign="top" align="center"><bold>30.01</bold>/<bold>0.8526</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Multi-Adds is computed according to a 1, 280 &#x000D7; 720 image.</p>
<p>The bold values represent the best performance.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Quantitative comparison of 4&#x000D7; super-resolution results obtained by different methods on RS-T1 and RS-T2 datasets.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center" colspan="2"><bold>PSNR/SSIM</bold></th>
</tr>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>RS-T1</bold></th>
<th valign="top" align="center"><bold>RS-T2</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SRCNN</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">52.7</td>
<td valign="top" align="center">28.87/0.7382</td>
<td valign="top" align="center">26.46/0.7296</td>
</tr> <tr>
<td valign="top" align="left">VDSR</td>
<td valign="top" align="center">666</td>
<td valign="top" align="center">612.6</td>
<td valign="top" align="center">29.33/0.7546</td>
<td valign="top" align="center">27.03/0.7525</td>
</tr> <tr>
<td valign="top" align="left">LGCNet</td>
<td valign="top" align="center">193</td>
<td valign="top" align="center">44.5</td>
<td valign="top" align="center">29.13/0.7481</td>
<td valign="top" align="center">26.76/0.7426</td>
</tr> <tr>
<td valign="top" align="left">LapSRN</td>
<td valign="top" align="center">543</td>
<td valign="top" align="center">139.6</td>
<td valign="top" align="center">29.51/0.7614</td>
<td valign="top" align="center">27.24/0.7600</td>
</tr> <tr>
<td valign="top" align="left">CARN-M</td>
<td valign="top" align="center">412</td>
<td valign="top" align="center">32.5</td>
<td valign="top" align="center">29.57/0.7624</td>
<td valign="top" align="center">27.37/0.7647</td>
</tr> <tr>
<td valign="top" align="left">IDN</td>
<td valign="top" align="center">553</td>
<td valign="top" align="center">32.3</td>
<td valign="top" align="center">29.56/0.7623</td>
<td valign="top" align="center">27.31/0.7627</td>
</tr> <tr>
<td valign="top" align="left">LESRCNN</td>
<td valign="top" align="center">774</td>
<td valign="top" align="center">241.6</td>
<td valign="top" align="center">29.62/0.7625</td>
<td valign="top" align="center">27.41/0.7646</td>
</tr> <tr>
<td valign="top" align="left">FeNet</td>
<td valign="top" align="center">366</td>
<td valign="top" align="center">20.4</td>
<td valign="top" align="center">29.70/0.7688</td>
<td valign="top" align="center">27.45/0.7672</td>
</tr> <tr>
<td valign="top" align="left">LBNet</td>
<td valign="top" align="center">742</td>
<td valign="top" align="center">38.9</td>
<td valign="top" align="center">29.78/0.7689</td>
<td valign="top" align="center">27.52/0.7732</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">336</td>
<td valign="top" align="center">21.4</td>
<td valign="top" align="center"><bold>29.85</bold>/<bold>0.7717</bold></td>
<td valign="top" align="center"><bold>27.67</bold>/<bold>0.7759</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>Multi-Adds is computed according to a 1, 280 &#x000D7; 720 image.</p>
<p>The bold values represent the best performance.</p>
</table-wrap-foot>
</table-wrap>
<p>In addition, the 2&#x000D7;, 3&#x000D7;, and 4&#x000D7; super-resolution reconstruction results of remote-sensing images are illustrated in <xref ref-type="fig" rid="F3">Figures 3</xref>&#x02013;<xref ref-type="fig" rid="F5">5</xref>, respectively. As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, when the magnification factor is 2, the visual effect of the proposed method on &#x0201C;overpass63,&#x0201D; &#x0201C;Sparseresidential10,&#x0201D; and &#x0201C;freeway41&#x0201D; is better than that of the comparison methods in terms of clarity, and the PSNR and SSIM values are also optimal. As shown in <xref ref-type="fig" rid="F4">Figure 4</xref>, the 3&#x000D7; reconstructed images of the proposed method achieve the optimal quality in terms of both structure and detailed texture, especially for the &#x0201C;Denseresidential46&#x0201D; image, where the comparison methods fail to recover the corner information. As shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, IDN, FeNet, LBNet, and our proposed method all achieve good visual results, while the reconstructed images of the remaining comparison methods are relatively blurry. Overall, as a lightweight super-resolution model, the proposed model achieves better quantitative and qualitative results than existing models.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Visual comparisons with different methods for 2&#x000D7; super-resolution on RS-T1 and RS-T2 datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0003.tif"/>
</fig>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Visual comparisons with different methods for 3&#x000D7; super-resolution on RS-T1 and RS-T2 datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Visual comparisons with different methods for 4&#x000D7; super-resolution on RS-T1 and RS-T2 datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0005.tif"/>
</fig>
</sec>
<sec>
<title>4.3. Experiments on super-resolution benchmark test sets</title>
<p>To further verify the generalization of the proposed model in this paper, we conduct comparison experiments on the benchmark test sets. The datasets Set5, Set14, BSD100, Urban100, and Manga109 are benchmark test sets for image super-resolution reconstruction, covering images of different scenes such as urban buildings, animals, plants, and animations. In order to verify the effectiveness of the proposed method, we compare the proposed method with state-of-the-arts methods [SRCNN (Dong et al., <xref ref-type="bibr" rid="B6">2014</xref>), FSRCNN (Dong et al., <xref ref-type="bibr" rid="B7">2016</xref>), LapSRN (Lai et al., <xref ref-type="bibr" rid="B16">2017</xref>), VDSR (Kim et al., <xref ref-type="bibr" rid="B14">2016a</xref>), LGCNet (Lei et al., <xref ref-type="bibr" rid="B18">2017</xref>), DRRN (Tai et al., <xref ref-type="bibr" rid="B36">2017a</xref>), CARN-M (Ahn et al., <xref ref-type="bibr" rid="B1">2018</xref>), IDN (Hui et al., <xref ref-type="bibr" rid="B12">2018</xref>), MemNet (Tai et al., <xref ref-type="bibr" rid="B37">2017b</xref>), LESRCNN (Tian et al., <xref ref-type="bibr" rid="B39">2020</xref>), MADNet (Lan et al., <xref ref-type="bibr" rid="B17">2021</xref>), FDIWN (Gao et al., <xref ref-type="bibr" rid="B8">2022a</xref>), LBNet (Gao et al., <xref ref-type="bibr" rid="B9">2022b</xref>)] on the five test sets mentioned above. It is worth noting that the models for the comparison methods are the already trained models provided by the original authors. The objective evaluation results of the 3&#x000D7; and 4&#x000D7; magnification factor super-resolution reconstruction experiments are shown in <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref>. The best values are highlighted in bold. As shown in <xref ref-type="table" rid="T4">Tables 4</xref>, <xref ref-type="table" rid="T5">5</xref>, the PSNR/SSIM values of the proposed method outperforms the others at the most metrics. Moreover, compared with the LBNet and FDIWN methods, which have comparable performance to the proposed method, they have more than twice the number of parameters and much larger Multi-Adds than those of the proposed method. Overall, the proposed model achieves a good balance among the number of parameters, complexity and performance.</p>
<table-wrap position="float" id="T4">
<label>Table 4</label>
<caption><p>Quantitative comparison of 3&#x000D7; super-resolution results obtained by different methods on super-resolution benchmark datasets.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center" colspan="5"><bold>PSNR/SSIM</bold></th>
</tr>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>Set5</bold></th>
<th valign="top" align="center"><bold>Set14</bold></th>
<th valign="top" align="center"><bold>BSD100</bold></th>
<th valign="top" align="center"><bold>Urban100</bold></th>
<th valign="top" align="center"><bold>Manga109</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SRCNN</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">52.7</td>
<td valign="top" align="center">32.75/0.9090</td>
<td valign="top" align="center">29.30/0.8215</td>
<td valign="top" align="center">28.41/0.7863</td>
<td valign="top" align="center">26.24/0.7989</td>
<td valign="top" align="center">30.48/0.9117</td>
</tr> <tr>
<td valign="top" align="left">FSRCNN</td>
<td valign="top" align="center">54</td>
<td valign="top" align="center">5.0</td>
<td valign="top" align="center">33.18/0.9140</td>
<td valign="top" align="center">29.37/0.8240</td>
<td valign="top" align="center">28.53/0.7910</td>
<td valign="top" align="center">26.43/0.8080</td>
<td valign="top" align="center">31.10/0.9210</td>
</tr> <tr>
<td valign="top" align="left">LapSRN</td>
<td valign="top" align="center">502</td>
<td valign="top" align="center">115.2</td>
<td valign="top" align="center">33.81/0.9220</td>
<td valign="top" align="center">29.79/0.8325</td>
<td valign="top" align="center">28.82/0.7980</td>
<td valign="top" align="center">27.07/0.8275</td>
<td valign="top" align="center">32.21/0.9350</td>
</tr> <tr>
<td valign="top" align="left">VDSR</td>
<td valign="top" align="center">666</td>
<td valign="top" align="center">612.6</td>
<td valign="top" align="center">33.66/0.9213</td>
<td valign="top" align="center">29.77/0.8314</td>
<td valign="top" align="center">28.82/0.7976</td>
<td valign="top" align="center">27.14/0.8279</td>
<td valign="top" align="center">32.01/0.9340</td>
</tr> <tr>
<td valign="top" align="left">LGCNet</td>
<td valign="top" align="center">193</td>
<td valign="top" align="center">79.0</td>
<td valign="top" align="center">33.32/0.9172</td>
<td valign="top" align="center">29.67/0.8289</td>
<td valign="top" align="center">28.63/0.7923</td>
<td valign="top" align="center">26.77/0.8180</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">DRRN</td>
<td valign="top" align="center">298</td>
<td valign="top" align="center">6796,9</td>
<td valign="top" align="center">34.03/0.9244</td>
<td valign="top" align="center">29.96/0.8349</td>
<td valign="top" align="center">28.95/0.8004</td>
<td valign="top" align="center">27.53/0.8378</td>
<td valign="top" align="center">32.71/0.9379</td>
</tr> <tr>
<td valign="top" align="left">CARN-M</td>
<td valign="top" align="center">412</td>
<td valign="top" align="center">46.1</td>
<td valign="top" align="center">33.99/0.9236</td>
<td valign="top" align="center">30.08/0.8367</td>
<td valign="top" align="center">28.91/0.8000</td>
<td valign="top" align="center">27.55/0.8385</td>
<td valign="top" align="center">32.78/0.9384</td>
</tr> <tr>
<td valign="top" align="left">IDN</td>
<td valign="top" align="center">553</td>
<td valign="top" align="center">56.3</td>
<td valign="top" align="center">34.11/0.9253</td>
<td valign="top" align="center">29.99/0.8354</td>
<td valign="top" align="center">28.95/0.8013</td>
<td valign="top" align="center">27.42/0.8359</td>
<td valign="top" align="center">32.71/0.9381</td>
</tr> <tr>
<td valign="top" align="left">MemNet</td>
<td valign="top" align="center">678</td>
<td valign="top" align="center">2662.4</td>
<td valign="top" align="center">34.09/0.9248</td>
<td valign="top" align="center">30.00/0.8350</td>
<td valign="top" align="center">28.96/0.8001</td>
<td valign="top" align="center">27.56/0.8376</td>
<td valign="top" align="center">32.51/0.9369</td>
</tr> <tr>
<td valign="top" align="left">LESRCNN</td>
<td valign="top" align="center">810</td>
<td valign="top" align="center">238.9</td>
<td valign="top" align="center">33.93/0.9231</td>
<td valign="top" align="center">30.12/0.8380</td>
<td valign="top" align="center">28.91/0.8005</td>
<td valign="top" align="center">27.70/0.8415</td>
<td valign="top" align="center">32.76/0.9389</td>
</tr> <tr>
<td valign="top" align="left">MADNet</td>
<td valign="top" align="center">930</td>
<td valign="top" align="center">88.4</td>
<td valign="top" align="center">34.14/0.9251</td>
<td valign="top" align="center">30.20/0.8395</td>
<td valign="top" align="center">28.98/0.8023</td>
<td valign="top" align="center">27.78/0.8439</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">LBNet</td>
<td valign="top" align="center">736</td>
<td valign="top" align="center">68.4</td>
<td valign="top" align="center">34.47/0.9277</td>
<td valign="top" align="center">30.38/0.8417</td>
<td valign="top" align="center">29.13/0.8061</td>
<td valign="top" align="center">28.42/0.8559</td>
<td valign="top" align="center">33.80/0.9430</td>
</tr> <tr>
<td valign="top" align="left">FDIWN</td>
<td valign="top" align="center">645</td>
<td valign="top" align="center">51.5</td>
<td valign="top" align="center"><bold>34.52</bold>/0.9281</td>
<td valign="top" align="center">30.42/0.8438</td>
<td valign="top" align="center">29.14/0.8065</td>
<td valign="top" align="center">28.35/0.8567</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">326</td>
<td valign="top" align="center">20.8</td>
<td valign="top" align="center">34.50/<bold>0.9283</bold></td>
<td valign="top" align="center"><bold>30.52</bold>/<bold>0.8455</bold></td>
<td valign="top" align="center"><bold>29.17</bold>/<bold>0.8085</bold></td>
<td valign="top" align="center"><bold>28.49</bold>/<bold>0.8586</bold></td>
<td valign="top" align="center"><bold>33.99</bold>/<bold>0.9470</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>&#x02212;Indicates that the result is unknown. Multi-Adds is computed according to a 1, 280 &#x000D7; 720 image.</p>
<p>The bold values represent the best performance.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T5">
<label>Table 5</label>
<caption><p>Quantitative comparison of 4&#x000D7; super-resolution results obtained by different methods on super-resolution benchmark datasets.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>Method</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center" colspan="5"><bold>PSNR/SSIM</bold></th>
</tr>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th/>
<th/>
<th/>
<th valign="top" align="center"><bold>Set5</bold></th>
<th valign="top" align="center"><bold>Set14</bold></th>
<th valign="top" align="center"><bold>BSD100</bold></th>
<th valign="top" align="center"><bold>Urban100</bold></th>
<th valign="top" align="center"><bold>Manga109</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">SRCNN</td>
<td valign="top" align="center">57</td>
<td valign="top" align="center">52.7</td>
<td valign="top" align="center">30.48/0.8626</td>
<td valign="top" align="center">27.50/0.7513</td>
<td valign="top" align="center">26.90/0.7101</td>
<td valign="top" align="center">24.52/0.7221</td>
<td valign="top" align="center">27.58/0.8555</td>
</tr> <tr>
<td valign="top" align="left">FSRCNN</td>
<td valign="top" align="center">54</td>
<td valign="top" align="center">4.6</td>
<td valign="top" align="center">30.72/0.8660</td>
<td valign="top" align="center">26.98/0.7150</td>
<td valign="top" align="center">26.98/0.7150</td>
<td valign="top" align="center">24.62/0.7280</td>
<td valign="top" align="center">27.90/0.8610</td>
</tr> <tr>
<td valign="top" align="left">LapSRN</td>
<td valign="top" align="center">543</td>
<td valign="top" align="center">139.6</td>
<td valign="top" align="center">31.54/0.8852</td>
<td valign="top" align="center">28.09/0.7700</td>
<td valign="top" align="center">27.32/0.7275</td>
<td valign="top" align="center">25.21/0.7562</td>
<td valign="top" align="center">29.09/0.8900</td>
</tr> <tr>
<td valign="top" align="left">VDSR</td>
<td valign="top" align="center">666</td>
<td valign="top" align="center">612.6</td>
<td valign="top" align="center">31.35/0.8838</td>
<td valign="top" align="center">28.01/0.7674</td>
<td valign="top" align="center">27.29/0.7251</td>
<td valign="top" align="center">25.18/0.7524</td>
<td valign="top" align="center">28.83/0.8870</td>
</tr> <tr>
<td valign="top" align="left">LGCNet</td>
<td valign="top" align="center">193</td>
<td valign="top" align="center">44.5</td>
<td valign="top" align="center">30.87/0.8746</td>
<td valign="top" align="center">27.82/0.7630</td>
<td valign="top" align="center">27.08/0.7186</td>
<td valign="top" align="center">24.82/0.7399</td>
<td valign="top" align="center">-</td>
</tr> <tr>
<td valign="top" align="left">DRRN</td>
<td valign="top" align="center">298</td>
<td valign="top" align="center">6796.9</td>
<td valign="top" align="center">31.68/0.8888</td>
<td valign="top" align="center">28.21/0.7720</td>
<td valign="top" align="center">27.38/0.7284</td>
<td valign="top" align="center">25.44/0.7638</td>
<td valign="top" align="center">29.45/0.8946</td>
</tr> <tr>
<td valign="top" align="left">CARN-M</td>
<td valign="top" align="center">412</td>
<td valign="top" align="center">32.5</td>
<td valign="top" align="center">31.92/0.8903</td>
<td valign="top" align="center">28.42/0.7762</td>
<td valign="top" align="center">27.44/0.7304</td>
<td valign="top" align="center">25.63/0.7688</td>
<td valign="top" align="center">29.80/0.8989</td>
</tr> <tr>
<td valign="top" align="left">IDN</td>
<td valign="top" align="center">553</td>
<td valign="top" align="center">32.3</td>
<td valign="top" align="center">31.82/0.8903</td>
<td valign="top" align="center">28.25/0.7730</td>
<td valign="top" align="center">27.41/0.7297</td>
<td valign="top" align="center">25.41/0.7632</td>
<td valign="top" align="center">29.41/0.8942</td>
</tr> <tr>
<td valign="top" align="left">MemNet</td>
<td valign="top" align="center">678</td>
<td valign="top" align="center">2662.4</td>
<td valign="top" align="center">31.74/0.8893</td>
<td valign="top" align="center">28.26/0.7723</td>
<td valign="top" align="center">27.40/0.7281</td>
<td valign="top" align="center">25.50/0.7630</td>
<td valign="top" align="center">29.42/0.8942</td>
</tr> <tr>
<td valign="top" align="left">LESRCNN</td>
<td valign="top" align="center">774</td>
<td valign="top" align="center">241.6</td>
<td valign="top" align="center">31.88/0.8903</td>
<td valign="top" align="center">28.44/0.7772</td>
<td valign="top" align="center">27.45/0.7313</td>
<td valign="top" align="center">25.77/0.7732</td>
<td valign="top" align="center">29.94/0.9002</td>
</tr>
<tr>
<td valign="top" align="left">MADNet</td>
<td valign="top" align="center">1002</td>
<td valign="top" align="center">54.1</td>
<td valign="top" align="center">32.01/0.8925</td>
<td valign="top" align="center">28.45/0.7781</td>
<td valign="top" align="center">27.47/0.7327</td>
<td valign="top" align="center">25.77/0.7751</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">LBNet</td>
<td valign="top" align="center">742</td>
<td valign="top" align="center">38.9</td>
<td valign="top" align="center"><bold>32.29</bold>/0.8960</td>
<td valign="top" align="center">28.68/0.7832</td>
<td valign="top" align="center">27.62/0.7382</td>
<td valign="top" align="center">26.27/0.7906</td>
<td valign="top" align="center">30.76/0.9111</td>
</tr> <tr>
<td valign="top" align="left">FDIWN</td>
<td valign="top" align="center">664</td>
<td valign="top" align="center">28.4</td>
<td valign="top" align="center">32.23/08955</td>
<td valign="top" align="center">28.66/07829</td>
<td valign="top" align="center">27.62/07380</td>
<td valign="top" align="center">26.28/07919</td>
<td valign="top" align="center">&#x02013;</td>
</tr> <tr>
<td valign="top" align="left">Proposed</td>
<td valign="top" align="center">336</td>
<td valign="top" align="center">21.4</td>
<td valign="top" align="center">32.27/<bold>0.8965</bold></td>
<td valign="top" align="center"><bold>28.70</bold>/<bold>0.7845</bold></td>
<td valign="top" align="center"><bold>27.66</bold>/<bold>0.7409</bold></td>
<td valign="top" align="center"><bold>26.53</bold>/<bold>0.7986</bold></td>
<td valign="top" align="center"><bold>30.98</bold>/<bold>0.9145</bold></td>
</tr></tbody>
</table>
<table-wrap-foot>
<p>&#x02212;Indicates that the result is unknown. Multi-Adds is computed according to a 1, 280 &#x000D7; 720 image.</p>
<p>The bold values represent the best performance.</p>
</table-wrap-foot>
</table-wrap>
<p>To evaluate the visual quality of the super-resolution reconstruction images, the 3&#x000D7; and 4&#x000D7; super-resolution images are shown in <xref ref-type="fig" rid="F6">Figures 6</xref>, <xref ref-type="fig" rid="F7">7</xref>, respectively. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, when the images are enlarged by 3 times, artifacts are introduced in the reconstruction results of the comparison methods. As shown in <xref ref-type="fig" rid="F7">Figure 7</xref>, the 4&#x000D7; super-resolution results of the &#x0201C;img073&#x0201D; and &#x0201C;img092&#x0201D; images in the Urban test set are closest to the Ground-Truth images, and achieve the best visual experience in terms of overall image clarity and detail texture. Other comparison methods exhibit visible artifacts, such as severe misalignment in the locally zoomed-in &#x0201C;img092&#x0201D; images restored by SRCNN, VDSR, LapSRN, CARN-M, IDN, LESRCNN, and FDIWM. Overall, compared with existing methods, the visual quality of the reconstructed images by the proposed method is optimal in terms of clarity and detailed texture.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Visual comparisons with different methods for 3&#x000D7; super-resolution on super-resolution benchmark datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0006.tif"/>
</fig>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Visual comparisons with different methods for 4&#x000D7; super-resolution on super-resolution benchmark datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0007.tif"/>
</fig>
</sec>
<sec>
<title>4.4. Ablation study</title>
<p>To verify the effectiveness of the global context extraction branch (GCEB), local context extraction branch (LCEB) and dynamic weight generation branch (DWGB) proposed in this work, we conduct ablation experiments on the RS-T1 dataset with a magnification factor of 2. The results of the ablation experiments are shown in <xref ref-type="table" rid="T6">Table 6</xref>. In the ablation experiment, this paper removes DWGB, GCEB, and LCEB one by one from the complete model, and then compare the performance of the modified model with the complete model. As shown in <xref ref-type="table" rid="T6">Table 6</xref>, the model performance all decreases when DWGB, GCEB, and LCEB are removed from the complete model. This indicates that DWGB, GCEB, and LCEB all have a positive effect on improving the model performance.</p>
<table-wrap position="float" id="T6">
<label>Table 6</label>
<caption><p>Ablation study of each module on the RS-T1 dataset with a magnification factor of 2.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>LCEB</bold></th>
<th valign="top" align="left"><bold>GCEB</bold></th>
<th valign="top" align="left"><bold>DWGB</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center"><bold>PSNR/SSIM</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="center">319</td>
<td valign="top" align="center">20.4</td>
<td valign="top" align="center">36.29/0.9343</td>
</tr> <tr>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x000D7;</td>
<td valign="top" align="center">304</td>
<td valign="top" align="center">15.9</td>
<td valign="top" align="center">36.11/0.9332</td>
</tr> <tr>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x000D7;</td>
<td valign="top" align="left">&#x000D7;</td>
<td valign="top" align="center">156</td>
<td valign="top" align="center">15.1</td>
<td valign="top" align="center">36.02/0.9328</td>
</tr> <tr>
<td valign="top" align="left">&#x000D7;</td>
<td valign="top" align="left">&#x02713;</td>
<td valign="top" align="left">&#x000D7;</td>
<td valign="top" align="center">256</td>
<td valign="top" align="center">15.3</td>
<td valign="top" align="center">36.06/0.9329</td>
</tr></tbody>
</table>
</table-wrap>
<p>We use the &#x0201C;denseresidential13&#x0201D; and &#x0201C;baseballdiamond98&#x0201D; images from the RS-T1 test set to verify the ablation experiment visually. From the local zoom-in visual results in <xref ref-type="fig" rid="F8">Figure 8</xref>, it can be seen that the image quality decreases when DWGB, GCEB and LCEB are removed from the complete model one by one. The combination of LCEB&#x0002B;GCEB&#x0002B;DWGB achieves the best visual performance, and when DWGB is removed from the complete model, the model&#x00027;s performance decreases, resulting in blurry images for &#x0201C;denseresidential13&#x0201D; and &#x0201C;baseballdiamond98.&#x0201D; This demonstrates the crucial role of the dynamic weight generation branch (DWGB) in adjusting global and local information within the overall network. When the model only has LCEB or GCEB, the reconstructed images is blurry. The visual results in <xref ref-type="fig" rid="F8">Figure 8</xref> confirm the effectiveness of the proposed modules. LCEB captures local details, GCEB extracts global information, and DWGB dynamically assigns weights and fuse the local and global features.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Visual results of 2&#x000D7; ablation experiments on remote sensing test set RS-T1.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fnbot-17-1220166-g0008.tif"/>
</fig>
<p>In addition, we analyze the effect of the number of CATBs on the performance of the proposed model. The experiment is conducted on the RS-T1 dataset with a magnification factor of 2. The results of the experiment are shown in <xref ref-type="table" rid="T7">Table 7</xref>. As shown in <xref ref-type="table" rid="T7">Table 7</xref>, the PSNR/SSIM values of the reconstructed images improve as the number of CATB blocks increases, but the number of parameters and the complexity of the model also increase. When the number of CATB blocks is 3, the model has the smallest number of parameters and computational complexity, but the PSNR and SSIM of the reconstructed images are also the lowest. When the number of CATB blocks is increased to 6, the best performance is achieved, but the number of model parameters and computational complexity are too large. Therefore, to balance the number of parameters and the performance of the model, we set the number of CATB blocks to 4.</p>
<table-wrap position="float" id="T7">
<label>Table 7</label>
<caption><p>Performance of the proposed model on RS-T1 with different number of CATBs.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:&#x00023;919498;color:&#x00023;ffffff">
<th valign="top" align="left"><bold>CATBs</bold></th>
<th valign="top" align="center"><bold>Params (K)</bold></th>
<th valign="top" align="center"><bold>Multi-Adds (G)</bold></th>
<th valign="top" align="center"><bold>PSNR/SSIM</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">245.7</td>
<td valign="top" align="center">16.8</td>
<td valign="top" align="center">36.18/0.9322</td>
</tr> <tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">319.8</td>
<td valign="top" align="center">20.4</td>
<td valign="top" align="center">36.30/0.9343</td>
</tr> <tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">406.2</td>
<td valign="top" align="center">26.8</td>
<td valign="top" align="center">36.38/0.9347</td>
</tr> <tr>
<td valign="top" align="left">6</td>
<td valign="top" align="center">478.3</td>
<td valign="top" align="center">30.7</td>
<td valign="top" align="center">36.42/0.9350</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="conclusions" id="s5">
<title>5. Conclusion</title>
<p>We propose a lightweight SISR network called CLASRN for the super-resolution reconstruction of low-resolution remote-sensing images. CLASRN combines the advantages of Transformer and CNN to better recover local details while emphasizing long-range information in images. Furthermore, the proposed network dynamically adjusts fusion weights between local and global features to enhance the network&#x00027;s feature extraction capability. Compared with other methods, the proposed method reconstructs high-quality images with a smaller number of parameters and lower computational complexity. Through the analysis of visual results, we have found that the proposed method has advantages over other comparison methods in restoring local details and global information of the image. Finally, experimental results on two remote-sensing datasets and five SR benchmark datasets demonstrate that our network can better achieve a balance between performance and model complexity.</p>
<p>To address the challenge of large model parameters that can hinder model deployment, we have developed a lightweight super-resolution reconstruction network that reduces computational complexity and model size while ensuring high-quality image reconstruction. In the future, we intend to investigate practical deployment techniques for lightweight super-resolution models, making them more compatible with lower-performance hardware devices, such as embedded and mobile devices. Additionally, there is still room for improvement in our model, particularly in real application scenarios. To improve the restoration quality of low-resolution images in real-world scenarios, we plan to explore the integration of blind super-resolution methods with supervised end-to-end training, aiming to design a model that can reconstruct super-resolution images in real-world situations.</p>
</sec>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>GP was responsible for paper scheme design, experiment, and paper writing. MX guided the paper scheme design, experiments, and wrote the papers. LF guided paper writing, revision, translation, and typesetting. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>This research was funded by the Yunnan Provincial Science and Technology Project (202205AG070008).</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ahn</surname> <given-names>N.</given-names></name> <name><surname>Kang</surname> <given-names>B.</given-names></name> <name><surname>Sohn</surname> <given-names>K.-A.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Fast, accurate, and lightweight super-resolution with cascading residual network,&#x0201D;</article-title> in <source>Proceedings of the European Conference on Computer Vision</source> (<publisher-loc>Milan</publisher-loc>: <publisher-name>ECCV</publisher-name>), <fpage>252</fpage>&#x02013;<lpage>268</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01249-6_16</pub-id></citation>
</ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bevilacqua</surname> <given-names>M.</given-names></name> <name><surname>Roumy</surname> <given-names>A.</given-names></name> <name><surname>Guillemot</surname> <given-names>C.</given-names></name> <name><surname>Morel</surname> <given-names>A. A.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Low-complexity single-image super-resolution based on nonnegative neighbor embedding,&#x0201D;</article-title> in <source>Proceedings of the British Machine Vision Conference</source>, eds <person-group person-group-type="editor"><name><surname>Bowden</surname> <given-names>R.</given-names></name> <name><surname>Collomosse</surname> <given-names>J.</given-names></name> <name><surname>Mikolajczyk</surname> <given-names>K.</given-names></name></person-group> (<publisher-name>BMVA Press</publisher-name>), <fpage>135.1</fpage>&#x02013;<lpage>135.10</lpage>. <pub-id pub-id-type="doi">10.5244/C.26.135</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Guo</surname> <given-names>T.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Deng</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Pre-trained image processing transformer,&#x0201D;</article-title> in <source>2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Nashville, TN</publisher-loc>), <fpage>12294</fpage>&#x02013;<lpage>12305</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01212</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>M.</given-names></name> <name><surname>Qian</surname> <given-names>Y.</given-names></name> <name><surname>Xiao</surname> <given-names>Z.</given-names></name> <name><surname>Zhong</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Remote sensing image super-resolution via residual aggregation and split attentional fusion network</article-title>. <source>IEEE J. Selec. Top. Appl. Earth Observ. Remote Sens</source>. <volume>14</volume>, <fpage>9546</fpage>&#x02013;<lpage>9556</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2021.3113658</pub-id></citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name> <name><surname>Dong</surname> <given-names>C.</given-names></name></person-group> (<year>2022</year>). <article-title>Activating more pixels in image super-resolution transformer</article-title>. <source>arXiv preprint arXiv:2205.04437</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2205.04437</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>C.</given-names></name> <name><surname>Loy</surname> <given-names>C. C.</given-names></name> <name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;Learning a deep convolutional network for image super-resolution,&#x0201D;</article-title> in <source>Computer Vision-ECCV 2014</source> (<publisher-loc>Zurich</publisher-loc>), <fpage>184</fpage>&#x02013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-10593-2_13</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>C.</given-names></name> <name><surname>Loy</surname> <given-names>C. C.</given-names></name> <name><surname>Tang</surname> <given-names>X.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Accelerating the super-resolution convolutional neural network,&#x0201D;</article-title> in <source>Computer Vision-ECCV 2016</source> (<publisher-loc>Amsterdam</publisher-loc>), <fpage>391</fpage>&#x02013;<lpage>407</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-46475-6_25</pub-id></citation>
</ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Wu</surname> <given-names>F.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name></person-group> (<year>2022a</year>). <article-title>&#x0201C;Feature distillation interaction weighting network for lightweight image super-resolution,&#x0201D;</article-title> in <source>Proceedings of the AAAI Conference on Artificial Intelligence</source> (<publisher-loc>Vancouver</publisher-loc>), <fpage>661</fpage>&#x02013;<lpage>669</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v36i1.19946</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gao</surname> <given-names>G.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Zeng</surname> <given-names>T.</given-names></name></person-group> (<year>2022b</year>). <article-title>Lightweight bimodal network for single-image super-resolution via symmetric CNN and recursive transformer</article-title>. <source>arXiv preprint arXiv:2204.13286</source>. <pub-id pub-id-type="doi">10.24963/ijcai.2022/128</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Georgescu</surname> <given-names>M.-I.</given-names></name> <name><surname>Ionescu</surname> <given-names>R. T.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>&#x0201C;Multimodal multi-head convolutional attention with various kernel sizes for medical image super-resolution,&#x0201D;</article-title> in <source>2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)</source> (<publisher-loc>Waikoloa, HI</publisher-loc>), <fpage>2194</fpage>&#x02013;<lpage>2204</lpage>. <pub-id pub-id-type="doi">10.1109/WACV56688.2023.00223</pub-id></citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.-B.</given-names></name> <name><surname>Singh</surname> <given-names>A.</given-names></name> <name><surname>Ahuja</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Single image super-resolution from transformed self-exemplars,&#x0201D;</article-title> in <source>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Boston, MA</publisher-loc>), <fpage>5197</fpage>&#x02013;<lpage>5206</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7299156</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Hui</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Gao</surname> <given-names>X.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Fast and accurate single image super-resolution via information distillation network,&#x0201D;</article-title> in <source>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Salt Lake City, UT</publisher-loc>), <fpage>723</fpage>&#x02013;<lpage>731</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00082</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>S.</given-names></name> <name><surname>Zhu</surname> <given-names>S.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Xu</surname> <given-names>M.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name></person-group> (<year>2023</year>). <article-title>Diffused convolutional neural network for hyperspectral image super-resolution</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>61</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3250640</pub-id></citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J. K.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2016a</year>). <article-title>&#x0201C;Accurate image super-resolution using very deep convolutional networks,&#x0201D;</article-title> in <source>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>1646</fpage>&#x02013;<lpage>1654</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.182</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>J. K.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2016b</year>). <article-title>&#x0201C;Deeply-recursive convolutional network for image super-resolution,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Las Vegas, NV</publisher-loc>), <fpage>1637</fpage>&#x02013;<lpage>1645</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.181</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lai</surname> <given-names>W.-S.</given-names></name> <name><surname>Huang</surname> <given-names>J.-B.</given-names></name> <name><surname>Ahuja</surname> <given-names>N.</given-names></name> <name><surname>Yang</surname> <given-names>M.-H.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Deep laplacian pyramid networks for fast and accurate super-resolution,&#x0201D;</article-title> in <source>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>624</fpage>&#x02013;<lpage>632</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.618</pub-id><pub-id pub-id-type="pmid">30106708</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lan</surname> <given-names>R.</given-names></name> <name><surname>Sun</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>H.</given-names></name> <name><surname>Pang</surname> <given-names>C.</given-names></name> <name><surname>Luo</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Madnet: a fast and lightweight network for single-image super resolution</article-title>. <source>IEEE Trans. Cybern</source>. <volume>51</volume>, <fpage>1443</fpage>&#x02013;<lpage>1453</lpage>. <pub-id pub-id-type="doi">10.1109/TCYB.2020.2970104</pub-id><pub-id pub-id-type="pmid">32149667</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lei</surname> <given-names>S.</given-names></name> <name><surname>Shi</surname> <given-names>Z.</given-names></name> <name><surname>Zou</surname> <given-names>Z.</given-names></name></person-group> (<year>2017</year>). <article-title>Super-resolution for remote sensing images via local&#x02013;global combined network</article-title>. <source>IEEE Geosci. Remote Sens. Lett</source>. <volume>14</volume>, <fpage>1243</fpage>&#x02013;<lpage>1247</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2017.2704122</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Cen</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2021</year>). <article-title>Different input resolutions and arbitrary output resolution: a meta learning-based deep framework for infrared and visible image fusion</article-title>. <source>IEEE Trans. Image Process</source>. <volume>30</volume>, <fpage>4070</fpage>&#x02013;<lpage>4083</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2021.3069339</pub-id><pub-id pub-id-type="pmid">33798086</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Tao</surname> <given-names>D.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name></person-group> (<year>2020</year>). <article-title>Attribute-aligned domain-invariant feature learning for unsupervised domain adaptation person re-identification</article-title>. <source>IEEE Trans. Inform. Forens. Sec</source>. <volume>16</volume>, <fpage>1480</fpage>&#x02013;<lpage>1494</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2020.3036800</pub-id><pub-id pub-id-type="pmid">33382653</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Gao</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>M.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2022a</year>). <article-title>Haze transfer and feature aggregation network for real-world single image dehazing</article-title>. <source>Knowl. Based Syst</source>. <volume>251</volume>:<fpage>109309</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109309</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name> <name><surname>Hu</surname> <given-names>Z.</given-names></name> <name><surname>Nie</surname> <given-names>F.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2023a</year>). <article-title>&#x0201C;Intermediary-guided bidirectional spatial-temporal aggregation network for video-based visible-infrared person re-identification,&#x0201D;</article-title> in <source>IEEE Transactions on Circuits and Systems for Video Technology (IEEE)</source>. <pub-id pub-id-type="doi">10.1109/TCSVT.2023.3246091</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>K.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2022b</year>). <article-title>Dual-stream reciprocal disentanglement learning for domain adaptation person re-identification</article-title>. <source>Knowl. Based Syst</source>. <volume>251</volume>:<fpage>109315</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109315</pub-id></citation>
</ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Mao</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). <article-title>Fractional differential and variational method for image fusion and super-resolution</article-title>. <source>Neurocomputing</source> <volume>171</volume>, <fpage>138</fpage>&#x02013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2015.06.035</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>G.</given-names></name></person-group> (<year>2023b</year>). <article-title>Feature dynamic alignment and refinement for infrared-visible image fusion: translation robust fusion</article-title>. <source>Inform. Fus</source>. <volume>95</volume>, <fpage>26</fpage>&#x02013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2023.02.011</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Li</surname> <given-names>F.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>Mutual prediction learning and mixed viewpoints for unsupervised-domain adaptation person re-identification on blockchain</article-title>. <source>Simul. Model. Pract. Theory</source> <volume>119</volume>:<fpage>102568</fpage>. <pub-id pub-id-type="doi">10.1016/j.simpat.2022.102568</pub-id></citation>
</ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>J.</given-names></name> <name><surname>Cao</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Van Gool</surname> <given-names>L.</given-names></name> <name><surname>Timofte</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>&#x0201C;Swinir: image restoration using swin transformer,&#x0201D;</article-title> in <source>2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>1833</fpage>&#x02013;<lpage>1844</lpage>. <pub-id pub-id-type="doi">10.1109/ICCVW54120.2021.00210</pub-id><pub-id pub-id-type="pmid">36633614</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lim</surname> <given-names>B.</given-names></name> <name><surname>Son</surname> <given-names>S.</given-names></name> <name><surname>Kim</surname> <given-names>H.</given-names></name> <name><surname>Nah</surname> <given-names>S.</given-names></name> <name><surname>Lee</surname> <given-names>K. M.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Enhanced deep residual networks for single image super-resolution,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>1132</fpage>&#x02013;<lpage>1140</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2017.151</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Xiong</surname> <given-names>Z.</given-names></name> <name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name></person-group> (<year>2023</year>). <article-title>Distilling knowledge from super resolution for efficient remote sensing salient object detection</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>57</volume>, <fpage>9791</fpage>&#x02013;<lpage>9809</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2023.3267271</pub-id></citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Feng</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zeng</surname> <given-names>T.</given-names></name></person-group> (<year>2023</year>). <article-title>Gradient prior dilated convolution network for remote sensing image super resolution</article-title>. <source>IEEE J. Selec. Top. Appl. Earth Observ. Remote Sens</source>. <volume>16</volume>, <fpage>3945</fpage>&#x02013;<lpage>3958</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2023.3252585</pub-id></citation>
</ref>
<ref id="B31">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>H.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>&#x0201C;Swin transformer: hierarchical vision transformer using shifted windows,&#x0201D;</article-title> in <source>2021 IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Montreal</publisher-loc>), <fpage>9992</fpage>&#x02013;<lpage>10002</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00986</pub-id></citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>Z.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Efficient transformer for single image super-resolution</article-title>. <source>arXiv preprint arXiv:2108.11084</source>. <pub-id pub-id-type="doi">10.1109/CVPRW56347.2022.00061</pub-id></citation>
</ref>
<ref id="B33">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>D.</given-names></name> <name><surname>Fowlkes</surname> <given-names>C.</given-names></name> <name><surname>Tal</surname> <given-names>D.</given-names></name> <name><surname>Malik</surname> <given-names>J.</given-names></name></person-group> (<year>2001</year>). <article-title>&#x0201C;A database of human segmented natural images and its application to evaluating segmentation algorithms and measuring ecological statistics,&#x0201D;</article-title> in <source>Proceedings Eighth IEEE International Conference on Computer Vision, ICCV 2001, Vol. 2</source> (<publisher-loc>Vancouver, BC</publisher-loc>), <fpage>416</fpage>&#x02013;<lpage>423</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2001.937655</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Matsui</surname> <given-names>Y.</given-names></name> <name><surname>Ito</surname> <given-names>K.</given-names></name> <name><surname>Aramaki</surname> <given-names>Y.</given-names></name> <name><surname>Fujimoto</surname> <given-names>A.</given-names></name> <name><surname>Ogawa</surname> <given-names>T.</given-names></name> <name><surname>Yamasaki</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Sketch-based manga retrieval using manga109 dataset</article-title>. <source>Multim. Tools Appl</source>. <volume>76</volume>, <fpage>21811</fpage>&#x02013;<lpage>21838</lpage>. <pub-id pub-id-type="doi">10.1007/s11042-016-4020-z</pub-id></citation>
</ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Mazur</surname> <given-names>N.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Malaviya</surname> <given-names>D.</given-names></name></person-group> (<year>2022</year>). <article-title>Small object detection method based on adaptive spatial parallel convolution and fast multi-scale fusion</article-title>. <source>Remote Sens</source>. <volume>14</volume>:<fpage>420</fpage>. <pub-id pub-id-type="doi">10.3390/rs14020420</pub-id></citation>
</ref>
<ref id="B36">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tai</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name></person-group> (<year>2017a</year>). <article-title>&#x0201C;Image super-resolution via deep recursive residual network,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>2790</fpage>&#x02013;<lpage>2798</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.298</pub-id></citation>
</ref>
<ref id="B37">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Tai</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name></person-group> (<year>2017b</year>). <article-title>&#x0201C;MemNet: a persistent memory network for image restoration,&#x0201D;</article-title> in <source>Proceedings of the IEEE International Conference on Computer Vision</source> (<publisher-loc>Venice</publisher-loc>), <fpage>4539</fpage>&#x02013;<lpage>4547</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.486</pub-id></citation>
</ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Yu</surname> <given-names>Z.</given-names></name></person-group> (<year>2023</year>). <article-title>Structure-embedded ghosting artifact suppression network for high dynamic range image reconstruction</article-title>. <source>Knowl. Based Syst</source>. <volume>263</volume>:<fpage>110278</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2023.110278</pub-id></citation>
</ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>C.</given-names></name> <name><surname>Zhuge</surname> <given-names>R.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Lightweight image super-resolution with enhanced cnn</article-title>. <source>Knowl. Based Syst</source>. <volume>205</volume>:<fpage>106235</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106235</pub-id></citation>
</ref>
<ref id="B40">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Timofte</surname> <given-names>R.</given-names></name> <name><surname>Agustsson</surname> <given-names>E.</given-names></name> <name><surname>Gool</surname> <given-names>L. V.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Ntire 2017 challenge on single image super-resolution: methods and results,&#x0201D;</article-title> in <source>2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</source> (<publisher-loc>Honolulu, HI</publisher-loc>), <fpage>1110</fpage>&#x02013;<lpage>1121</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2017.149</pub-id></citation>
</ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tu</surname> <given-names>J.</given-names></name> <name><surname>Mei</surname> <given-names>G.</given-names></name> <name><surname>Ma</surname> <given-names>Z.</given-names></name> <name><surname>Piccialli</surname> <given-names>F.</given-names></name></person-group> (<year>2022</year>). <article-title>Swcgan: Generative adversarial network combining swin transformer and CNN for remote sensing image super-resolution</article-title>. <source>IEEE J. Selec. Top. Appl. Earth Observ. Remote Sens</source>. <volume>15</volume>, <fpage>5662</fpage>&#x02013;<lpage>5673</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2022.3190322</pub-id></citation>
</ref>
<ref id="B42">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>&#x0201C;Attention is all you need,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 30</source> (<publisher-loc>Long Beach, CA</publisher-loc>).</citation>
</ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Yang</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Lu</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2023</year>). <article-title>MSAGAN: a new super-resolution algorithm for multispectral remote sensing image based on a multiscale attention GAN network</article-title>. <source>IEEE Geosci. Remote Sens. Lett</source>. <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/LGRS.2023.3258965</pub-id></citation>
</ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Bovik</surname> <given-names>A.</given-names></name> <name><surname>Sheikh</surname> <given-names>H.</given-names></name> <name><surname>Simoncelli</surname> <given-names>E.</given-names></name></person-group> (<year>2004</year>). <article-title>Image quality assessment: from error visibility to structural similarity</article-title>. <source>IEEE Trans. Image Process</source>. <volume>13</volume>, <fpage>600</fpage>&#x02013;<lpage>612</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2003.819861</pub-id><pub-id pub-id-type="pmid">15376593</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Xue</surname> <given-names>Y.</given-names></name> <name><surname>Jiang</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>K.</given-names></name> <name><surname>Ma</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>FeNEt: feature enhancement network for lightweight remote-sensing image super-resolution</article-title>. <source>IEEE Trans. Geosci. Remote Sens</source>. <volume>60</volume>, <fpage>1</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1109/TGRS.2022.3168787</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiao</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>F.</given-names></name> <name><surname>Jin</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>Heterogeneous knowledge distillation for simultaneous infrared-visible image fusion and super-resolution</article-title>. <source>IEEE Trans. Instrum. Measure</source>. 71, 5004015. <pub-id pub-id-type="doi">10.1109/TIM.2022.3149101</pub-id></citation>
</ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>P.</given-names></name> <name><surname>Tang</surname> <given-names>H.</given-names></name> <name><surname>Ge</surname> <given-names>J.</given-names></name> <name><surname>Feng</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>Espc_nasunet: an end-to-end super-resolution semantic segmentation network for mapping buildings from remote sensing images</article-title>. <source>IEEE J. Selec. Top. Appl. Earth Observ. Remote Sens</source>. <volume>14</volume>, <fpage>5421</fpage>&#x02013;<lpage>5435</lpage>. <pub-id pub-id-type="doi">10.1109/JSTARS.2021.3079459</pub-id></citation>
</ref>
<ref id="B48">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yan</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Xie</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name> <name><surname>ZhengtaoYu</surname></name></person-group> (<year>2022</year>). <article-title>Cross-domain person re-identification with pose-invariant feature decomposition and hypergraph structure alignment</article-title>. <source>Neurocomputing</source> <volume>467</volume>, <fpage>229</fpage>&#x02013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2021.09.054</pub-id></citation>
</ref>
<ref id="B49">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zeyde</surname> <given-names>R.</given-names></name> <name><surname>Elad</surname> <given-names>M.</given-names></name> <name><surname>Protter</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;On single image scale-up using sparse-representations,&#x0201D;</article-title> in <source>Curves and Surfaces: 7th International Conference</source> (<publisher-loc>Avignon</publisher-loc>), <fpage>711</fpage>&#x02013;<lpage>730</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-642-27413-8_47</pub-id></citation>
</ref>
<ref id="B50">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Cross-compatible embedding and semantic consistent feature construction for sketch re-identification,&#x0201D;</article-title> in <source>Proceedings of the 30th ACM International Conference on Multimedia (MM&#x00027;22)</source> (<publisher-loc>Lisbon</publisher-loc>), <fpage>3347</fpage>&#x02013;<lpage>3355</lpage>. <pub-id pub-id-type="doi">10.1145/3503161.3548224</pub-id></citation>
</ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>M.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wei</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Image dehazing by an artificial image fusion method based on adaptive structure decomposition</article-title>. <source>IEEE Sensors J</source>. <volume>20</volume>, <fpage>8062</fpage>&#x02013;<lpage>8072</lpage>. <pub-id pub-id-type="doi">10.1109/JSEN.2020.2981719</pub-id></citation>
</ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Meng</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Mazur</surname> <given-names>N.</given-names></name></person-group> (<year>2021a</year>). <article-title>Remote sensing image defogging networks based on dual self-attention boost residual octave convolution</article-title>. <source>Remote Sens</source>. 13, 3104. <pub-id pub-id-type="doi">10.3390/rs13163104</pub-id></citation>
</ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Wei</surname> <given-names>H.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Mazur</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2021b</year>). <article-title>Atmospheric light estimation based remote sensing image dehazing</article-title>. <source>Remote Sens</source>. 13, 2432. <pub-id pub-id-type="doi">10.3390/rs13132432</pub-id><pub-id pub-id-type="pmid">34329312</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Wei</surname> <given-names>H.</given-names></name> <name><surname>Hu</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Qi</surname> <given-names>G.</given-names></name> <name><surname>Mazur</surname> <given-names>N.</given-names></name></person-group> (<year>2021c</year>). <article-title>A novel fast single image dehazing algorithm based on artificial multiexposure image fusion</article-title>. <source>IEEE Trans. Instrum. Measure</source>. <volume>70</volume>, <fpage>1</fpage>&#x02013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/TIM.2020.3024335</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>