<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mar. Sci.</journal-id>
<journal-title>Frontiers in Marine Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mar. Sci.</abbrev-journal-title>
<issn pub-type="epub">2296-7745</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmars.2024.1382147</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Marine Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Learning degradation-aware visual prompt for maritime image restoration under adverse weather conditions</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>He</surname><given-names>Xin</given-names>
</name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>*</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/2612506"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/formal-analysis/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Jia</surname><given-names>Tong</given-names>
</name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/visualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname><given-names>Junjie</given-names>
</name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Basic Sciences for Aviation, Naval Aviation University</institution>, <addr-line>Yantai</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Art and Design, Yantai Institute of Science and Technology</institution>, <addr-line>Yantai</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Electromechanical and Automotive Engineering, Yantai University</institution>, <addr-line>Yantai</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Xinyu Zhang, Dalian Maritime University, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Man Zhou, Nanyang Technological University, Singapore</p>
<p>Pengpeng Li, Nanjing University of Science and Technology, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Xin He, <email xlink:href="mailto:hexin6770@163.com">hexin6770@163.com</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>09</day>
<month>04</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>11</volume>
<elocation-id>1382147</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>02</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>03</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2024 He, Jia and Li</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>He, Jia and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Adverse weather conditions such as rain and haze often lead to a degradation in the quality of maritime images, which is crucial for activities like navigation, fishing, and search and rescue. Therefore, it is of great interest to develop an effective algorithm to recover high-quality maritime images under adverse weather conditions. This paper proposes a prompt-based learning method with degradation perception for maritime image restoration, which contains two key components: a restoration module and a prompting module. The former is employed for image restoration, whereas the latter encodes weather-related degradation-specific information to modulate the restoration module, enhancing the recovery process for improved results. Inspired by the recent trend of prompt learning in artificial intelligence, this paper adopts soft-prompt technology to generate learnable visual prompt parameters for better perceiving the degradation-conditioned cues. Extensive experimental results on several benchmarks show that our approach achieves superior restoration performance in maritime image dehazing and deraining tasks.</p>
</abstract>
<kwd-group>
<kwd>maritime image</kwd>
<kwd>image restoration</kwd>
<kwd>image deraining</kwd>
<kwd>image dehazing</kwd>
<kwd>prompt learning</kwd>
<kwd>deep learning</kwd>
<kwd>visual transformer</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="5"/>
<equation-count count="12"/>
<ref-count count="51"/>
<page-count count="13"/>
<word-count count="5836"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Ocean Observation</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Adverse weather conditions, including rain and haze, frequently occur in our everyday environment. These conditions result in diminished visual quality in captured images and significantly affect the effectiveness of numerous maritime vision systems, such as autonomous ships for ocean observation (<xref ref-type="bibr" rid="B48">Zheng et&#xa0;al., 2024</xref>). In maritime navigation and transportation, correctly identifying and interpreting environmental information from images is vital for safety. <xref ref-type="fig" rid="f1"><bold>Figure&#xa0;1</bold></xref> shows the physical imaging process of different adverse weather conditions. Thus, image processing under adverse weather conditions contributes to enhancing the safety of maritime traffic and navigation by reducing accidents and collisions (<xref ref-type="bibr" rid="B29">Lu et&#xa0;al., 2021</xref>).</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>The physical imaging process of different adverse weather conditions.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g001.tif"/>
</fig>
<p>To solve image restoration under adverse weather conditions, early algorithms are predominantly based on traditional prior models. In the context of image dehazing, one common approach was the atmospheric scattering model (<xref ref-type="bibr" rid="B12">He et&#xa0;al., 2010</xref>), which assumed that haze in an image could be represented as a result of light scattering due to atmospheric particles (<xref ref-type="bibr" rid="B23">Li et&#xa0;al., 2018a</xref>). These algorithms typically aim to estimate and remove the haze from images, enhancing visibility. On the other hand, for image deraining, a prevalent technique was the linear superposition model. This model assumed that the observed image under rainy conditions could be expressed as a linear combination of the clean background scene and the rain streaks (<xref ref-type="bibr" rid="B4">Chen et&#xa0;al., 2023b</xref>). These early deraining algorithms focus on separating the rain streaks from the desired scene, thus improving the clarity of the image. However, these early prior-based algorithms struggled to adapt to complex and rapidly changing scenes, as they relied heavily on predefined models that could not effectively account for the wide range of scenarios encountered in real-world environments.</p>
<p>With the rise of big data and artificial intelligence, a plethora of image restoration methods based on deep learning have emerged. These techniques aim to learn the mapping relationship between degraded images and their corresponding clear counterparts. Convolutional Neural Networks (CNNs) have emerged as a powerful tool for image restoration due to their inherent ability to capture and learn complex hierarchical features from data. We have witnessed the rapid advancement of CNNs in image dehazing and deraining (<xref ref-type="bibr" rid="B24">Li et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B50">Zhou et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B5">Chen et&#xa0;al., 2022</xref>). However, due to the inherent characteristics of convolution operations, specifically the use of local receptive fields and the independence of input content, CNNs struggle to effectively model spatially-long feature dependencies of images (<xref ref-type="bibr" rid="B6">Chen et&#xa0;al., 2023c</xref>).</p>
<p>Later, Transformer-based models (<xref ref-type="bibr" rid="B35">Vaswani et&#xa0;al., 2017</xref>) originally bring significant breakthroughs to the natural language processing (NLP) field. The vision Transformer (ViT), as a new network backbone, has been widely applied to various tasks. It has also been utilized in image restoration tasks <xref ref-type="bibr" rid="B43">Zamir et&#xa0;al. (2022)</xref> and has achieved better performance compared to CNNs due to its ability to model non-local features effectively. Albeit these approaches have achieved commendable restoration performance in the given weather situation, they often exhibit suboptimal results when applied to maritime images.</p>
<p>The reason behind this can be summarized as follows: (1) Maritime images are often captured over large bodies of water, which introduces additional complexities due to the presence of reflective surfaces, varying water conditions, and dynamic backgrounds. These factors can exacerbate the impact of weather-related degradation. (2) Weather conditions at sea can change rapidly, with haze and rain appearing and dissipating quickly. This dynamic nature poses challenges for image restoration, as algorithms must adapt to evolving weather conditions. Thus, effective image restoration techniques tailored to the unique characteristics of maritime environments are essential for ensuring safe and efficient maritime activities (<xref ref-type="bibr" rid="B47">Zheng et&#xa0;al., 2020</xref>).</p>
<p>This raises a question: how to better help image restoration in adapting to the complex and ever-changing maritime scenes under adverse weather conditions? Recent trends of prompt learning (<xref ref-type="bibr" rid="B37">Wang et&#xa0;al., 2022</xref>) in artificial intelligence, may offer a potential solution. Prompt learning empowers deep models to adapt swiftly to complex and dynamically changing environments. It allows for the creation of tailored prompts that can capture the intricacies of specific situations, ensuring the model&#x2019;s responsiveness to various challenges. Therefore, this motivates us to introduce prompt learning to better encode degradation features of different weather conditions. This paper proposes a prompt-based learning method with degradation perception for maritime image restoration. The proposed method comprises two essential components: a restoration module and a prompting module. The restoration module is employed for image restoration, while the prompting module encodes weather-related degradation-specific information to modulate the restoration module. Specifically, the main contributions of this paper are as follows:</p>
<list list-type="bullet">
<list-item>
<p>This paper presents a new solution for image restoration in adverse weather conditions for maritime images. By incorporating prompt learning into the Transformer-based restoration network, it enhance the adaptability of deep networks to various weather degradation characteristics, enabling our model to adaptively learn more useful features to facilitate better restoration.</p>
</list-item>
<list-item>
<p>This paper employs a prompt creation block to generate a set of learnable parameters by implicitly predicting degradation-conditioned soft prompts. In addition, this paper further introduces a prompt fusion block to guide the restoration process by interacting with the network backbone.</p>
</list-item>
<list-item>
<p>Quantitative and qualitative experiments demonstrate that our proposed method achieves favorable performance on multiple benchmark datasets, and can better reconstruct clear images and restore image details compared to previous methods.</p>
</list-item>
</list>
</sec>
<sec id="s2">
<label>2</label>
<title>Related work</title>
<p>In this section, this paper presents a review of recent work related to maritime image restoration and prompt learning.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Maritime image restoration</title>
<p>To deal with the uncertain task of maritime image restoration, considerable efforts have been made. Existing approaches can be categorized into strategies based on priors and learning-based strategies. Hu et&#xa0;al (<xref ref-type="bibr" rid="B14">Hu et&#xa0;al., 2019</xref>). proposed a haze removal method based on illumination decomposition. This method decomposes the hazy image into a haze layer by separating the glow layer. It estimates the transmission rate using haze-line prior, thereby restoring the haze-free image. Luo et&#xa0;al (<xref ref-type="bibr" rid="B29">Lu et&#xa0;al., 2021</xref>). proposed a novel CNN-based visibility dehazing framework aimed at enhancing the visual quality of images captured by maritime cameras under hazy conditions. This framework comprises two subnetworks: the coarse feature extraction module and the fine feature fusion module. Hu et&#xa0;al. (<xref ref-type="bibr" rid="B15">Hu et&#xa0;al., 2021</xref>) proposed a deep learning-based variational optimization method for reconstructing haze-free images from observed hazy images. This method fully leverages a unified denoising framework and strong deep learning representation capabilities. Guo et&#xa0;al. (<xref ref-type="bibr" rid="B11">Guo et&#xa0;al., 2021</xref>) designed a heterogeneous twin birth haze removal network, HTDNet, to enhance maritime surveillance capabilities in haze environments. The network consists of a twin feature extraction module for learning coarse haze features and a feature fusion module for integration and enhancement.</p>
<p>Van et&#xa0;al. (<xref ref-type="bibr" rid="B34">Van Nguyen et&#xa0;al., 2021</xref>) proposed a haze removal algorithm for maritime environment images based on texture and structure priors in illumination decomposition. This method utilizes a haze removal algorithm to eliminate the haze component from the glow-free layer and employs illumination compensation to restore natural illumination in the glow layer. Yang et&#xa0;al. (<xref ref-type="bibr" rid="B40">Yang et&#xa0;al., 2022</xref>) proposed a multi-head pyramid large kernel encoder-decoder network (LKEDN-MHP) for denoising tasks in maritime images. This method utilizes the transmission map extracted from the guidance image as an additional input to improve the network performance. Liu et&#xa0;al. (<xref ref-type="bibr" rid="B27">Liu et&#xa0;al., 2022</xref>) proposed a CNN-based dual-channel two-stage image dehazing network, which utilizes an attention mechanism to achieve adaptive fusion of multi-channel features. Hu et&#xa0;al. (<xref ref-type="bibr" rid="B16">Hu et&#xa0;al., 2022</xref>) proposed a maritime video dehazing algorithm based on spatiotemporal information fusion and improved dark channel prior. This method utilizes an enhanced dark channel prior model to restore each frame image, thereby achieving video dehazing.</p>
<p>Recently, Huang et&#xa0;al. (<xref ref-type="bibr" rid="B17">Huang et&#xa0;al., 2023</xref>) proposed an improved convex optimization model based on an atmospheric scattering model to achieve image dehazing. This method integrates simplified atmospheric light value estimation and the V channel in the HSV color space to obtain more local information. He et&#xa0;al. (<xref ref-type="bibr" rid="B13">He and Ji, 2023</xref>) improved MID-GAN is capable of training with non-paired adversarial learning. It consists of a CycleGAN cycle framework with two constraint branches. And it introduced an effective attentionrecursive feature extraction module to gradually extract haze components in an unsupervised manner. Chen et&#xa0;al. (<xref ref-type="bibr" rid="B5">Chen et&#xa0;al., 2022</xref>) introduced a contrastive learning mechanism based on the CycleGAN framework to improve dehazing performance. However, due to the limited performance of the aforementioned methods in the task of maritime image restoration and the relative saturation of model capabilities, there is a need to explore an effective approach to address these issues.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Prompt learning</title>
<p>Prompt learning was initially introduced in the field of natural language processing (NLP) and has proven to be highly effective, it has been applied to various vision-related tasks. Prompt learning is divided into two different methods: hard prompts and soft prompts. Hard prompts refer to explicit and predefined instructions given to the model during training. These prompts provide specific information and guide the model to produce the desired output. Soft prompts are more flexible and adaptive. They are not explicitly defined but rather generate prompt information based on input data or learn from the training process. Soft prompts allow the model to dynamically adjust its behavior based on input and context, enabling it to capture more specific and nuanced contextual information.</p>
<p>Recently, Zhou et&#xa0;al. (<xref ref-type="bibr" rid="B51">Zhou et&#xa0;al., 2022</xref>) demonstrated that a simple design based on conditional prompt learning performs exceptionally well in various problem scenarios, including generalization from base classes to novel classes, cross-dataset prompt transfer, and domain generalization. Potlapalli et&#xa0;al. (<xref ref-type="bibr" rid="B32">Potlapalli et&#xa0;al., 2023</xref>) demonstrated the effectiveness of their designed prompt block in integrated image restoration by integrating it into state-of-the-art restoration models. The prompt block can interact with input features, dynamically adjust representations, and adapt the restoration process to the relevant degradation. Li et&#xa0;al. (<xref ref-type="bibr" rid="B20">Li et&#xa0;al., 2023b</xref>) proposed a novel prompt-in-prompt learning for universal image restoration. The method involves simultaneous learning of high-level degradation-aware prompts and low-level basic restoration prompts to generate effective universal restoration prompts. By utilizing a selective promptfeature interaction module to modulate features most relevant to the degradation. Ai et&#xa0;al. (<xref ref-type="bibr" rid="B1">Ai et&#xa0;al., 2023</xref>) proposed a multi-modal prompt learning method called MPerceiver, which includes cross-modal adapters and image restoration adapters to learn holistic and multiscale detail representations. The adaptability of text and visual prompts is dynamically adjusted based on degradation prediction, enabling effective adaptation to various unknown degradations. Kong et&#xa0;al. (<xref ref-type="bibr" rid="B19">Kong et&#xa0;al., 2024</xref>) proposed sequential learning strategy and prompt learning strategy, respectively. These two strategies are effective for both CNN and Transformer backbones, and they can complement each other to learn effective image representations.</p>
<p>Inspired by these methods, this paper proposes a prompt-based learning approach to guide the maritime image restoration process, facilitating the integration and communication of information.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed method</title>
<p>In this section, this paper first describes the overall pipeline of the model. Then, this paper provides details of the restoration module and prompting module, which serve as the fundamental building blocks of the approach. The restoration module mainly comprises two key elements: multi-head self-attention (MHSA) and dual gated feed-forward network (DGFN). The prompting module mainly consists of two key elements: prompt creation block (PCB) and prompt fusion block (PFB).</p>
<sec id="s3_1">
<label>3.1</label>
<title>Overall pipeline</title>
<p>The overall pipeline of the proposed model, as illustrated in the <xref ref-type="fig" rid="f2"><bold>Figure&#xa0;2</bold></xref>, is based on a hierarchical encoder-decoder framework (<xref ref-type="bibr" rid="B3">Chen et&#xa0;al., 2023a</xref>). Given a maritime degraded image <italic>I</italic><sub>rain</sub> &#x2208; &#x211d;<italic><sup>H</sup>
</italic><sup>&#xd7;</sup><italic><sup>W</sup>
</italic><sup>&#xd7;3</sup>, where <italic>H</italic> &#xd7; <italic>W</italic> denotes the spatial resolution of the feature maps, and <italic>C</italic> represents the channels, this paper performs feature projection embedding using a 3 &#xd7; 3 convolution. On the network backbone, this paper stacks 4 levels of hierarchical encoder-decoders, where the encoder-decoder serves as the restoration module of the model, extracting rich spatially variant degradation distribution features. To extract multiscale representations from degradation information, each level of the restoration module covers its specific spatial resolution and channel dimensions. Beginning with high-resolution input, the restoration module aims to progressively decrease spatial resolution while enhancing channel capacity, resulting in a low-resolution latent representation <italic>F</italic> &#x2208; &#x211d;<italic><sup>H/</sup>
</italic><sup>8&#xd7;</sup><italic><sup>W/</sup>
</italic><sup>8&#xd7;8</sup><italic><sup>C</sup>
</italic>. During the stage of high-resolution image restoration, this paper incorporates a prompting module into the framework to generate prompts and enrich input features for dynamically guiding the restoration process of the restoration module. This paper also introduces skip connections (<xref ref-type="bibr" rid="B21">Li et&#xa0;al., 2023a</xref>) to bridge consecutive intermediate features, ensuring stable training. Next, this paper provides a detailed description of the proposed restoration module, prompting module, and their core building blocks.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>The overall architecture of the proposed network.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g002.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Restoration module</title>
<p>This paper develops a restoration module as a feature extraction unit, which can be used to encode degradation information to recover output clean restored images. Formally, given the input features of the (<italic>l</italic> &#x2212; 1)-th block <italic>X<sub>l</sub>
</italic><sub>&#x2212;1</sub>, the encoding of the restoration module process can be represented as <xref ref-type="disp-formula" rid="eq1">Equations (1</xref>, <xref ref-type="disp-formula" rid="eq2">2)</xref>:</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>'</mml:mo>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>H</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>'</mml:mo>
</mml:msubsup>
<mml:mo>+</mml:mo>
<mml:mi>D</mml:mi>
<mml:mi>G</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>'</mml:mo>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>LN</italic> denotes the layer normalization, <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msubsup>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>'</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represent the outputs of MHSA and DGFN, which are described below.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Multi-head self-attention</title>
<p>In reviewing the standard self-attention mechanism in Transformers (<xref ref-type="bibr" rid="B43">Zamir et&#xa0;al., 2022</xref>), given queries <italic>Q</italic>, keys <italic>K</italic>, and values <italic>V</italic>, the output of dot-product attention is typically represented as <xref ref-type="disp-formula" rid="eq3">Equation (3)</xref>:</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:msup>
<mml:mi>n</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>Q</mml:mi>
<mml:mo>&#x22a4;</mml:mo>
</mml:msup>
<mml:mi>K</mml:mi>
</mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xb7;</mml:mo>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>&#x3b1;</italic> is a learnable parameter; <italic>Q</italic>, <italic>K</italic>, and <italic>V</italic> represent the matrix forms of <italic>Q</italic>, <italic>K</italic>, and <italic>V</italic>, respectively. It is noted that computing self-attention using the Softmax function may lead to unstable gradients due to the presence of exponential functions, which could also limit the network&#x2019;s ability for nonlinear fitting. This work replaces it with the ReLU activation function, which can alleviate the issues of gradient vanishing or exploding, and aid in learning better feature representations. Specifically, this work starts by aggregating pixel-level cross-channel context through the application of a 1 &#xd7; 1 convolution. Subsequently, a 3 &#xd7; 3 depthwise convolution is applied to encode channel-wise context. This work employs bias-free convolution layers in the network. Next, it reshapes the projections of queries and keys to allow their dot product interaction to generate a transposed attention map of size <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>C</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>&#xd7;</mml:mo>
<mml:mover accent="true">
<mml:mi>C</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, rather than a massive regular attention map of size <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>H</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mover accent="true">
<mml:mi>W</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>&#xd7;</mml:mo>
<mml:mover accent="true">
<mml:mi>H</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mover accent="true">
<mml:mi>W</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. Then, the attention map is further interacted with the reshaped projections of values to complete the self-attention computation. Overall, the MHSA process is defined as <xref ref-type="disp-formula" rid="eq4">Equations (4</xref>, <xref ref-type="disp-formula" rid="eq5">5)</xref>:</p>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>U</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>Q</mml:mi>
<mml:mo>&#x22a4;</mml:mo>
</mml:msup>
<mml:mi>K</mml:mi>
</mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xb7;</mml:mo>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq5">
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>X</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>X</italic> and <inline-formula>
<mml:math display="inline" id="im5">
<mml:mover accent="true">
<mml:mi>X</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> are the input and output feature maps, <italic>Conv</italic><sub>1&#xd7;1</sub>(&#xb7;) denotes 1 &#xd7; 1 convolution.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Dual gated feed-forward network</title>
<p>To enhance the enrichment of contextual information, this paper introduces a dual gated feed-forward network that operates on each pixel. It incorporates two branches based on a gating mechanism. They initially undergo feature <italic>Y</italic> transformation by using 1 &#xd7; 1 convolutions, followed by 3 &#xd7; 3 depth-wise convolutions to encode information from spatially adjacent pixel positions, facilitating the learning of local image details for effective restoration. One branch extends the feature channels, while the other branch, activated with GELU non-linearity, reduces the channels back to the original input dimensionality, enabling the discovery of non-linear contextual information in hidden layers. The DGFN is formulated as <xref ref-type="disp-formula" rid="eq6">Equations (6</xref>, <xref ref-type="disp-formula" rid="eq7">7)</xref>:</p>
<disp-formula id="eq6">
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>G</mml:mi>
<mml:mi>E</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>U</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2299;</mml:mo>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq7">
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mtext>Con</mml:mtext>
<mml:msub>
<mml:mtext>v</mml:mtext>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>Y</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>, <italic>Conv</italic><sub>3&#xd7;3</sub>(&#xb7;), &#x2299; denote outputs, 3 &#xd7; 3 depthwise convolution and element-wise multiplication, respectively. Overall, compared to MHSA, DGFN plays a distinctly different role by governing the flow of information across various levels in our pipeline, thereby enabling each level to focus on fine details complementary to other levels.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Prompting module</title>
<p>Different from (<xref ref-type="bibr" rid="B49">Zhou et&#xa0;al., 2023</xref>) that models global features, this paper further proposes a prompting module designed to perceive features of interfering information in degraded images and dynamically generate valuable prompts to guide high-quality maritime image restoration. Given the input features <italic>F</italic>, the prompting module first employs PCB to generate prompts for the distribution of degraded features. Subsequently, PFB collaboratively fuses the input features <italic>F</italic> with the generated prompt information to obtain the output features <italic>F</italic><sub>b</sub>. The overall procedure of prompting module is defined as <xref ref-type="disp-formula" rid="eq8">Equation (8)</xref>:</p>
<disp-formula id="eq8">
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>F</mml:mi>
<mml:mi>B</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>B</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>F<sub>pc</sub>
</italic> represents the learnable prompt components, the prompt creation block and prompt fusion block are described below.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Prompt creation block</title>
<p>The PCB dynamically captures the distribution characteristics in degraded maritime images. This capability enhances the provision of useful prompt information for the restoration process. Here, this paper employs a soft prompt (<xref ref-type="bibr" rid="B32">Potlapalli et&#xa0;al., 2023</xref>) to generate a set of learnable parameters, which encode distinctive deteriorative features related to various weather conditions. For input features <italic>F</italic>, the network first applies global average pooling, followed by a 1 &#xd7; 1 convolution to obtain a compact feature vector. Subsequently, a Softmax function is applied to derive the prompt weights <italic>W</italic> &#x2208; &#x211d;<italic><sub>N</sub>
</italic>, where the value of <italic>N</italic> is determined by the prompt length. In soft prompt learning, the length of the prompt determines the amount of information and guidance provided to the model. In fact, it&#x2019;s essential to strike a balance between the impact of prompt length, choosing an appropriate length that balances information, guidance, and diversity of generation to achieve satisfactory results. This paper will analyzes its impact in Section 4.4.2. Next, the weights are adjusted within the prompting components using a linear combination. Finally, a 3 &#xd7; 3 convolution is applied to obtain the conditional prompt <italic>P</italic>. The overall procedure of PCB is defined as <xref ref-type="disp-formula" rid="eq9">Equations (9</xref>, <xref ref-type="disp-formula" rid="eq10">10)</xref>:</p>
<disp-formula id="eq9">
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq10">
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#xb7;</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>GAP</italic>(&#xb7;), <italic>Linear</italic>(&#xb7;) denote the global average pooling and linear combination, respectively.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Prompt fusion block</title>
<p>To facilitate more effective interaction between the prompt information and input features for better guidance in the restoration process, this paper designs a PFB. In this module, the input features <italic>F</italic> are concatenated with the degradation distribution prompt <italic>P</italic> along the channel dimension. Subsequently, the concatenated information is further processed by the restoration module to generate the transformation of the input features. Finally, through operations of 1 &#xd7; 1 convolution and 3 &#xd7; 3 convolution, the features are smoothed and mapped to the output features <inline-formula>
<mml:math display="inline" id="im7">
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
</mml:math>
</inline-formula>. The overall procedure of PFB is defined as <xref ref-type="disp-formula" rid="eq11">Equation (11)</xref>:</p>
<disp-formula id="eq11">
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>F</mml:mi>
<mml:mo>^</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>Conv</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>&#x211b;</mml:mi>
<mml:mi>&#x2133;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#xb7;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>&#x211b;</mml:mi>
<mml:mi>&#x2133;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>(&#xb7;) and <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mi mathvariant="script">C</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>(&#xb7;) denote the process of involving restoration module and concatenation, respectively.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Loss function</title>
<p>To supervise the training progress of our network, this paper employs the <italic>L</italic><sub>1</sub> pixel loss function. The final output is the restored image <italic>I<sub>rec</sub>
</italic>. It is obtained by adding the residual image <italic>I<sub>res</sub>
</italic> to the input degraded image <italic>I<sub>deg</sub>
</italic>, where <italic>I<sub>res</sub>
</italic> &#x2208; &#x211d;<italic><sup>H</sup>
</italic><sup>&#xd7;</sup><italic><sup>W</sup>
</italic><sup>&#xd7;3</sup>. During training, the network minimizes the loss function ,which is defined as <xref ref-type="disp-formula" rid="eq12">Equation (12)</xref>:</p>
<disp-formula id="eq12">
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>I<sub>gt</sub>
</italic> represents the ground truth image, and &#x2225; &#xb7; &#x2225;<sub>1</sub> denotes the <italic>L</italic><sub>1</sub>-norm.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<p>In this section, this paper evaluates our method on benchmarks for image dehazing and image deraining tasks. The main experiments are conducted using PyTorch and trained on 4 TESLA V100 GPUs.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental settings</title>
<p>Following (<xref ref-type="bibr" rid="B43">Zamir et&#xa0;al., 2022</xref>), our method employs a 4-level encoder-decoder framework. From level-1 to level-4, the number of restoration module is [4,6,6,8], attention heads are [1,2,4,8], and number of channels is [48,96,192,384]. The prompt length in the prompting module is 5. The batch size and patch size are configured as 16 and 128, respectively. Our model is trained using the AdamW optimizer for a total of 300,000 iterations with cosine annealing scheme (<xref ref-type="bibr" rid="B28">Loshchilov and Hutter, 2016</xref>) to gradually decrease the initial learning rate from 3&#xd7;10<sup>&#x2212;4</sup> to 1&#xd7;10<sup>&#x2212;6</sup>. Specifically, the learning rate is initially set at 3&#xd7;10<sup>&#x2212;4</sup> for the 92,000 iterations and subsequently reduced to 1 &#xd7; 10<sup>&#x2212;6</sup> over the next 208,000 iterations.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental results on image dehazing</title>
<p>To evaluate the dehazing performance of the method, this paper trains on the RESIDE SOTS-Outdoor (<xref ref-type="bibr" rid="B23">Li et&#xa0;al., 2018a</xref>), but since this dataset lacks maritime scenes, this paper conducts testing on real maritime hazy images. It should be noted that there are no ground truth images for real maritime hazy images. Here, this paper compares our method with 6 popular image dehazing algorithms, including DCP (<xref ref-type="bibr" rid="B12">He et&#xa0;al., 2010</xref>), DehazeNet (<xref ref-type="bibr" rid="B2">Cai et&#xa0;al., 2016</xref>), AODNet (<xref ref-type="bibr" rid="B22">Li et&#xa0;al., 2017</xref>), GridDehazeNet <xref ref-type="bibr" rid="B26">Liu et&#xa0;al. (2019)</xref>, MSBDN (<xref ref-type="bibr" rid="B7">Dong et&#xa0;al., 2020</xref>), and DeHamer (<xref ref-type="bibr" rid="B10">Guo et&#xa0;al., 2022</xref>). For fair comparison, all comparison algorithms use consistent pre-trained weights trained on the training set. In the case of hazy maritime images without ground truth data, this paper employs non-reference metrics such as NIQE (Naturalness Image Quality Evaluator) (<xref ref-type="bibr" rid="B31">Mittal et&#xa0;al., 2012b</xref>) and BRISQUE (Blind/Referenceless Image Spatial Quality Evaluator) (<xref ref-type="bibr" rid="B30">Mittal et&#xa0;al., 2012a</xref>). When the NIQE or BRISQUE scores are lower, it means that the image is considered to have higher quality in terms of naturalness (for NIQE) or overall spatial quality (for BRISQUE).</p>
<p>As presented in <xref ref-type="table" rid="T1"><bold>Table&#xa0;1</bold></xref>, our proposed method achieves notably lower NIQE and BRISQUE scores, indicating that it produces high-quality outputs characterized by clearer content and superior perceptual quality when compared to other models in maritime scenarios. This underscores the effectiveness of our model in enhancing image quality within maritime contexts. To provide compelling evidence, this paper illustrates a visual quality comparison between two samples generated by recent methods in <xref ref-type="fig" rid="f3"><bold>Figure&#xa0;3</bold></xref>. It is observed that the performance of DCP (<xref ref-type="bibr" rid="B12">He et&#xa0;al., 2010</xref>) is suboptimal, particularly in the sky region, where undesirable halo effects occur. DehazeNet (<xref ref-type="bibr" rid="B2">Cai et&#xa0;al., 2016</xref>) and AODNet (<xref ref-type="bibr" rid="B22">Li et&#xa0;al., 2017</xref>), are found to exhibit insufficient learning capabilities, resulting in their inability to effectively remove haze from images. The limitations in their learning capabilities become particularly evident when dealing with challenging and complex hazy scenes, such as those with dense fog or severe haze. The dehazing results obtained from GridDehazeNet <xref ref-type="bibr" rid="B26">Liu et&#xa0;al. (2019)</xref>, MSBDN (<xref ref-type="bibr" rid="B7">Dong et&#xa0;al., 2020</xref>), and DeHamer (<xref ref-type="bibr" rid="B10">Guo et&#xa0;al., 2022</xref>) still exhibit residual haze, indicating limitations in the generalization of their algorithms to maritime images. These residual haziness issues suggest that their models may struggle to effectively adapt to the unique challenges posed by maritime scenarios. In contrast, our method demonstrates the capability to recover significantly clearer images, particularly in the sailboat regions. This suggests that the introduction of prompt learning cues can substantially improve the adaptation of dehazing algorithms to more challenging maritime images. The enhanced performance underscores the effectiveness of incorporating prompt learning, which enables better haze removal and results in visually superior outcomes in maritime scenes.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Quantitative comparisons of different methods on hazy maritime images.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="bottom" align="center">Method</th>
<th valign="bottom" align="center">DCP</th>
<th valign="bottom" align="center">DehazeNet</th>
<th valign="bottom" align="center">AODNet</th>
<th valign="bottom" align="center">GridDehazeNet</th>
<th valign="bottom" align="center">FFA Net</th>
<th valign="bottom" align="center">DeHamer</th>
<th valign="bottom" align="center">Ours</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">NIQE</td>
<td valign="bottom" align="center">3.464</td>
<td valign="bottom" align="center">3.466</td>
<td valign="bottom" align="center">3.433</td>
<td valign="bottom" align="center">3.504</td>
<td valign="bottom" align="center">3.530</td>
<td valign="bottom" align="center">3.483</td>
<td valign="bottom" align="center">3.015</td>
</tr>
<tr>
<td valign="bottom" align="center">BRISQUE</td>
<td valign="bottom" align="center">22.149</td>
<td valign="bottom" align="center">25.884</td>
<td valign="bottom" align="center">25.006</td>
<td valign="bottom" align="center">24.113</td>
<td valign="bottom" align="center">24.847</td>
<td valign="bottom" align="center">25.054</td>
<td valign="bottom" align="center">22.028</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>Image dehazing comparisons for different methods on hazy maritime images.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g003.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Experimental results on image deraining</title>
<p>To evaluate the deraining performance of the method, this paper carries out comprehensive experiments using the Rain13K dataset (<xref ref-type="bibr" rid="B18">Jiang et&#xa0;al., 2020</xref>), comprising 13,700 pairs of clean and rainy images. To evaluate our approach, this paper employs 4 benchmarks [Test100 (<xref ref-type="bibr" rid="B46">Zhang et&#xa0;al., 2019</xref>), Rain100H (<xref ref-type="bibr" rid="B41">Yang et&#xa0;al., 2017</xref>), Rain100L (<xref ref-type="bibr" rid="B41">Yang et&#xa0;al., 2017</xref>), and Test2800 (<xref ref-type="bibr" rid="B9">Fu et&#xa0;al., 2017b</xref>)] for testing purposes. Here, this paper compares our method with 9 popular image deraining algorithms, including DerainNet (<xref ref-type="bibr" rid="B8">Fu et&#xa0;al., 2017a</xref>), SEMI (<xref ref-type="bibr" rid="B38">Wei et&#xa0;al., 2019</xref>), DIDMDN (<xref ref-type="bibr" rid="B45">Zhang and Patel, 2018</xref>), UMRL (<xref ref-type="bibr" rid="B42">Yasarla and Patel, 2019</xref>), RESCAN (<xref ref-type="bibr" rid="B25">Li et&#xa0;al., 2018b</xref>), PReNet (<xref ref-type="bibr" rid="B33">Ren et&#xa0;al., 2019</xref>), MSPFN (<xref ref-type="bibr" rid="B18">Jiang et&#xa0;al., 2020</xref>), MPRNet (<xref ref-type="bibr" rid="B44">Zamir et&#xa0;al., 2021</xref>), and IDT (<xref ref-type="bibr" rid="B39">Xiao et&#xa0;al., 2022</xref>). Here, this paper uses full-reference image evaluation metrics, as there are ground truth images available. This paper quantitatively assesses our results by computing both PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) scores (<xref ref-type="bibr" rid="B36">Wang et&#xa0;al., 2004</xref>), specifically using the Y channel within the YCbCr color space. This allows us to make objective comparisons, focusing on luminance information, which is crucial for evaluating image quality accurately. These metrics provide valuable insights into the fidelity and structural similarity between the deraining images and their corresponding ground truth references.</p>
<p>
<xref ref-type="table" rid="T2"><bold>Table&#xa0;2</bold></xref> presents the quantitative results of various algorithms for image deraining. In comparison to the recent IDT method (<xref ref-type="bibr" rid="B39">Xiao et&#xa0;al., 2022</xref>), our approach demonstrates an average improvement of 1dB across the four datasets. This significant enhancement highlights our method&#x2019;s ability to adapt more effectively to diverse rainy conditions. Our results suggest that our approach outperforms existing methods by providing superior deraining performance across a range of challenging rain scenarios. <xref ref-type="fig" rid="f4"><bold>Figures&#xa0;4</bold></xref> and <xref ref-type="fig" rid="f5"><bold>5</bold></xref> show the visual results on the Rain100H and Test100 datasets. Observing these visual results, it becomes evident that DerainNet (<xref ref-type="bibr" rid="B8">Fu et&#xa0;al., 2017a</xref>), SEMI (<xref ref-type="bibr" rid="B38">Wei et&#xa0;al., 2019</xref>), and DIDMDN (<xref ref-type="bibr" rid="B45">Zhang and Patel, 2018</xref>) struggle to effectively remove heavy rain artifacts. UMRL (<xref ref-type="bibr" rid="B42">Yasarla and Patel, 2019</xref>), RESCAN (<xref ref-type="bibr" rid="B25">Li et&#xa0;al., 2018b</xref>), and PReNet (<xref ref-type="bibr" rid="B33">Ren et&#xa0;al., 2019</xref>) still leave residual rain streaks in their recovery results. MSPFN (<xref ref-type="bibr" rid="B18">Jiang et&#xa0;al., 2020</xref>), MPRNet (<xref ref-type="bibr" rid="B44">Zamir et&#xa0;al., 2021</xref>), and IDT (<xref ref-type="bibr" rid="B39">Xiao et&#xa0;al., 2022</xref>) exhibit shortcomings in preserving local image details, especially in regions such as the ship&#x2019;s hull. In contrast, our approach stands out due to its incorporation of prompt learning, enabling adaptive feature extraction. As a result, it excels in eliminating rain streaks while effectively retaining intricate image structures. This demonstrates the robustness and effectiveness of our method in addressing the challenges posed by rainy conditions and preserving fine-grained image details.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Quantitative comparisons of different methods on the Rain13K dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" rowspan="2" align="center">Datasets<break/>Method</th>
<th valign="top" colspan="2" align="center">Test100</th>
<th valign="top" colspan="2" align="center">Rain100H</th>
<th valign="top" colspan="2" align="center">Rain100L</th>
<th valign="top" colspan="2" align="center">Test2800</th>
<th valign="top" colspan="2" align="center">Average</th>
</tr>
<tr>
<th valign="top" align="center">PSNR</th>
<th valign="top" align="center">SSIM</th>
<th valign="top" align="center">PSNR</th>
<th valign="top" align="center">SSIM</th>
<th valign="top" align="center">PSNR</th>
<th valign="top" align="center">SSIM</th>
<th valign="top" align="center">PSNR</th>
<th valign="top" align="center">SSIM</th>
<th valign="top" align="center">PSNR</th>
<th valign="top" align="center">SSIM</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center">DerainNet</td>
<td valign="top" align="center">22.77</td>
<td valign="top" align="center">0.810</td>
<td valign="top" align="center">14.92</td>
<td valign="top" align="center">0.592</td>
<td valign="top" align="center">27.03</td>
<td valign="top" align="center">0.884</td>
<td valign="top" align="center">24.31</td>
<td valign="top" align="center">0.861</td>
<td valign="top" align="center">22.26</td>
<td valign="top" align="center">0.787</td>
</tr>
<tr>
<td valign="top" align="center">SEMI</td>
<td valign="top" align="center">22.35</td>
<td valign="top" align="center">0.788</td>
<td valign="top" align="center">16.56</td>
<td valign="top" align="center">0.486</td>
<td valign="top" align="center">25.03</td>
<td valign="top" align="center">0.842</td>
<td valign="top" align="center">24.43</td>
<td valign="top" align="center">0.782</td>
<td valign="top" align="center">22.09</td>
<td valign="top" align="center">0.725</td>
</tr>
<tr>
<td valign="top" align="center">DIDMDN</td>
<td valign="top" align="center">22.56</td>
<td valign="top" align="center">0.818</td>
<td valign="top" align="center">17.35</td>
<td valign="top" align="center">0.524</td>
<td valign="top" align="center">25.23</td>
<td valign="top" align="center">0.741</td>
<td valign="top" align="center">28.13</td>
<td valign="top" align="center">0.867</td>
<td valign="top" align="center">23.32</td>
<td valign="top" align="center">0.738</td>
</tr>
<tr>
<td valign="top" align="center">UMRL</td>
<td valign="top" align="center">24.41</td>
<td valign="top" align="center">0.829</td>
<td valign="top" align="center">26.01</td>
<td valign="top" align="center">0.832</td>
<td valign="top" align="center">29.18</td>
<td valign="top" align="center">0.923</td>
<td valign="top" align="center">29.97</td>
<td valign="top" align="center">0.905</td>
<td valign="top" align="center">27.39</td>
<td valign="top" align="center">0.872</td>
</tr>
<tr>
<td valign="top" align="center">RESCAN</td>
<td valign="top" align="center">25.00</td>
<td valign="top" align="center">0.835</td>
<td valign="top" align="center">26.36</td>
<td valign="top" align="center">0.786</td>
<td valign="top" align="center">29.80</td>
<td valign="top" align="center">0.881</td>
<td valign="top" align="center">31.29</td>
<td valign="top" align="center">0.904</td>
<td valign="top" align="center">28.11</td>
<td valign="top" align="center">0.852</td>
</tr>
<tr>
<td valign="top" align="center">PReNet</td>
<td valign="top" align="center">24.81</td>
<td valign="top" align="center">0.851</td>
<td valign="top" align="center">26.77</td>
<td valign="top" align="center">0.858</td>
<td valign="top" align="center">32.44</td>
<td valign="top" align="center">0.950</td>
<td valign="top" align="center">31.75</td>
<td valign="top" align="center">0.916</td>
<td valign="top" align="center">28.94</td>
<td valign="top" align="center">0.894</td>
</tr>
<tr>
<td valign="top" align="center">MSPFN</td>
<td valign="top" align="center">27.50</td>
<td valign="top" align="center">0.876</td>
<td valign="top" align="center">28.66</td>
<td valign="top" align="center">0.860</td>
<td valign="top" align="center">32.40</td>
<td valign="top" align="center">0.933</td>
<td valign="top" align="center">32.82</td>
<td valign="top" align="center">0.930</td>
<td valign="top" align="center">30.35</td>
<td valign="top" align="center">0.900</td>
</tr>
<tr>
<td valign="top" align="center">MPRNet</td>
<td valign="top" align="center">30.27</td>
<td valign="top" align="center">0.897</td>
<td valign="top" align="center">30.41</td>
<td valign="top" align="center">0.890</td>
<td valign="top" align="center">36.40</td>
<td valign="top" align="center">0.965</td>
<td valign="top" align="center">33.64</td>
<td valign="top" align="center">0.938</td>
<td valign="top" align="center">32.68</td>
<td valign="top" align="center">0.923</td>
</tr>
<tr>
<td valign="top" align="center">IDT</td>
<td valign="top" align="center">29.69</td>
<td valign="top" align="center">0.905</td>
<td valign="top" align="center">29.95</td>
<td valign="top" align="center">0.898</td>
<td valign="top" align="center">37.01</td>
<td valign="top" align="center">0.971</td>
<td valign="top" align="center">33.38</td>
<td valign="top" align="center">0.937</td>
<td valign="top" align="center">32.51</td>
<td valign="top" align="center">0.928</td>
</tr>
<tr>
<td valign="top" align="center">Ours</td>
<td valign="top" align="center">31.08</td>
<td valign="top" align="center">0.907</td>
<td valign="top" align="center">30.85</td>
<td valign="top" align="center">0.900</td>
<td valign="top" align="center">38.30</td>
<td valign="top" align="center">0.974</td>
<td valign="top" align="center">34.02</td>
<td valign="top" align="center">0.941</td>
<td valign="top" align="center"><bold>33.56</bold>
</td>
<td valign="top" align="center"><bold>0.930</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold indicates the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Image deraining comparisons for different methods on the Rain100H dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g004.tif"/>
</fig>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Image deraining comparisons for different methods on the Test100 dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g005.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Ablation experiment and analysis</title>
<p>To further analysis the effectiveness of our proposed method, this paper conducts additional ablation experiments. This paper focuses on evaluating the performance of our approach in the image deraining task, specifically examining the effectiveness of prompting modules and the effect of prompt length in prompting modules.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Effectiveness of prompting modules</title>
<p>In this experiment, this paper aims to gauge the contributions of the prompting modules within our methodology. This paper designs three distinct model variants to comprehensively analyze the impact of prompting modules on our approach. These variants encompass models with or without prompting modules, models with or without PFB, and models with prompting modules placed at different positions within the network architecture. The quantitative results for these various model variants are documented in <xref ref-type="table" rid="T3"><bold>Table&#xa0;3</bold></xref>, shedding light on their respective contributions to the image restoration task.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Ablation experimental result on the effectiveness of prompting modules.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="bottom" rowspan="2" align="center">Model</th>
<th valign="bottom" colspan="3" align="center">Prompting module</th>
<th valign="bottom" rowspan="2" align="center">Restoration<break/>module</th>
<th valign="bottom" rowspan="2" align="center">PSNR/SSIM</th>
</tr>
<tr>
<th valign="bottom" align="center">PCB</th>
<th valign="bottom" align="center">PFB</th>
<th valign="bottom" align="center">Position</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">(a)</td>
<td valign="bottom" align="center">&#xd7;</td>
<td valign="bottom" align="center">&#xd7;</td>
<td valign="bottom" align="center">&#x2013;</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center">33.15/0.924</td>
</tr>
<tr>
<td valign="bottom" align="center">(b)</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="bottom" align="center">&#xd7;</td>
<td valign="bottom" align="center">Enc.</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center">33.39/0.927</td>
</tr>
<tr>
<td valign="bottom" align="center">(c)</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center">Enc.</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center">33.48/0.929</td>
</tr>
<tr>
<td valign="bottom" align="center">(d)</td>
<td valign="top" align="center">&#x2713;</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center">Dec.</td>
<td valign="bottom" align="center">&#x2713;</td>
<td valign="bottom" align="center"><bold>33.56/0.930</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold indicates the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Starting with Model (a), the performance starkly deteriorates in the absence of prompting modules. This striking decline underscores the pivotal role that prompting modules play in the image restoration process. They serve as critical components in guiding the model&#x2019;s understanding of the task and assisting it in achieving high-quality restoration results. Comparing Model (b) and Model (c), it becomes evident that the absence of PFB in the prompting modules results in a disconnect between the guidance provided by the prompts and the features processed in the main restoration modules. This lack of synchronization hinders the model&#x2019;s ability to effectively utilize the provided prompts, consequently restricting its overall performance. Further comparison between Model (c) and Model (d) reveals the advantage of introducing prompting modules during the network decoding phase. This strategic placement allows the model to better harness the clear image features, facilitating improved performance and yielding the best network performance among the variants.</p>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Effect of prompt length in prompting modules</title>
<p>To delve deeper into the impact of prompts in our prompting modules, this paper conducts experiments to investigate the effect of prompt length. This paper varies the length of prompts while keeping other parameters constant and examined how it influenced the model&#x2019;s performance. <xref ref-type="table" rid="T4"><bold>Table&#xa0;4</bold></xref> presents the quantitative results. Our findings reveal that prompt length plays a crucial role in image restoration process. Specifically, shorter prompts tend to yield faster convergence and better overall results, while longer prompts sometimes lead to overfitting or increased computational complexity. This analysis allows us to optimize the prompt length within our prompting modules to achieve the best balance between performance and efficiency. Finally, this paper determines that a prompt length of 5 is the most suitable configuration for our model.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>Ablation experimental result on the effect of prompt length in prompting modules.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="center">Length</th>
<th valign="top" align="center">3</th>
<th valign="top" align="center">4</th>
<th valign="top" align="center">5</th>
<th valign="top" align="center">6</th>
<th valign="top" align="center">7</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center">PSNR/SSIM</td>
<td valign="top" align="center">33.41/0.928</td>
<td valign="top" align="center">33.49/0.929</td>
<td valign="top" align="center"><bold>33.56/0.930</bold>
</td>
<td valign="top" align="center">33.52/0.930</td>
<td valign="top" align="center">33.44/0.927</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>Bold indicates the best results.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>User study</title>
<p>This paper conducts a user study to assess the outcomes of various image restoration techniques. These studies are based on restoring real foggy maritime images. A portion of the participants invited to our user study are professionals engaged in maritime activities to ensure the professionalism of our work. Additionally, in order to enhance the universality of our work, this paper has also invited non-maritime professionals to participate in our user study. Participants are presented with a set of images and asked to select the one with the best visual clarity. To maintain fairness, the methods remain anonymous, and the images within each set are randomly ordered. This paper distributes the questionnaire widely to online users and collect responses. Finally, this paper receives responses from 46 human evaluators. <xref ref-type="fig" rid="f6"><bold>Figure&#xa0;6</bold></xref> depicts the average selection percentage for each method. Based on the majority of human evaluators&#x2019; feedback, our method consistently outperforms the others.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Averaged selection percentage of user study.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmars-11-1382147-g006.tif"/>
</fig>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Limitations</title>
<p>While our proposed method introduces prompt learning to enhance restoration performance under adverse weather conditions, it comes with relatively high parameter count and computational complexity. The comparison results are presented in the <xref ref-type="table" rid="T5"><bold>Table&#xa0;5</bold></xref>, compared to other methods, our approach has a higher parameter count. As a result, significant computational resources are also required, which to some extent limits the applicability of the model. In order to enable our method to be more rapidly and conveniently applied in maritime operations, this paper plans to address this issue through strategies such as model compression and pruning, aiming to make the model more suitable for maritime vision applications.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Comparison of model efficiency.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="center"/>
<th valign="top" align="center">Uformer</th>
<th valign="top" align="center">Restormer</th>
<th valign="top" align="center">IDT</th>
<th valign="top" align="center">DRSformer</th>
<th valign="top" align="center">Ours</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="center">Paramaters (M)</td>
<td valign="top" align="center">50.8</td>
<td valign="top" align="center">26.1</td>
<td valign="top" align="center">16.4</td>
<td valign="top" align="center">33.7</td>
<td valign="top" align="center">35.6</td>
</tr>
<tr>
<td valign="top" align="center">Flops (G)</td>
<td valign="top" align="center">45.9</td>
<td valign="top" align="center">174.7</td>
<td valign="top" align="center">61.9</td>
<td valign="top" align="center">242.9</td>
<td valign="top" align="center">43.2</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusions</title>
<p>This paper has proposed a prompting image restoration approach by learning degradation-aware visual prompt for maritime surveillance. Our proposed approach possesses the capability to interact with input features, allowing for dynamic adjustments in weather-related representations. This adaptability ensures that the restoration process is tailored to the specific degradation being addressed. This paper has validated the effectiveness of our method on extensive experimental datasets, enhancing its restoration performance in various weather conditions, including rain removal and haze removal in maritime images. In future work, this paper plans to explore leveraging text models such as CLIP as alternative prompts to further guide the image restoration process.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <uri xlink:href="https://sites.google.com/view/reside-dehaze-datasets">https://sites.google.com/view/reside-dehaze-datasets</uri>, <uri xlink:href="https://www.deraining.tech/benchmark.html">https://www.deraining.tech/benchmark.html</uri>.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>XH: Data curation, Formal analysis, Methodology, Writing &#x2013; original draft, Writing &#x2013; review &amp; editing. TJ: Data curation, Software, Visualization, Writing &#x2013; original draft. JL: Investigation, Writing &#x2013; review &amp; editing.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was partly supported by the Research Project of the Naval Staff Navigation Assurance Bureau (2023(1)).</p>
</sec>
<sec id="s9" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ai</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>He</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Multimodal prompt perceiver: Empower adaptiveness, generalizability and fidelity for all-in-one image restoration</article-title>. arXiv [preprint] arXiv:2312.02918.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Qing</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Dehazenet: An end-to-end system for single image haze removal</article-title>. <source>IEEE Trans. image Process.</source> <volume>25</volume>, <fpage>5187</fpage>&#x2013;<lpage>5198</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.2016.2598681</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2023</year>a). &#x201c;<article-title>Learning a sparse transformer network for effective image deraining</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>Canada</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5896</fpage>&#x2013;<lpage>5905</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00571</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2023</year>b). <article-title>Towards unified deep image deraining: A survey and a new benchmark</article-title>. arXiv [preprint] arXiv:2310.03535.</citation>
</ref>
<ref id="B5">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Kong</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). &#x201c;<article-title>Unpaired deep image deraining using dual contrastive learning</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2017</fpage>&#x2013;<lpage>2026</lpage>
</citation>
</ref>
<ref id="B6">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Fan</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2023</year>c). &#x201c;<article-title>Hybrid cnn-transformer feature fusion for single image deraining</article-title>,&#x201d; in <conf-name>Proceedings of the AAAI Conference on Artificial Intelligence</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>AAAI Press</publisher-name>), Vol. <volume>37</volume>. <fpage>378</fpage>&#x2013;<lpage>386</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1609/aaai.v37i1.25111</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dong</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xiang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>F.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). &#x201c;<article-title>Multi-scale boosted dehazing network with dense feature fusion</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2157</fpage>&#x2013;<lpage>2167</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Paisley</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>a). <article-title>Clearing the skies: A deep network architecture for single-image rain removal</article-title>. <source>IEEE Trans. Image Process.</source> <volume>26</volume>, <fpage>2944</fpage>&#x2013;<lpage>2956</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.2017.2691802</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Paisley</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>b). &#x201c;<article-title>Removing rain from single images via a deep detail network</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3855</fpage>&#x2013;<lpage>3863</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2017.186</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Guo</surname> <given-names>C.-L.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Anwar</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Cong</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Image dehazing transformer with transmission-aware 3d position embedding</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5812</fpage>&#x2013;<lpage>5820</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Guo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>R. W.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). &#x201c;<article-title>Heterogeneous twin dehazing network for visibility enhancement in maritime video surveillance</article-title>,&#x201d; in <conf-name>2021 IEEE International Intelligent Transportation Systems Conference (ITSC)</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>2875</fpage>&#x2013;<lpage>2880</lpage>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Tang</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Single image haze removal using dark channel prior</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>33</volume>, <fpage>2341</fpage>&#x2013;<lpage>2353</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Single maritime image dehazing using unpaired adversarial learning</article-title>. <source>Signal Image Video Process.</source> <volume>17</volume>, <fpage>593</fpage>&#x2013;<lpage>600</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11760-022-02265-5</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>H.-M.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Single image defogging based on illumination decomposition for visual maritime surveillance</article-title>. <source>IEEE Trans. Image Process.</source> <volume>28</volume>, <fpage>2882</fpage>&#x2013;<lpage>2897</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.83</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Tong</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep learning-enabled variational optimization method for image dehazing in maritime intelligent transportation systems</article-title>. <source>J. Advanced Transportation</source> <volume>2021</volume>, <fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2021/6658763</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Maritime video defogging based on spatialtemporal information fusion and an improved dark channel prior</article-title>. <source>Multimedia Tools Appl.</source> <volume>81</volume>, <fpage>24777</fpage>&#x2013;<lpage>24798</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11042-022-11921-4</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Niu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Miah</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>A sea fog image defogging method based on the improved convex optimization model</article-title>. <source>J. Mar. Sci. Eng.</source> <volume>11</volume>, <fpage>1775</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/jmse11091775</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yi</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). &#x201c;<article-title>Multi-scale progressive fusion network for single image deraining</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8346</fpage>&#x2013;<lpage>8355</lpage>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kong</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Dong</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Towards effective multiple-in-one image restoration: A sequential and prompt learning strategy</article-title>. arXiv [preprint] arXiv:2401.03379.</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lei</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Shan</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2023</year>b). <article-title>Prompt-in-prompt learning for universal image restoration</article-title>. arXiv [preprint] arXiv:2312.05038.</citation>
</ref>
<ref id="B21">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2023</year>a). &#x201c;<article-title>Dilated convolutional transformer for highquality image deraining</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (Canada: IEEE), <fpage>4198</fpage>&#x2013;<lpage>4206</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPRW59228.2023.00442</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Aod-net: All-in-one dehazing network</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>. (<publisher-loc>Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4770</fpage>&#x2013;<lpage>4778</lpage>.</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Tao</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>W.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>a). <article-title>Benchmarking single-image dehazing and beyond</article-title>. <source>IEEE Trans. Image Process.</source> <volume>28</volume>, <fpage>492</fpage>&#x2013;<lpage>505</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.83</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>R. T.</given-names>
</name>
<name>
<surname>Cheong</surname> <given-names>L.-F.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>All in one bad weather removal using architectural search</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3175</fpage>&#x2013;<lpage>3185</lpage>.</citation>
</ref>
<ref id="B25">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zha</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>b). &#x201c;<article-title>Recurrent squeeze-and-excitation context aggregation net for single image deraining</article-title>,&#x201d; in <conf-name>Proceedings of the European conference on computer vision (ECCV)</conf-name>. (<publisher-loc>Germany</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>254</fpage>&#x2013;<lpage>269</lpage>.</citation>
</ref>
<ref id="B26">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Griddehazenet: Attention-based multi-scale network for image dehazing</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF international conference on computer vision</conf-name>. <publisher-loc>Korea (South)</publisher-loc>: <publisher-name>(IEEE)</publisher-name>, <fpage>7314</fpage>&#x2013;<lpage>7323</lpage>.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Dual-channel and two-stage dehazing network for promoting ship detection in visual perception system</article-title>. <source>Math. Problems Eng.</source> <volume>2022</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2022/8998743</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Loshchilov</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Hutter</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Sgdr: Stochastic gradient descent with warm restarts</article-title>. arXiv [preprint] arXiv:1608.03983.</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Cnn-enabled visibility enhancement framework for vessel detection under haze environment</article-title>. <source>J. advanced transportation</source> <volume>2021</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1155/2021/7649214</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mittal</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Moorthy</surname> <given-names>A. K.</given-names>
</name>
<name>
<surname>Bovik</surname> <given-names>A. C.</given-names>
</name>
</person-group> (<year>2012</year>a). <article-title>No-reference image quality assessment in the spatial domain</article-title>. <source>IEEE Trans. image Process.</source> <volume>21</volume>, <fpage>4695</fpage>&#x2013;<lpage>4708</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.2012.2214050</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mittal</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Soundararajan</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Bovik</surname> <given-names>A. C.</given-names>
</name>
</person-group> (<year>2012</year>b). <article-title>Making a &#x201c;completely blind&#x201d; image quality analyzer</article-title>. <source>IEEE Signal Process. Lett.</source> <volume>20</volume>, <fpage>209</fpage>&#x2013;<lpage>212</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Potlapalli</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Zamir</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>F. S.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Promptir: Prompting for all-in-one blind image restoration</article-title>. arXiv [preprint] arXiv:2306.13090.</citation>
</ref>
<ref id="B33">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ren</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zuo</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Progressive image deraining networks: A better and simpler baseline</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3937</fpage>&#x2013;<lpage>3946</lpage>.</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Van Nguyen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Mai</surname> <given-names>T. T. N.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Single maritime image defogging based on illumination decomposition using texture and structure priors</article-title>. <source>IEEE Access</source> <volume>9</volume>, <fpage>34590</fpage>&#x2013;<lpage>34603</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ACCESS.2021.3060439</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Shazeer</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Parmar</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Uszkoreit</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Gomez</surname> <given-names>A. N.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Attention is all you need</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>30</volume>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Bovik</surname> <given-names>A. C.</given-names>
</name>
<name>
<surname>Sheikh</surname> <given-names>H. R.</given-names>
</name>
<name>
<surname>Simoncelli</surname> <given-names>E. P.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Image quality assessment: from error visibility to structural similarity</article-title>. <source>IEEE Trans. image Process.</source> <volume>13</volume>, <fpage>600</fpage>&#x2013;<lpage>612</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TIP.2003.819861</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>C.-Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). &#x201c;<article-title>Learning to prompt for continual learning</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>139</fpage>&#x2013;<lpage>149</lpage>.</citation>
</ref>
<ref id="B38">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wei</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Semi-supervised transfer learning for image rain removal</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>3877</fpage>&#x2013;<lpage>3886</lpage>.</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zha</surname> <given-names>Z.-J.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Image de-raining transformer</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell</source>.</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A novel approach to maritime image dehazing based on a large kernel encoder&#x2013;decoder network with multihead pyramids</article-title>. <source>Electronics</source> <volume>11</volume>, <fpage>3351</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/electronics11203351</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>R. T.</given-names>
</name>
<name>
<surname>Feng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Deep joint rain detection and removal from a single image</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1357</fpage>&#x2013;<lpage>1366</lpage>.</citation>
</ref>
<ref id="B42">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yasarla</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Patel</surname> <given-names>V. M.</given-names>
</name>
</person-group> (<year>2019</year>). &#x201c;<article-title>Uncertainty guided multi-scale residual learning-using a cycle spinning cnn for single image de-raining</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>8405</fpage>&#x2013;<lpage>8414</lpage>.</citation>
</ref>
<ref id="B43">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zamir</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Arora</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hayat</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>F. S.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>M.-H.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Restormer: Efficient transformer for high-resolution image restoration</article-title>,&#x201d; In <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>5728</fpage>&#x2013;<lpage>5739</lpage>.</citation>
</ref>
<ref id="B44">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zamir</surname> <given-names>S. W.</given-names>
</name>
<name>
<surname>Arora</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hayat</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Khan</surname> <given-names>F. S.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>M.-H.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Multi-stage progressive image restoration</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, virtual</conf-name>. (<publisher-loc>virtual</publisher-loc>: <publisher-name>IEEE</publisher-name>) <fpage>14821</fpage>&#x2013;<lpage>14831</lpage>.</citation>
</ref>
<ref id="B45">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Patel</surname> <given-names>V. M.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Density-aware single image de-raining using a multi-stream dense network</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>695</fpage>&#x2013;<lpage>704</lpage>.</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Sindagi</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Patel</surname> <given-names>V. M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Image de-raining using a conditional generative adversarial network</article-title>. <source>IEEE Trans. circuits Syst. video Technol.</source> <volume>30</volume>, <fpage>3943</fpage>&#x2013;<lpage>3956</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TCSVT.76</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Qi</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). &#x201c;<article-title>Overwater image dehazing via cycle-consistent generative adversarial network</article-title>,&#x201d; in <conf-name>Proceedings of the Asian Conference on Computer Vision</conf-name>. (<publisher-loc>Japan</publisher-loc>: <publisher-name>Springer</publisher-name>), doi:&#xa0;<pub-id pub-id-type="doi">10.3390/electronics9111877</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Adaptive collision avoidance decisions in autonomous ship encounter scenarios through rule-guided vision supervised learning</article-title>. <source>Ocean Eng.</source> <volume>297</volume>, <fpage>117096</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.oceaneng.2024.117096</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>C.-L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2023</year>). &#x201c;<article-title>Fourmer: An efficient global modeling paradigm for image restoration</article-title>,&#x201d; in <conf-name>International conference on machine learning</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>PMLR</publisher-name>), <fpage>42589</fpage>&#x2013;<lpage>42601</lpage>.</citation>
</ref>
<ref id="B50">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). &#x201c;<article-title>Image de-raining via continual learning</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>virtual</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>4907</fpage>&#x2013;<lpage>4916</lpage>.</citation>
</ref>
<ref id="B51">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Loy</surname> <given-names>C. C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2022</year>). &#x201c;<article-title>Conditional prompt learning for vision-language models</article-title>,&#x201d; in <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. (<publisher-loc>USA</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>16816</fpage>&#x2013;<lpage>16825</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>
