<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Phys.</journal-id>
<journal-title>Frontiers in Physics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Phys.</abbrev-journal-title>
<issn pub-type="epub">2296-424X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1121311</article-id>
<article-id pub-id-type="doi">10.3389/fphy.2023.1121311</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Physics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Cascaded information enhancement and cross-modal attention feature fusion for multispectral pedestrian detection</article-title>
<alt-title alt-title-type="left-running-head">Yang et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphy.2023.1121311">10.3389/fphy.2023.1121311</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>Yang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Xu</surname>
<given-names>Kaixiong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2135585/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wang</surname>
<given-names>Kaizheng</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2135236/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Faculty of Information Engineering and Automation</institution>, <institution>Kunming University of Science and Technology</institution>, <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Faculty of Electrical Engineering</institution>, <institution>Kunming University of Science and Technology</institution>, <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/669638/overview">Bo Xiao</ext-link>, Imperial College London, United Kingdom</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1664211/overview">Guanqiu Qi</ext-link>, Buffalo State College, United States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2140551/overview">Jian Sun</ext-link>, Southwest University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Kaizheng Wang, <email>kz.wang@foxmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Radiation Detectors and Imaging, a section of the journal Frontiers in Physics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>18</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1121311</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Yang, Xu and Wang.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Yang, Xu and Wang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Multispectral pedestrian detection is a technology designed to detect and locate pedestrians in Color and Thermal images, which has been widely used in automatic driving, video surveillance, etc. So far most available multispectral pedestrian detection algorithms only achieved limited success in pedestrian detection because of the lacking take into account the confusion of pedestrian information and background noise in Color and Thermal images. Here we propose a multispectral pedestrian detection algorithm, which mainly consists of a cascaded information enhancement module and a cross-modal attention feature fusion module. On the one hand, the cascaded information enhancement module adopts the channel and spatial attention mechanism to perform attention weighting on the features fused by the cascaded feature fusion block. Moreover, it multiplies the single-modal features with the attention weight element by element to enhance the pedestrian features in the single-modal and thus suppress the interference from the background. On the other hand, the cross-modal attention feature fusion module mines the features of both Color and Thermal modalities to complement each other, then the global features are constructed by adding the cross-modal complemented features element by element, which are attentionally weighted to achieve the effective fusion of the two modal features. Finally, the fused features are input into the detection head to detect and locate pedestrians. Extensive experiments have been performed on two improved versions of annotations (sanitized annotations and paired annotations) of the public dataset KAIST. The experimental results show that our method demonstrates a lower pedestrian miss rate and more accurate pedestrian detection boxes compared to the comparison method. Additionally, the ablation experiment also proved the effectiveness of each module designed in this paper.</p>
</abstract>
<kwd-group>
<kwd>multispectral pedestrian detection</kwd>
<kwd>attention mechanism</kwd>
<kwd>feature fusion</kwd>
<kwd>convolutional neural network</kwd>
<kwd>background noise</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Pedestrian detection, parsing visual content to identify and locate pedestrians on an image/video, has been viewed as an essential and central task within the computer vision field and widely employed in various applications, e.g. autonomous driving, video surveillance and person re-identification [<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B7">7</xref>]. The performance of such technology has greatly advanced through the facilitation of convolutional neural networks (CNN). Typically, pedestrian detectors take Color images as input and try to retrieve the pedestrian information from them. However, the quality of Color images highly depends on the light condition. Missing recognition of pedestrians occurs frequently when pedestrian detectors process Color images with poor resolution and contrast caused by unfavorable lighting. Consequently, the use of such models has been limited for the application of all-weather devices.</p>
<p>Thermal imaging is related to the infrared radiation of pedestrians, barely affected by changes in ambient light. The technique of combining Color and Thermal images has been explored in recent years [<xref ref-type="bibr" rid="B8">8</xref>&#x2013;<xref ref-type="bibr" rid="B16">16</xref>]. These methods has been shown to exhibit positive effects on pedestrian detection performance in complex environments as it could retrieve more pedestrian information. However, despite important initial success, there remain two major challenges. First, as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, the image of pedestrians tends to blend with the background for nighttime Color images resulting from insufficient light [<xref ref-type="bibr" rid="B17">17</xref>], and for daytime Thermal images as well due to similar temperatures between the human body and the ambient environment [<xref ref-type="bibr" rid="B18">18</xref>]. Second, there is an essential difference between Color images and Thermal images the former displays the color and texture detail information of pedestrians while the latter shows the temperature information. Therefore, solutions needed to be taken to augment the pedestrian features in Color and Thermal modalities in order to suppress background interference, and enable better integration and understanding of both Color and Thermal images to improve the accuracy of pedestrian detection in complex environments.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Example of color and thermal images of pedestrians in daytime and nighttime scenes.</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g001.tif"/>
</fig>
<p>To address the challenges above, the researches [<xref ref-type="bibr" rid="B19">19</xref>,<xref ref-type="bibr" rid="B20">20</xref>] designed illumination-aware networks to obtain illumination-measured parameters of Color and Thermal images respectively, which were used as fusion weights for Color and Thermal features in order to realize a self-adaptively fuse of two modal features. However, the acquisition of illumination-measured parameters relied heavily on the classification scores, the accuracy of which was limited by the performance of the classifier. [<xref ref-type="bibr" rid="B21">21</xref>] reported confidence-aware networks to predict the confidence of detection boxes for each modal, and then Dempster-Sheffer theory combination rules were employed to fuse the results of different branches based on uncertainty. Nevertheless, the accuracy of predicting the detection boxes&#x2019; confidence is also affected by the performance of the confidence-aware network. A cyclic fusion and refinement scheme was introduced by [<xref ref-type="bibr" rid="B22">22</xref>] for the sake of gradually improving the quality of Color and Thermal features and automatically adjusting the complementary and consistent information balance of the two modalities to effectively utilize the information of both modalities. However, this method only used a simple feature cascade operation to fuse Color and Thermal features and failed to fully exploit the complementary features of these two modalities.</p>
<p>To tackle the problems aforementioned, we propose a multispectral pedestrian detection algorithm with cascaded information enhancement and cross-modal attention feature fusion. The cascaded information enhancement module (CIEM) is designed to enhance the pedestrian information suppressed by the background in the Color and Thermal images. CIEM uses a cascaded feature fusion block to fuse Color and Thermal features to obtain fused features of both modalities. Since the fused features contain the consistency and complementary information of Color and Thermal modalities, the fused features can be used to enhance Color and Thermal features respectively to reduce the interference of background on pedestrian information. Inspired by the attention mechanism, the attention weights of the fused features are sequentially obtained by channel and spatial attention learning, and the Color and Thermal features are multiplied with the attention weights element by element, respectively. In this way, the single-modal features have the combined information of the two modalities, and the single-modal information is enhanced from the perspective of the fused features. Although CIEM enriches single-modal pedestrian features, simple feature fusion of the enhanced single-modal features is still insufficient for robust multispectral pedestrian detection. Thus, we design the cross-modal attention feature fusion module (CAFFM) to efficiently fuse Color and Thermal features. Cross-modal attention is used in this module to implement the differentiation of pedestrian features between different modalities. In order to supplement the pedestrian information of the other modality to the local modality, the attention of the other modality is adopted to augment the pedestrian features of the local modality. A global feature is constructed by adding the Color and Thermal features after performing cross-modal feature enhancement, and the global feature is used to guide the fusion of the Color and Thermal features. Overall, the method presented in this paper enables more comprehensive pedestrian features acquisition through cascaded information enhancement and cross-modal attention feature fusion, which effectively enhances the accuracy of multispectral image pedestrian detection. The main contributions of this paper are summarized as follows.<list list-type="simple">
<list-item>
<p>(1) A cascaded information enhancement module is proposed. From the perspective of fused features, it reduces the interference from the background of Color and Thermal modalities on pedestrian detection and augments the pedestrian features of Color and Thermal modalities separately through an attention mechanism.</p>
</list-item>
<list-item>
<p>(2) The designed cross-modal attention feature fusion module first mines the features of both Color and Thermal modalities separately through a cross-modal attention network and adds them to the other modality for cross-modal feature enhancement. Meanwhile, the cross-modal enhanced Color and Thermal features are used to construct global features to guide the feature fusion of the two modalities.</p>
</list-item>
<list-item>
<p>(3) Numerous experiments are conducted on the public dataset KAIST to demonstrate the effectiveness and superiority of the proposed method. In addition, the ablation experiments also demonstrate the effectiveness of the proposed modules.</p>
</list-item>
</list>
</p>
</sec>
<sec id="s2">
<title>2 Related works</title>
<sec id="s2-1">
<title>2.1 Multispectral pedestrian detection</title>
<p>Multispectral sensors can obtain paired Color-Thermal images to provide complementary information about pedestrian targets. A large multispectral pedestrian detection (KAIST) dataset was constructed by [<xref ref-type="bibr" rid="B8">8</xref>]. Meanwhile, by combining the traditional aggregated channel feature (ACF) pedestrian detector [<xref ref-type="bibr" rid="B23">23</xref>] with the HOG algorithm [<xref ref-type="bibr" rid="B24">24</xref>], an extended ACF (ACF &#x2b; T &#x2b; THOG) method was proposed to fuse Color and Thermal features. In 2016, [<xref ref-type="bibr" rid="B9">9</xref>] proposed four fusion modalities of low-layer feature, middle-layer feature, high-layer feature, and confidence fraction fusion with VGG16 as the backbone network, and the middle-layer feature fusion was proved to offer the maximum integration capability of Color and Thermal features. Inspired by this, [<xref ref-type="bibr" rid="B25">25</xref>] developed a multispectral region candidate network with Faster RCNN (Region with CNN features, RCNN) [<xref ref-type="bibr" rid="B26">26</xref>] as the architecture and replaced the original classifier in Faster RCNN with an enhanced decision tree classifier to reduce the missed and false detection of pedestrians. Recently,[<xref ref-type="bibr" rid="B27">27</xref>] deployed the EfficientDet as the backbone network and proposed an EfficientDet-based fusion framework for multispectral pedestrian detection to improve the detection accuracy of pedestrians in Color and Thermal images by adding and cascading the Color and Thermal features. Although the studies [<xref ref-type="bibr" rid="B8">8</xref>,<xref ref-type="bibr" rid="B9">9</xref>,<xref ref-type="bibr" rid="B25">25</xref>,<xref ref-type="bibr" rid="B27">27</xref>] fused Color and Thermal features for pedestrian detection, they mainly focused on exploring the impact of different stages of fusion on pedestrian detection, and only adopted simple feature fusion and not focusing on the case of pedestrian and background confusion.</p>
<p>In 2019, [<xref ref-type="bibr" rid="B28">28</xref>] observed a weak alignment problem of pedestrian position between Color and Thermal images, for which the KAIST dataset was re-annotated and Aligned Region CNN (AR-CNN) was proposed to handle weakly aligned multispectral pedestrian detection data in an end-to-end manner. But the deployment of this algorithm requires pairs of annotations, and the annotation of the dataset is a time-consuming and labor-intensive task, which makes the algorithm difficult to be applied in realistic scenes. [<xref ref-type="bibr" rid="B29">29</xref>] proposed a new single-stage multispectral pedestrian detection framework. This framework used multi-label learning to learn input state-aware features based on the state of the input image pair by assigning an individual label (if the pedestrian is visible in only one image of the image pair, the label vector is assigned as <italic>y</italic>
<sub>1</sub> &#x2208; [0, 1] or <italic>y</italic>
<sub>2</sub> &#x2208; [1, 0]; if the pedestrian is visible in both images of the image pair, the label vector is assigned as <italic>y</italic>
<sub>3</sub> &#x2208; [1, 1]) to solve the problem of weak alignment of pedestrian locations between Color and Thermal images, but the model still requires pairs of annotations during training. [<xref ref-type="bibr" rid="B19">19</xref>] designed illumination-aware networks to obtain illumination-measured parameters for Color and Thermal images separately and used them as the fusion weights for Color and Thermal features. [<xref ref-type="bibr" rid="B20">20</xref>] designed a differential modality perception fusion module to guide the features of the two modalities to become similar, and then used the illumination perception network to assign fusion weights to the Color and Thermal features. [<xref ref-type="bibr" rid="B30">30</xref>] reported an uncertainty-aware cross-modal guidance (UCG) module to guide the distribution of modal features with high prediction uncertainty to align with the distribution of modal features with low prediction uncertainty. The researches [<xref ref-type="bibr" rid="B19">19</xref>,<xref ref-type="bibr" rid="B20">20</xref>] noticed that the pedestrians in Color and Thermal images are easily confused with the background and used illumination-aware networks to assign fusion weights to Color and Thermal features. However, the acquisition of illumination-measured parameters relied heavily on the classification scores, whose accuracy was limited by the performance of the classifier. In contrast, the method proposed in this paper not only considers the confusion of pedestrians and background in Color and Thermal images but also effectively fuses the two modal features.</p>
</sec>
<sec id="s2-2">
<title>2.2 Attention mechanisms</title>
<p>Attention mechanisms [<xref ref-type="bibr" rid="B31">31</xref>] utilized in computer vision are aimed to perform the processing of visual information. Currently, attention mechanisms have been widely used in semantic segmentation [<xref ref-type="bibr" rid="B32">32</xref>], image captioning [<xref ref-type="bibr" rid="B33">33</xref>], image fusion [<xref ref-type="bibr" rid="B34">34</xref>,<xref ref-type="bibr" rid="B35">35</xref>], image dehazing [<xref ref-type="bibr" rid="B36">36</xref>], saliency target detection [<xref ref-type="bibr" rid="B37">37</xref>], person re-identification [<xref ref-type="bibr" rid="B38">38</xref>&#x2013;<xref ref-type="bibr" rid="B40">40</xref>], etc. [<xref ref-type="bibr" rid="B41">41</xref>] introduced the idea of a squeeze and excitation network (SENet) to simulate the interdependence between feature channels in order to generate channel attention to recalibrate the feature mapping of channel directions. [<xref ref-type="bibr" rid="B42">42</xref>] employed the use of a selective kernel unit (SKNet) to adaptively fuse branches with different kernel sizes based on input information. A work inspired by this was from [<xref ref-type="bibr" rid="B43">43</xref>]. They designed a multi-scale channel attention feature fusion network that used channel attention mechanisms to replace simple fusion operations such as feature cascades or summations in feature fusion to produce richer feature representations. However, this recent progress in multispectral pedestrian detection has also been limited to two main challenges the interference caused by background and the difference of fundamental characteristics in Color and Thermal images. Therefore, we propose a multispectral pedestrian detection algorithm with cascaded information enhancement and cross-modal attention feature fusion based on the attention mechanism.</p>
</sec>
</sec>
<sec sec-type="methods" id="s3">
<title>3 Methods</title>
<p>The overall network framework of the proposed algorithm is shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The network consists of an encoder, a cascaded information enhancement module (CIEM), a cross-modal attentional feature fusion module (CAFFM) and a detection head. Specifically, ResNet-101 [<xref ref-type="bibr" rid="B44">44</xref>] is used as the backbone network of the encoder to encode the features of the input Color images <bold>
<italic>X</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and Thermal images <bold>
<italic>X</italic>
</bold>
<sub>
<italic>t</italic>
</sub> to obtain the corresponding feature maps <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> &#x2208; R<sup>
<italic>W</italic>&#xd7;<italic>H</italic>&#xd7;<italic>C</italic>
</sup> and <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub> &#x2208; R<sup>
<italic>W</italic>&#xd7;<italic>H</italic>&#xd7;<italic>C</italic>
</sup> (<italic>W</italic>, <italic>H</italic>, <italic>C</italic> represent the width, height and the number of channels of the feature maps, respectively). CIEM enhances single-modal information from the perspective of fused features by cascading feature fusion blocks to fuse <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub>, and attention weighting the fused features to enrich pedestrian features. CAFFM complements the features of different modalities by mining the complementary features between the two modalities and constructs global features to guide the effective fusion of the two modal features. The detection head is employed for pedestrian recognition and localization of the final fused features.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Overall framework of the proposed algorithm.</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g002.tif"/>
</fig>
<sec id="s3-1">
<title>3.1 Cascaded information enhancement module</title>
<p>Considering the confusion of pedestrians with the backgrounds in Color and Thermal images, we design a cascaded information enhancement module (CIEM) to augment the pedestrian features of both modalities to mitigate the effect of background interference on pedestrian detection. Specifically, a cascaded feature fusion block is used to fuse the Color features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and Thermal features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub>. The cascaded feature fusion block consists of feature cascade, 1 &#xd7; 1 convolution, 3 &#xd7; 3 convolution, <italic>BN</italic> layer, and <italic>ReLu</italic> activation function. The feature cascade operation splice <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub> along the direction of channels. 1 &#xd7; 1 convolution is conducive to cross-channel feature interaction in the channel dimension and reducing the number of channels in the splice feature map, while 3 &#xd7; 3 convolution expands the field of perception and makes a more comprehensive fusion of features for generating fusion features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub>:<disp-formula id="e1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>u</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>B</mml:mi>
<mml:mi>N</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>BN</italic> denotes batch normalization, <inline-formula id="inf1">
<mml:math id="m2">
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x22c5;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="0.3333em"/>
</mml:math>
</inline-formula> denotes a convolution kernel with kernel size <italic>n</italic> &#xd7; <italic>n</italic>, [&#x22c5;, &#x22c5;] denotes the cascade of features along the channel direction, <italic>ReLu</italic>(&#x22c5;) represents <italic>ReLu</italic> activation function. Fusion feature <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> is used to enhance the single-modal information because <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> combines the consistency and complementarity of the Color features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and Thermal features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub>. The use of <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> for enhancing the single-modal feature can reduce the interference of the noise in the single-modal features (for example, it is difficult to distinguish between the pedestrian information and the background noise).</p>
<p>In order to effectively enhance pedestrian features, the fusion feature <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> is sent into the channel attention module (CAM) and spatial attention module (PAM) [<xref ref-type="bibr" rid="B45">45</xref>] to make the network pay attention to pedestrian features. The network structure of CAM and PAM is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> first learns the channel attention weight <bold>
<italic>w</italic>
</bold>
<sub>
<italic>ca</italic>
</sub> &#x2208; R<sup>1&#xd7;1&#xd7;<italic>C</italic>
</sup> by CAM, then uses <bold>
<italic>w</italic>
</bold>
<sub>
<italic>ca</italic>
</sub> to weight <bold>
<italic>F</italic>
</bold>
<sub>
<italic>ct</italic>
</sub>, and the spatial attention weight <bold>
<italic>w</italic>
</bold>
<sub>
<italic>pa</italic>
</sub> &#x2208; R<sup>
<italic>W</italic>&#xd7;<italic>H</italic>&#xd7;1</sup> is obtained from the weighted features by PAM.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Network structure of channel attention and spatial attention.</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g003.tif"/>
</fig>
<p>The single-modal Color features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>c</italic>
</sub> and Thermal features <bold>
<italic>F</italic>
</bold>
<sub>
<italic>t</italic>
</sub> are multiplied element by element with the attention weights <bold>
<italic>w</italic>
</bold>
<sub>
<italic>ca</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>
<italic>pa</italic>
</sub> to enhance the single-modal features from the perspective of fused features. The whole process can be described as follows:<disp-formula id="e2">
<mml:math id="m3">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2297;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2297;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(2)</label>
</disp-formula>
<disp-formula id="e3">
<mml:math id="m4">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2297;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2297;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf2">
<mml:math id="m5">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf3">
<mml:math id="m6">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denote the Color features and Thermal features obtained by the cascaded information enhancement module, respectively. &#x2297; represents the element by element multiplication.</p>
</sec>
<sec id="s3-2">
<title>3.2 Cross-modal attention feature fusion module</title>
<p>There is an essential difference between Color and Thermal images, Color images reflect the color and texture detail information of pedestrians while Thermal images contain the temperature information of pedestrians, however, they also have some complementary information. In order to explore the complementary features of different image modalities and fuse them effectively, we design a cross-modal attention feature fusion module.</p>
<p>Specifically, the Color features <inline-formula id="inf4">
<mml:math id="m7">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and Thermal features <inline-formula id="inf5">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> enhanced by CIEM are first mapped into feature vectors <bold>
<italic>v</italic>
</bold>
<sub>
<italic>c</italic>
</sub> &#x2208; R<sup>1&#xd7;1&#xd7;<italic>C</italic>
</sup> and <bold>
<italic>v</italic>
</bold>
<sub>
<italic>t</italic>
</sub> &#x2208; R<sup>1&#xd7;1&#xd7;<italic>C</italic>
</sup>, respectively, by using global average pooling operation. The cross-modal attention network consists of a set of symmetric 1 &#xd7; 1 convolutions, <italic>ReLu</italic> activation functions, and <italic>Sigmoid</italic> activation functions. In order to obtain the complementary features of the two modalities, more pedestrian features need to be mined from the single-modal. The feature vectors <bold>
<italic>v</italic>
</bold>
<sub>
<italic>t</italic>
</sub> and <bold>
<italic>v</italic>
</bold>
<sub>
<italic>c</italic>
</sub> are learned to the respective modal attention weights <bold>
<italic>w</italic>
</bold>
<sub>
<italic>t</italic>
</sub> &#x2208; R<sup>1&#xd7;1&#xd7;<italic>C</italic>
</sup> and <bold>
<italic>w</italic>
</bold>
<sub>
<italic>c</italic>
</sub> &#x2208; R<sup>1&#xd7;1&#xd7;<italic>C</italic>
</sup> by a cross-modal attention network, and then the Color features <inline-formula id="inf6">
<mml:math id="m9">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are multiplied element by element with the attention weights <bold>
<italic>w</italic>
</bold>
<sub>
<italic>t</italic>
</sub> of the Thermal modality, and the Thermal features <inline-formula id="inf7">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are multiplied element by element with the attention weights <bold>
<italic>w</italic>
</bold>
<sub>
<italic>c</italic>
</sub> of the Color modality to complement the features of the other modality into the present modality. The specific process can be expressed as follows.<disp-formula id="e4">
<mml:math id="m11">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>u</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(4)</label>
</disp-formula>
<disp-formula id="e5">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2297;</mml:mo>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>u</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(6)</label>
</disp-formula>
<disp-formula id="e7">
<mml:math id="m14">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2297;</mml:mo>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>P</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf8">
<mml:math id="m15">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denotes Color features after supplementation with Thermal features, <inline-formula id="inf9">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denotes Thermal features after supplementation with Color features, <italic>GAP</italic>(&#x22c5;) denotes global average pooling operation, <italic>Conv</italic>
<sub>1</sub>(&#x22c5;) denotes convolution with convolution kernel size 1 &#xd7; 1, <italic>ReLu</italic>(&#x22c5;) denotes <italic>ReLu</italic> activation operation, and <italic>Sigmoid</italic> (&#x22c5;) denotes <italic>Sigmoid</italic> activation operation.</p>
<p>In order to efficiently fuse the two modal features, the features <inline-formula id="inf10">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf11">
<mml:math id="m18">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are subjected to an element by element addition operation to obtain a global feature vector containing Color and Thermal features. Then, the features <inline-formula id="inf12">
<mml:math id="m19">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf13">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are added element by element and multiplied with the attention weight <bold>
<italic>w</italic>
</bold>
<sub>
<italic>ct</italic>
</sub> of the global feature vector element by element to guide the fusion of Color and Thermal features from the perspective of global features to obtain the final fused feature <bold>
<italic>F</italic>
</bold>. The fused feature <bold>
<italic>F</italic>
</bold> is input to the detection head to obtain the pedestrian detection results. The feature fusion process can be expressed as follows:<disp-formula id="e8">
<mml:math id="m21">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>L</mml:mi>
<mml:mi>u</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2295;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(8)</label>
</disp-formula>
<disp-formula id="e9">
<mml:math id="m22">
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2297;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2295;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(9)</label>
</disp-formula>where &#x2295; denotes element by element addition.</p>
</sec>
<sec id="s3-3">
<title>3.3 Loss function</title>
<p>The loss function in this paper is consistent with the literature [<xref ref-type="bibr" rid="B26">26</xref>] and uses the Region Proposal Network (RPN) loss function <italic>L</italic>
<sub>
<italic>RPN</italic>
</sub> and Fast RCNN [<xref ref-type="bibr" rid="B46">46</xref>] loss function <italic>L</italic>
<sub>
<italic>FR</italic>
</sub> to jointly optimize the network:<disp-formula id="e10">
<mml:math id="m23">
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>Both <italic>L</italic>
<sub>
<italic>RPN</italic>
</sub> and <italic>L</italic>
<sub>
<italic>FR</italic>
</sub> consist of classification loss <italic>L</italic>
<sub>
<italic>cls</italic>
</sub> and bounding box regression loss <italic>L</italic>
<sub>
<italic>reg</italic>
</sub>:<disp-formula id="e11">
<mml:math id="m24">
<mml:mi>L</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>reg&#x2009;</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:munder>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(11)</label>
</disp-formula>Where, <italic>N</italic>
<sub>
<italic>cls</italic>
</sub> is the number of anchors, <italic>N</italic>
<sub>
<italic>reg</italic>
</sub> is the sum of positive and negative sample number, <italic>p</italic>
<sub>
<italic>i</italic>
</sub> is the probability that the <italic>i</italic>-th anchor is predicted to be the target, <inline-formula id="inf14">
<mml:math id="m25">
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is 1 when the anchor is a positive sample, otherwise it is 0, <italic>t</italic>
<sub>
<italic>i</italic>
</sub> denotes the bounding box regression parameter predicting the <italic>i</italic>-th anchor, and <inline-formula id="inf15">
<mml:math id="m26">
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denotes the GT bounding box parameter of the <italic>i</italic>-th anchor, <italic>&#x3bb;</italic> &#x3d; 1.</p>
<p>The difference between the classification loss of RPN network and Fast RCNN network is that the RPN network focuses only on the foreground and background when classifying, so its loss is a binary cross-entropy loss, while the Fast RCNN classification is focused to the target category and is a multi-category cross-entropy loss:<disp-formula id="e12">
<mml:math id="m27">
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>log</mml:mi>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
<p>The bounding box regression loss of RPN network and Fast RCNN network uses <inline-formula id="inf16">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:mtext>Smooth&#x2009;</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> loss:<disp-formula id="e13">
<mml:math id="m29">
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>reg&#x2009;</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(13)</label>
</disp-formula>Where, R denotes <inline-formula id="inf17">
<mml:math id="m30">
<mml:msub>
<mml:mrow>
<mml:mtext>Smooth&#x2009;</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> function,<disp-formula id="e14">
<mml:math id="m31">
<mml:msub>
<mml:mrow>
<mml:mtext>&#x2009;Smooth&#x2009;</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="array">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mtext>&#x2009;if&#x2009;</mml:mtext>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x3c;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>0.5</mml:mn>
</mml:mtd>
<mml:mtd columnalign="center">
<mml:mtext>&#x2009;otherwise&#x2009;</mml:mtext>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>The difference between the bounding box regression loss of RPN loss and the regression loss of Fast RCNN loss is that the RPN network is trained when <italic>&#x3c3;</italic> &#x3d; 3 and the Fast RCNN network is trained when <italic>&#x3c3;</italic> &#x3d; 1.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experimental results and analysis</title>
<sec id="s4-1">
<title>4.1 Datasets</title>
<p>This paper evaluates the algorithm performance on the KAIST pedestrian dataset [<xref ref-type="bibr" rid="B8">8</xref>], which is composed of 95,328 pairs of Color and Thermal images captured during daytime and nighttime. It is the most widely used multispectral pedestrian detection dataset at present. The dataset is labeled with four categories including person, people, person?, and cyclist. Considering the application areas of multispectral pedestrian detection (e.g., automatic driving), all four categories are treated as positive examples for detection in this paper. To address the problem of the annotation errors and missing annotations in the original annotation of the KAIST dataset, studies [<xref ref-type="bibr" rid="B9">9</xref>,<xref ref-type="bibr" rid="B28">28</xref>,<xref ref-type="bibr" rid="B47">47</xref>] performed data cleaning and re-annotation of the original data. Given that the annotations used in various studies are not consistent, we use 7601 pairs of Color and Thermal images from synthetic annotation (SA) [<xref ref-type="bibr" rid="B47">47</xref>] and 8892 pairs of Color and Thermal images from paired annotation (PA) [<xref ref-type="bibr" rid="B28">28</xref>] for model training. The test set consists of 2252 pairs of Color and Thermal images, of which 1455 pairs are from the daytime and 797 pairs are from the nighttime. For a fair comparison with other methods, the test experiments were performed according to the reasonable settings proposed in the literature [<xref ref-type="bibr" rid="B8">8</xref>].</p>
</sec>
<sec id="s4-2">
<title>4.2 Evaluation indexes</title>
<p>In this paper, Log-average Miss Rate (MR) proposed by [<xref ref-type="bibr" rid="B48">48</xref>] is employed as an evaluation index and combined with the plotting of the Miss Rate-FPPI curve to assess the effectiveness of the algorithm. The horizontal coordinate of the Miss Rate-FPPI curve indicates the average number of False Positives Per Image (FPPI), and the vertical coordinate represents the Miss Rate (MR), which is expressed as:<disp-formula id="e15">
<mml:math id="m32">
<mml:mtext>&#x2009;MissRate&#x2009;</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(15)</label>
</disp-formula>
<disp-formula id="e16">
<mml:math id="m33">
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>I</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>&#x2009;Total&#x2009;</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>&#x2009;images&#x2009;</mml:mtext>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(16)</label>
</disp-formula>where <italic>FN</italic> denotes False Negative, <italic>TP</italic> denotes True Positive, <italic>FP</italic> denotes False Positive, the sum of <italic>TP</italic> and <italic>FN</italic> is the number of all positive samples, and Total (images) denotes the total number of predicted images. It is worth noting that the lower the Miss Rate-FPPI curve trend, the better the detection performance; the smaller the MR value, the better the detection performance. In order to calculate MR, in logarithmic space, nine points are taken from the horizontal coordinate (limited value range is <inline-formula id="inf18">
<mml:math id="m34">
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>) of Miss Rate-FPPI curve, and then there are nine corresponding vertical coordinates <italic>m</italic>
<sub>1</sub>, <italic>m</italic>
<sub>2</sub>,&#x2026;<italic>m</italic>
<sub>9</sub>. By averaging these values, MR can be obtained as follows:<disp-formula id="e17">
<mml:math id="m35">
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">R</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="true">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(17)</label>
</disp-formula>where <italic>n</italic> is 9.</p>
</sec>
<sec id="s4-3">
<title>4.3 Implementation details</title>
<p>In this paper, the deep learning framework pytorch1.7 is adopted. The experimental platform is the ubuntu18.04 operating system and a single NVIDIA GeForce RTX 2080Ti GPU. Stochastic Gradient Descent (SGD) algorithm is used to optimize the network during model training, with momentum value of 0.9, weight attenuation value 5 &#xd7; 10<sup>&#x2013;4</sup>, and initial learning rate is 1 &#xd7; 10<sup>&#x2013;3</sup>. The model is iterated for five epochs with the batch size of 4, and the learning rate decay to 1 &#xd7; 10<sup>&#x2013;4</sup> after the 3rd epoch.</p>
</sec>
<sec id="s4-4">
<title>4.4 Experimental results and analysis</title>
<sec id="s4-4-1">
<title>4.4.1 Construction of the baseline</title>
<p>This work constructs a baseline algorithm architecture based on ResNet-101 backbone network and Faster RCNN detection head. Simple characteristic fusion (feature cascade, element by element addition and element by element multiplication) of the Color and Thermal features output by the backbone network is carried out in three sets of experiments. The fused feature is used as the input of the detection head. In order to ensure the high efficiency of the build baseline algorithm, synthesis annotation is employed to train and test the baseline. The test results are shown in <xref ref-type="table" rid="T1">Table 1</xref>. The MR values using feature cascade, element by element addition and element by element multiplication in the all-weather scene are 14.62%, 13.84% and 14.26%, respectively. By comparing these three results, it can be seen that the feature element by element addition demonstrates the best performance. Therefore, we adopt the method of adding features element by element as the baseline integration method.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Experimental results of baseline under different fusion modes.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Fusion modes</th>
<th align="center">All-weather</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">feature cascade</td>
<td align="center">14.62</td>
</tr>
<tr>
<td align="center">element by element multiplication</td>
<td align="center">14.26</td>
</tr>
<tr>
<td align="center">element by element addition</td>
<td align="center">
<bold>13.84</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values in highlight the optimal results for this column.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4-4-2">
<title>4.4.2 Performance comparison of different methods</title>
<p>The performance of this method is compared with several other state-of-the-art methods. The compared methods include hand-represented methods, e.g., ACT &#x2b; T &#x2b; THOG [<xref ref-type="bibr" rid="B8">8</xref>] and deep learning-based methods, e.g., Halfway Fusion [<xref ref-type="bibr" rid="B9">9</xref>], CMT_CNN[<xref ref-type="bibr" rid="B49">49</xref>], CIAN[<xref ref-type="bibr" rid="B50">50</xref>], IAF R-CNN[<xref ref-type="bibr" rid="B51">51</xref>], IATDNN &#x2b; IAMSS[<xref ref-type="bibr" rid="B19">19</xref>], CS-RCNN [<xref ref-type="bibr" rid="B52">52</xref>], IT-MN [<xref ref-type="bibr" rid="B53">53</xref>], and DCRD [<xref ref-type="bibr" rid="B54">54</xref>]. Here, the model is trained using 7601 pairs of Color and Thermal images from SA and 8892 pairs of Color and Thermal images from PA, respectively. Besides, 2252 pairs of Color and Thermal images from the test set are used for model testing. <xref ref-type="table" rid="T2">Table 2</xref> lists the experimental results.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>MRs of different methods on KAIST datasets.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">Methods</th>
<th colspan="3" align="center">SA</th>
<th colspan="3" align="center">PA(Color)</th>
<th colspan="3" align="center">PA(Thermal)</th>
</tr>
<tr>
<th align="center">All-weather</th>
<th align="center">Day</th>
<th align="center">Night</th>
<th align="center">All-weather</th>
<th align="center">Day</th>
<th align="center">Night</th>
<th align="center">All-weather</th>
<th align="center">Day</th>
<th align="center">Night</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">ACF &#x2b; T &#x2b; THOG</td>
<td align="center">41.65</td>
<td align="center">39.18</td>
<td align="center">48.29</td>
<td align="center">41.74</td>
<td align="center">39.30</td>
<td align="center">49.52</td>
<td align="center">41.36</td>
<td align="center">38.74</td>
<td align="center">48.30</td>
</tr>
<tr>
<td align="center">Halfway Fusion</td>
<td align="center">25.75</td>
<td align="center">24.88</td>
<td align="center">26.59</td>
<td align="center">25.10</td>
<td align="center">24.29</td>
<td align="center">26.12</td>
<td align="center">25.51</td>
<td align="center">25.20</td>
<td align="center">24.90</td>
</tr>
<tr>
<td align="center">CMT_CNN</td>
<td align="center">36.83</td>
<td align="center">34.56</td>
<td align="center">41.82</td>
<td align="center">36.25</td>
<td align="center">34.12</td>
<td align="center">41.21</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
</tr>
<tr>
<td align="center">IAF R-CNN</td>
<td align="center">15.73</td>
<td align="center">14.55</td>
<td align="center">18.26</td>
<td align="center">15.65</td>
<td align="center">14.95</td>
<td align="center">18.11</td>
<td align="center">16.00</td>
<td align="center">15.22</td>
<td align="center">17.56</td>
</tr>
<tr>
<td align="center">IATDNN &#x2b; IAMSS</td>
<td align="center">14.95</td>
<td align="center">14.67</td>
<td align="center">15.72</td>
<td align="center">15.14</td>
<td align="center">14.82</td>
<td align="center">15.87</td>
<td align="center">15.08</td>
<td align="center">15.02</td>
<td align="center">15.20</td>
</tr>
<tr>
<td align="center">CIAN</td>
<td align="center">14.12</td>
<td align="center">14.77</td>
<td align="center">11.13</td>
<td align="center">14.64</td>
<td align="center">15.13</td>
<td align="center">12.43</td>
<td align="center">14.68</td>
<td align="center">16.21</td>
<td align="center">9.88</td>
</tr>
<tr>
<td align="center">CS-RCNN</td>
<td align="center">11.43</td>
<td align="center">11.86</td>
<td align="center">8.82</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
</tr>
<tr>
<td align="center">IT-MN</td>
<td align="center">14.19</td>
<td align="center">14.30</td>
<td align="center">13.98</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
</tr>
<tr>
<td align="center">DCRD</td>
<td align="center">12.58</td>
<td align="center">13.12</td>
<td align="center">11.65</td>
<td align="center">13.64</td>
<td align="center">13.15</td>
<td align="center">13.98</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
<td align="center">&#x2013;</td>
</tr>
<tr>
<td align="center">Ours</td>
<td align="center">10.71</td>
<td align="center">13.09</td>
<td align="center">8.45</td>
<td align="center">11.11</td>
<td align="center">12.85</td>
<td align="center">8.77</td>
<td align="center">10.98</td>
<td align="center">13.07</td>
<td align="center">8.53</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="table" rid="T2">Table 2</xref> shows that when the model is trained with SA, the MRs of the method proposed in this paper are 10.71%, 13.09% and 8.45% for all-weather, daytime and nighttime scenes, respectively, which are 0.72%, &#x2212;1.23% and 0.37% lower than the compared method CS-RCNN with the best performance, respectively. The PA (Color) and PA (Thermal) in <xref ref-type="table" rid="T2">Table 2</xref> represent the Color annotation and Thermal annotation in the pairwise annotation PA, respectively, for the purpose of training the model. It can be seen from two that the MRs of the method in this paper are 11.11% and 10.98% when using Color annotation and Thermal annotation in the all-weather scene, which are 2.53% and 3.70%, respectively, lower than those of compared method with the best performance. In addition, by analyzing the experimental results of two improved versions of annotations, it can be found that pedestrian detection results are different when using different annotations, indicating the importance of annotations.</p>
</sec>
<sec id="s4-4-3">
<title>4.4.3 Analysis of ablation experiments</title>
<sec id="s4-4-3-1">
<title>4.4.3.1 Complementarity and importance of color and thermal features</title>
<p>This section compares the effect of different input sources on pedestrian detection performance. In order to eliminate the impact of the proposed module on detection performance, three sets of experiments are conducted on baseline: 1) the combination of Color and Thermal images as the input source (the input of the two branches of the backbone network are respectively Color and Thermal images); 2) dual-stream Color image as the input source (use Color images to replace Thermal images, that is, the backbone network input source is Color images); 3) dual-stream Thermal images as the input source (use Thermal images to replace Color images, that is, the backbone network input source is Thermal images).The training set of the model here is 7061 pairs of images of SA, and the test set is 2252 pairs of Color and Thermal images. <xref ref-type="table" rid="T3">Table 3</xref> shows the MRs of these three input sources for the all-weather, daytime, and nighttime scenes. It can be seen from <xref ref-type="table" rid="T3">Table 3</xref> that the MRs obtained using Color and Thermal images as input to the network are 13.84%, 15.35% and 12.48% for the all-weather, daytime and nighttime scenes, respectively, which are 11.53%, 3.96%, 18.70% and 3.71%, 7.46%, 0.13% lower than using Color images and Thermal images as input alone. The experimental results prove that the detection network combining Color and Thermal features delivers better performance, indicating that Color and Thermal features are important for pedestrian detection.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>MRs of different modal inputs.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Input</th>
<th align="center">All-weather</th>
<th align="center">Day</th>
<th align="center">Night</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">dual-stream Color images</td>
<td align="center">25.37</td>
<td align="center">19.31</td>
<td align="center">31.18</td>
</tr>
<tr>
<td align="center">dual-stream Thermal images</td>
<td align="center">17.55</td>
<td align="center">22.81</td>
<td align="center">12.61</td>
</tr>
<tr>
<td align="center">Color images &#x2b; Thermal images</td>
<td align="center">
<bold>13.84</bold>
</td>
<td align="center">
<bold>15.35</bold>
</td>
<td align="center">
<bold>12.48</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values in highlight the optimal results for this column.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>
<xref ref-type="fig" rid="F4">Figure 4</xref> shows the Miss Rate-FPPI curves of the detection results for these three input sources in the all-weather, daytime, and nighttime scenes (blue, red and green curves indicate dual-stream Thermal images, dual-stream Color images, and Color and Thermal images, respectively). By analyzing the Miss Rate-FPPI curve trend and combining with the experimental data in <xref ref-type="table" rid="T3">Table 3</xref>, it can be seen that the detection effect of Color images as the input source is better than that of Thermal images in the daytime scene while the result is the opposite for the night scene, and the detection effect of Color and Thermal images combined as the input source is better than that of single-modal input in both daytime and nighttime. It shows that there are complementary features between Color and Thermal modalities, and the fusion of the two modal features can improve the pedestrian detection performance.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The Miss Rate-FPPI curves of the detection results of the three groups of input sources in the All-weather, Daytime and Nighttime scenes (From left to right, All-weather, Daytime and Nighttime Miss Rate-FPPI curves are shown in the figure).</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g004.tif"/>
</fig>
</sec>
<sec id="s4-4-3-2">
<title>4.4.3.2 Ablation experiments</title>
<p>In this section, ablation experiments are conducted to demonstrate the effectiveness of the proposed cascaded information enhancement module (CIEM) and cross-modal attentional feature fusion module (CAFFM). Here, 7061 pairs of SA images are used to train the model, and 2252 pairs of Color and Thermal images in the test set are used to test the model.</p>
<p>Effectiveness of CIEM: CIEM is used to enhance the pedestrian features in Color and Thermal images to reduce the interference from the background. The experimental results are shown in <xref ref-type="table" rid="T4">Table 4</xref>. The MRs of baseline on SA are 13.84%, 15.35% and 12.48% for all-weather, daytime and nighttime scenes, respectively. When CIEM is additionally employed, the MRs are 11.21%, 13.15% and 9.07% for all-weather, daytime and nighttime scenes, respectively, which are reduced by 2.63%, 2.20% and 3.41% compared to the baseline, respectively. It is shown that the proposed CIEM effectively enhances the pedestrian features in both modalities, reduces the interference of background, and improves the pedestrian detection performance.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>MRs for ablation studies of the proposed method on SA.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="center">Methods</th>
<th align="center">All-weather</th>
<th align="center">Day</th>
<th align="center">Night</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">Baseline</td>
<td align="center">13.84</td>
<td align="center">15.35</td>
<td align="center">12.48</td>
</tr>
<tr>
<td align="center">baseline &#x2b; CIEM</td>
<td align="center">11.21</td>
<td align="center">13.15</td>
<td align="center">9.07</td>
</tr>
<tr>
<td align="center">baseline &#x2b; CAFFM</td>
<td align="center">11.68</td>
<td align="center">13.81</td>
<td align="center">9.50</td>
</tr>
<tr>
<td align="center">Overall model</td>
<td align="center">
<bold>10.71</bold>
</td>
<td align="center">
<bold>13.09</bold>
</td>
<td align="center">
<bold>8.45</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The bold values in highlight the optimal results for this column.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Validity of CAFFM: CAFFM is used to effectively fuse Color and Thermal features. The experimental results are shown in <xref ref-type="table" rid="T4">Table 4</xref>. On the SA, when the baseline is used with CAFFM, the MRs are 11.68%, 13.81% and 9.50% in all-weather, daytime and nighttime scenes, respectively, which are reduced by 2.16%, 1.54% and 2.98% compared baseline, respectively. It shows that the proposed CAFFM effectively fuses the two modal features to achieve robust multispectral pedestrian detection.</p>
<p>Overall effectiveness: The proposed CIEM and CAFFM are additionally used on the basis of baseline. Experimental results show a reduction of 3.13%, 2.26% and 4.03% in MRs for all-weather, daytime and nighttime scenes, respectively, compared to the baseline, indicating the overall effectiveness of the proposed method. A closer look reveals that with additional employment of CIEM and CAFFM alone, MRs are decreased by 2.63% and 2.16%, respectively, in the all-weather scene, but the MR of the overall model is reduced by 3.13%. It demonstrates that there is some orthogonal complementarity in the role of the proposed two modules.</p>
<p>
<xref ref-type="fig" rid="F5">Figure 5</xref> shows the Miss Rate-FPPI curves for CIEM and CAFFM ablation studies in all-weather, daytime and nighttime scenes (blue, red, orange and green curves represent baseline, baseline &#x2b; CIEM, baseline &#x2b; CAFFM and overall model, respectively). It is clear that the curve trends of each module and the overall model are both lower than that of the baseline, which further proves the effectiveness of the method presented in this work.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The Miss Rate-FPPI curves of CIEM and CAFFM ablation studies in All-weather, Daytime and Nighttime scenes (From left to right, All-weather, Daytime and Nighttime Miss Rate-FPPI curves are shown in the figure).</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g005.tif"/>
</fig>
<p>Furthermore, in order to qualitatively analyze the effectiveness of the proposed CIEM and CAFFM, four pairs of Color and Thermal images (two pairs of images are taken from daytime and two pairs of images are taken from nighttime) are selected from the test set for testing. The pedestrian detection results of the baseline and each proposed module are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. The first row is the visualization results of labeled boxes for Color and Thermal images, and the second to the fifth rows are the visualization results of the labeled and prediction boxes for baseline, baseline &#x2b; CIEM, baseline &#x2b; CAFFM, and the overall model pedestrian detection with the green and red boxes representing the labeled and prediction boxes, respectively. It can be seen that the proposed method successfully addresses the problem of pedestrian missing detection in complex environments and achieves more accurate detection boxes. For example, the second row, pedestrian detection missing happens in the first, third, and fourth pairs of images in the baseline detection result, however, the pedestrian miss detection problem is properly solved with CIEM and CAFFM added to the baseline and the overall model produces more accurate pedestrian detection boxes.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>In this paper, each module and baseline pedestrian detection results (The first row is the visualization results of labeled boxes for Color and Thermal images, and the second to the fifth rows are the visualization results of the labeled and prediction boxes for baseline, baseline &#x2b; CIEM, baseline &#x2b; CAFFM and the overall model pedestrian detection with the green and red boxes representing the labeled and prediction boxes, respectively).</p>
</caption>
<graphic xlink:href="fphy-11-1121311-g006.tif"/>
</fig>
</sec>
</sec>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>5 Conclusion</title>
<p>In this paper, we propose a multispectral pedestrian detection algorithm including cascaded information enhancement module and cross-modal attention feature fusion module. The proposed method improves the accuracy of pedestrian detection in multispectral images (Color and Thermal images) by effectively fusing the features from the two modules and augmenting the pedestrian features. Specifically, on the one hand, a cascaded information enhancement module (CIEM) is designed to enhance single-modal features to enrich the pedestrian features and suppress interference from the background noise. On the other hand, unlike previous methods that simply splice Color and Thermal features directly, a cross-modal attention feature fusion module (CAFFM) is introduced to mine the features of both Color and Thermal modalities and to complement each other, then complementary enhanced modal features are used to construct global features. Extensive experiments have been conducted on two improved annotations of the public dataset KAIST. The experimental results show that the proposed method is conducive to obtain more comprehensive pedestrian features and improve the accuracy of multispectral image pedestrian detection.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://gitcode.net/mirrors/soonminhwang/rgbt-ped-detection?utm_source=csdn_github_accelerator">https://gitcode.net/mirrors/soonminhwang/rgbt-ped-detection?utm_source&#x3d;csdn_github_accelerator</ext-link>.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>YY responsible for scheme design, experiment and writing of the paper. KX guide the scheme design and experiment of the paper. KW guide experimental data analysis, paper writing and modification.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the National Natural Science Foundation of China (No. 52107017) and Fundamental Research Fund of Science and Technology Department of Yunnan Province(No.202201AU070172).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jeong</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Ko</surname>
<given-names>BC</given-names>
</name>
<name>
<surname>Nam</surname>
<given-names>J-Y</given-names>
</name>
</person-group>. <article-title>Early detection of sudden pedestrian crossing for safe driving during summer nights</article-title>. <source>IEEE Trans Circuits Syst Video Technol</source> (<year>2017</year>) <volume>27</volume>:<fpage>1368</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2016.2539684</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>Y</given-names>
</name>
<etal/>
</person-group> <article-title>Pedestrian search in surveillance videos by learning discriminative deep features</article-title>. <source>Neurocomputing</source> (<year>2018</year>) <volume>283</volume>:<fpage>120</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2017.12.042</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>Unsupervised domain adaptive person re-identification guided by low-rank priori</article-title>. <source>J Chongqing Univ</source> (<year>2021</year>) <volume>44</volume>:<fpage>57</fpage>&#x2013;<lpage>70</lpage>. <pub-id pub-id-type="doi">10.11835/j.issn.1000-582X.2021.11.008</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Attribute-aligned domain-invariant feature learning for unsupervised domain adaptation person re-identification</article-title>. <source>IEEE Trans Inf Forensics Security</source> (<year>2021</year>) <volume>16</volume>:<fpage>1480</fpage>&#x2013;<lpage>94</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2020.3036800</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Triple adversarial learning and multi-view imaginative reasoning for unsupervised domain adaptation person re-identification</article-title>. <source>IEEE Trans Circuits Syst Video Technol</source> (<year>2022</year>) <volume>32</volume>:<fpage>2814</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2021.3099943</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Mutual prediction learning and mixed viewpoints for unsupervised-domain adaptation person re-identification on blockchain</article-title>. <source>Simulation Model Pract Theor</source> (<year>2022</year>) <volume>119</volume>:<fpage>102568</fpage>. <pub-id pub-id-type="doi">10.1016/j.simpat.2022.102568</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Occluded person re-identification via defending against attacks from obstacles</article-title>. <source>IEEE Trans Inf Forensics Security</source> (<year>2023</year>) <volume>18</volume>:<fpage>147</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2022.3218449</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Hwang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Kweon</surname>
<given-names>IS</given-names>
</name>
</person-group>. <article-title>Multispectral pedestrian detection: Benchmark dataset and baseline</article-title>. In: <conf-name>2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <conf-date>07-12 June 2015</conf-date>; <conf-loc>Boston, MA, USA</conf-loc> (<year>2015</year>). p. <fpage>1037</fpage>&#x2013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2015.7298706</pub-id>
</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Metaxas</surname>
<given-names>DN</given-names>
</name>
</person-group>. <article-title>Multispectral deep neural networks for pedestrian detection</article-title>. In: <conf-name>Proceedings of the British Machine Vision Conference 2016</conf-name>; <conf-date>19-22 September 2016</conf-date>; <conf-loc>York, UK</conf-loc> (<year>2016</year>).</citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gonz&#xe1;lez</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Socarras</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Serrat</surname>
<given-names>J</given-names>
</name>
<name>
<surname>V&#xe1;zquez</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J</given-names>
</name>
<etal/>
</person-group> <article-title>Pedestrian detection at day/night time with visible and fir cameras: A comparison</article-title>. <source>Sensors</source> (<year>2016</year>) <volume>16</volume>:<fpage>820</fpage>. <pub-id pub-id-type="doi">10.3390/s16060820</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Analysis-synthesis dictionary pair learning and patch saliency measure for image fusion</article-title>. <source>Signal Process.</source> (<year>2020</year>) <volume>167</volume>:<fpage>107327</fpage>. <pub-id pub-id-type="doi">10.1016/j.sigpro.2019.107327</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>Multi-focus image fusion: A survey of the state of the art</article-title>. <source>Inf Fusion</source> (<year>2020</year>) <volume>64</volume>:<fpage>71</fpage>&#x2013;<lpage>91</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2020.06.013</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>He</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R</given-names>
</name>
</person-group>. <article-title>Joint medical image fusion, denoising and enhancement via discriminative low-rank sparse dictionaries learning</article-title>. <source>Pattern Recognition</source> (<year>2018</year>) <volume>79</volume>:<fpage>130</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2018.02.005</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D</given-names>
</name>
</person-group>. <article-title>Discriminative dictionary learning-based multiple component decomposition for detail-preserving noisy image fusion</article-title>. <source>IEEE Trans Instrumentation Meas</source> (<year>2020</year>) <volume>69</volume>:<fpage>1082</fpage>&#x2013;<lpage>102</lpage>. <pub-id pub-id-type="doi">10.1109/tim.2019.2912239</pub-id>
</citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>A unified framework for damaged image fusion and completion based on low-rank and sparse decomposition</article-title>. <source>Signal Processing: Image Commun</source> (<year>2021</year>) <volume>29</volume>:<fpage>116400</fpage>. <pub-id pub-id-type="doi">10.1016/j.image.2021.116400</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Tao</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Key point-aware occlusion suppression and semantic alignment for occluded person re-identification</article-title>. <source>Inf Sci</source> (<year>2022</year>) <volume>606</volume>:<fpage>669</fpage>&#x2013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2022.05.077</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Mazur</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>C</given-names>
</name>
<etal/>
</person-group> <article-title>Camera style transformation with preserved self-similarity and domain-dissimilarity in unsupervised person re-identification</article-title>. <source>J Vis Commun Image Representation</source> (<year>2021</year>) <volume>80</volume>:<fpage>103303</fpage>. <pub-id pub-id-type="doi">10.1016/j.jvcir.2021.103303</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Qian</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>M</given-names>
</name>
</person-group>. <article-title>Baanet: Learning bi-directional adaptive attention gates for multispectral pedestrian detection</article-title>. In: <conf-name>2022 International Conference on Robotics and Automation (ICRA)</conf-name>; <conf-date>23-27 May 2022</conf-date>; <conf-loc>Philadelphia, PA, USA</conf-loc> (<year>2022</year>). p. <fpage>2920</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA46639.2022.9811999</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guan</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>MY</given-names>
</name>
</person-group>. <article-title>Fusion of multispectral data through illumination-aware deep neural networks for pedestrian detection</article-title>. <source>Inf Fusion</source> (<year>2019</year>) <volume>50</volume>:<fpage>148</fpage>&#x2013;<lpage>57</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2018.11.017</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>Improving multispectral pedestrian detection by addressing modality imbalance problems</article-title>. In: <source>European conference on computer vision</source>. <publisher-loc>Berlin, Germany</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2020</year>). p. <fpage>787</fpage>&#x2013;<lpage>803</lpage>.</citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Confidence-aware fusion using dempster-shafer theory for multispectral pedestrian detection</article-title>. <source>IEEE Trans Multimedia</source> (<year>2022</year>) <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/tmm.2022.3160589</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Fromont</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Lefevre</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Avignon</surname>
<given-names>B</given-names>
</name>
</person-group>. <article-title>Multispectral fusion for object detection with cyclic fuse-and-refine blocks</article-title>. In: <conf-name>2020 IEEE International Conference on Image Processing (ICIP)</conf-name>; <conf-date>25-28 October 2020</conf-date>; <conf-loc>Abu Dhabi, United Arab Emirates</conf-loc> (<year>2020</year>). p. <fpage>276</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP40778.2020.9191080</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doll&#xe1;r</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Appel</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Belongie</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Perona</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Fast feature pyramids for object detection</article-title>. <source>IEEE Trans Pattern Anal Machine Intelligence</source> (<year>2014</year>) <volume>36</volume>:<fpage>1532</fpage>&#x2013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2014.2300479</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dalal</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Triggs</surname>
<given-names>B</given-names>
</name>
</person-group>. <article-title>Histograms of oriented gradients for human detection</article-title>. In: <conf-name>2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR&#x2019;05)</conf-name>; <conf-date>20-25 June 2005</conf-date>; <conf-loc>San Diego, CA, USA</conf-loc> (<year>2005</year>). p. <fpage>886</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2005.177</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>K&#xf6;nig</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Adam</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Jarvers</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Layher</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Neumann</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Teutsch</surname>
<given-names>M</given-names>
</name>
</person-group>. <article-title>Fully convolutional region proposal networks for multispectral person detection</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)</conf-name>; <conf-date>21-26 July 2017</conf-date>; <conf-loc>Honolulu, HI, USA</conf-loc> (<year>2017</year>). p. <fpage>243</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1109/CVPRW.2017.36</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ren</surname>
<given-names>S</given-names>
</name>
<name>
<surname>He</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Girshick</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Faster r-cnn: Towards real-time object detection with region proposal networks</article-title>. <source>IEEE Trans Pattern Anal Machine Intelligence</source> (<year>2017</year>) <volume>39</volume>:<fpage>1137</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2577031</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>A fusion framework for multi-spectral pedestrian detection using efficientdet</article-title>. In: <conf-name>2021 21st International Conference on Control, Automation and Systems (ICCAS)</conf-name>; <conf-date>12-15 October 2021</conf-date>; <conf-loc>Jeju, Korea</conf-loc> (<year>2021</year>). p. <fpage>1111</fpage>&#x2013;<lpage>3</lpage>. <pub-id pub-id-type="doi">10.23919/ICCAS52745.2021.9650057</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Weakly aligned cross-modal learning for multispectral pedestrian detection</article-title>. In: <conf-name>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</conf-name>; <conf-date>27 October 2019 - 02 November 2019</conf-date>; <conf-loc>Seoul, Korea (South)</conf-loc> (<year>2019</year>). p. <fpage>5126</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2019.00523</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Mlpd: Multi-label pedestrian detector in multispectral domain</article-title>. <source>IEEE Robotics Automation Lett</source> (<year>2021</year>) <volume>6</volume>:<fpage>7846</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1109/LRA.2021.3099870</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>JU</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Ro</surname>
<given-names>YM</given-names>
</name>
</person-group>. <article-title>Uncertainty-guided cross-modal learning for robust multispectral pedestrian detection</article-title>. <source>IEEE Trans Circuits Syst Video Technol</source> (<year>2022</year>) <volume>32</volume>:<fpage>1510</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2021.3076466</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Vaswani</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Shazeer</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Parmar</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Uszkoreit</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Gomez</surname>
<given-names>AN</given-names>
</name>
<etal/>
</person-group> <article-title>Attention is all you need</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>; <conf-date>December 4-9, 2017</conf-date>; <conf-loc>Long Beach California USA</conf-loc>. <publisher-name>Curran Associates Inc.</publisher-name> (<year>2017</year>).</citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Attention-based multi-modal fusion network for semantic scene completion</article-title>. <source>Proc AAAI Conf Artif Intelligence</source> (<year>2020</year>) <volume>34</volume>:<fpage>11402</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v34i07.6803</pub-id>
</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Image emotion caption based on visual attention mechanisms</article-title>. In: <conf-name>2020 IEEE 6th International Conference on Computer and Communications (ICCC)</conf-name>; <conf-date>11-14 December 2020</conf-date>; <conf-loc>Chengdu, China</conf-loc>. <publisher-name>IEEE</publisher-name> (<year>2020</year>). p. <fpage>1456</fpage>&#x2013;<lpage>60</lpage>.</citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiao</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Heterogeneous knowledge distillation for simultaneous infrared-visible image fusion and super-resolution</article-title>. <source>IEEE Trans Instrumentation Meas</source> (<year>2022</year>) <volume>71</volume>:<fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1109/tim.2022.3149101</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Cen</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Different input resolutions and arbitrary output resolution: A meta learning-based deep framework for infrared and visible image fusion</article-title>. <source>IEEE Trans Image Process</source> (<year>2021</year>) <volume>30</volume>:<fpage>4070</fpage>&#x2013;<lpage>83</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2021.3069339</pub-id>
</citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Haze transfer and feature aggregation network for real-world single image dehazing</article-title>. <source>Knowledge-Based Syst</source> (<year>2022</year>) <volume>251</volume>:<fpage>109309</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109309</pub-id>
</citation>
</ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Multi-stream attention-aware graph convolution network for video salient object detection</article-title>. <source>IEEE Trans Image Process</source> (<year>2021</year>) <volume>30</volume>:<fpage>4183</fpage>&#x2013;<lpage>97</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2021.3070200</pub-id>
</citation>
</ref>
<ref id="B38">
<label>38.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
</person-group>. <article-title>Dual-stream reciprocal disentanglement learning for domain adaptation person re-identification</article-title>. <source>Knowledge-Based Syst</source> (<year>2022</year>) <volume>251</volume>:<fpage>109315</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109315</pub-id>
</citation>
</ref>
<ref id="B39">
<label>39.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Cross-compatible embedding and semantic consistent feature construction for sketch re-identification</article-title>. In: <conf-name>MM &#x2019;22. Proceedings of the 30th ACM International Conference on Multimedia</conf-name>; <conf-date>10 October 2022</conf-date>; <conf-loc>New York, NY, USA</conf-loc>. <publisher-name>Association for Computing Machinery</publisher-name> (<year>2022</year>). p. <fpage>3347</fpage>&#x2013;<lpage>55</lpage>. <pub-id pub-id-type="doi">10.1145/3503161.3548224</pub-id>
</citation>
</ref>
<ref id="B40">
<label>40.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Chai</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Body part-level domain alignment for domain-adaptive person re-identification with transformer framework</article-title>. <source>IEEE Trans Inf Forensics Security</source> (<year>2022</year>) <volume>17</volume>:<fpage>3321</fpage>&#x2013;<lpage>34</lpage>. <pub-id pub-id-type="doi">10.1109/TIFS.2022.3207893</pub-id>
</citation>
</ref>
<ref id="B41">
<label>41.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Albanie</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>E</given-names>
</name>
</person-group>. <article-title>Squeeze-and-excitation networks</article-title>. <source>IEEE Trans Pattern Anal Machine Intelligence</source> (<year>2020</year>) <volume>42</volume>:<fpage>2011</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2019.2913372</pub-id>
</citation>
</ref>
<ref id="B42">
<label>42.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Selective kernel networks</article-title>. In: <conf-name>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <conf-date>15-20 June 2019</conf-date>; <conf-loc>Long Beach, CA, USA</conf-loc> (<year>2019</year>). p. <fpage>510</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2019.00060</pub-id>
</citation>
</ref>
<ref id="B43">
<label>43.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dai</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Gieseke</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Oehmcke</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Barnard</surname>
<given-names>K</given-names>
</name>
</person-group>. <article-title>Attentional feature fusion</article-title>. In: <conf-name>2021 IEEE Winter Conference on Applications of Computer Vision (WACV)</conf-name>; <conf-date>03-08 January 2021</conf-date>; <conf-loc>Waikoloa, HI, USA</conf-loc> (<year>2021</year>). p. <fpage>3559</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.1109/WACV48630.2021.00360</pub-id>
</citation>
</ref>
<ref id="B44">
<label>44.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <conf-date>27-30 June 2016</conf-date>; <conf-loc>NV, USA</conf-loc> (<year>2016</year>). p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
</citation>
</ref>
<ref id="B45">
<label>45.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Woo</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>J-Y</given-names>
</name>
<name>
<surname>Kweon</surname>
<given-names>IS</given-names>
</name>
</person-group>. <article-title>Cbam: Convolutional block attention module</article-title>. In: <conf-name>Proceedings of the European conference on computer vision (ECCV)</conf-name> (<year>2018</year>). p. <fpage>3</fpage>&#x2013;<lpage>19</lpage>.</citation>
</ref>
<ref id="B46">
<label>46.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Girshick</surname>
<given-names>R</given-names>
</name>
</person-group>. <article-title>Fast r-cnn</article-title>. In: <conf-name>2015 IEEE International Conference on Computer Vision (ICCV)</conf-name>; <conf-date>07-13 December 2015</conf-date>; <conf-loc>Santiago, Chile</conf-loc> (<year>2015</year>). p. <fpage>1440</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2015.169</pub-id>
</citation>
</ref>
<ref id="B47">
<label>47.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>M</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Multispectral pedestrian detection via simultaneous detection and segmentation</article-title>. <comment>
<italic>arXiv preprint arXiv:1808.04818</italic>
</comment>
</citation>
</ref>
<ref id="B48">
<label>48.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dollar</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Wojek</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Schiele</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Perona</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Pedestrian detection: An evaluation of the state of the art</article-title>. <source>IEEE Trans Pattern Anal Machine Intelligence</source> (<year>2012</year>) <volume>34</volume>:<fpage>743</fpage>&#x2013;<lpage>61</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2011.155</pub-id>
</citation>
</ref>
<ref id="B49">
<label>49.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Ouyang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Ricci</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Sebe</surname>
<given-names>N</given-names>
</name>
</person-group>. <article-title>Learning cross-modal deep representations for robust pedestrian detection</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</conf-name>; <conf-date>26 Jul 2017</conf-date>; <conf-loc>Hawaii</conf-loc> (<year>2017</year>). p. <fpage>4236</fpage>&#x2013;<lpage>44</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.451</pub-id>
</citation>
</ref>
<ref id="B50">
<label>50.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Qiao</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>K</given-names>
</name>
<etal/>
</person-group> <article-title>Cross-modality interactive attention network for multispectral pedestrian detection</article-title>. <source>Inf Fusion</source> (<year>2019</year>) <volume>50</volume>:<fpage>20</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1016/j.inffus.2018.09.015</pub-id>
</citation>
</ref>
<ref id="B51">
<label>51.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>M</given-names>
</name>
</person-group>. <article-title>Illumination-aware faster r-cnn for robust multispectral pedestrian detection</article-title>. <source>Pattern Recognition</source> (<year>2019</year>) <volume>85</volume>:<fpage>161</fpage>&#x2013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/j.patcog.2018.08.005</pub-id>
</citation>
</ref>
<ref id="B52">
<label>52.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Yin</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Nie</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Attention based multi-layer fusion of multispectral images for pedestrian detection</article-title>. <source>IEEE Access</source> (<year>2020</year>) <volume>8</volume>:<fpage>165071</fpage>&#x2013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3022623</pub-id>
</citation>
</ref>
<ref id="B53">
<label>53.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhuang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Pu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Illumination and temperature-aware multispectral networks for edge-computing-enabled pedestrian detection</article-title>. <source>IEEE Trans Netw Sci Eng</source> (<year>2022</year>) <volume>9</volume>:<fpage>1282</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/TNSE.2021.3139335</pub-id>
</citation>
</ref>
<ref id="B54">
<label>54.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>T</given-names>
</name>
<name>
<surname>Lam</surname>
<given-names>K-M</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Qiu</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Deep cross-modal representation learning and distillation for illumination-invariant pedestrian detection</article-title>. <source>IEEE Trans Circuits Syst Video Technol</source> (<year>2022</year>) <volume>32</volume>:<fpage>315</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2021.3060162</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>