<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="brief-report">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Comput. Sci.</journal-id>
<journal-title>Frontiers in Computer Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Comput. Sci.</abbrev-journal-title>
<issn pub-type="epub">2624-9898</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcomp.2022.1062792</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Computer Science</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Real-time multiple target segmentation with multimodal few-shot learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Khoshboresh-Masouleh</surname> <given-names>Mehdi</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1961258/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Shah-Hosseini</surname> <given-names>Reza</given-names></name>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2096807/overview"/>
</contrib>
</contrib-group>
<aff><institution>School of Surveying and Geospatial Engineering, College of Engineering, University of Tehran</institution>, <addr-line>Tehran</addr-line>, <country>Iran</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mingming Gong, The University of Melbourne, Australia</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Haopeng Li, University of Melbourne, Australia; Zhaoqing Wang, The University of Sydney, Australia; Dongting Hu, The University of Melbourne, Australia</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Mehdi Khoshboresh-Masouleh <email>m.khoshboresh&#x00040;ut.ac.ir</email></corresp>
<corresp id="c002">Reza Shah-Hosseini <email>rshahosseini&#x00040;ut.ac.ir</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Computer Vision, a section of the journal Frontiers in Computer Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>11</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>4</volume>
<elocation-id>1062792</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>11</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Khoshboresh-Masouleh and Shah-Hosseini.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Khoshboresh-Masouleh and Shah-Hosseini</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Deep learning-based target segmentation requires a big training dataset to achieve good results. In this regard, few-shot learning a model that quickly adapts to new targets with a few labeled support samples is proposed to tackle this issue. In this study, we introduce a new multimodal few-shot learning [e.g., red-green-blue (RGB), thermal, and depth] for real-time multiple target segmentation in a real-world application with a few examples based on a new squeeze-and-attentions mechanism for multiscale and multiple target segmentation. Compared to the state-of-the-art methods (HSNet, CANet, and PFENet), the proposed method demonstrates significantly better performance on the PST900 dataset with 32 time-series sets in both Hand-Drill, and Survivor classes.</p>
</abstract>
<kwd-group>
<kwd>few-shot learning</kwd>
<kwd>multimodal images</kwd>
<kwd>target detection</kwd>
<kwd>real-time processing</kwd>
<kwd>squeeze-and-attention CNN</kwd>
</kwd-group>
<counts>
<fig-count count="3"/>
<table-count count="3"/>
<equation-count count="3"/>
<ref-count count="22"/>
<page-count count="7"/>
<word-count count="3569"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>Real-time multiple target segmentation is an important task for real-world applications (Morelande et al., <xref ref-type="bibr" rid="B8">2007</xref>; Wagner et al., <xref ref-type="bibr" rid="B14">2009</xref>), such as search and rescue robots. In this regard, the quadruped mobile robot based on multimodal images can provide more comprehensive spatial and spectral information for the development of real-time multiple target segmentation (Rahman et al., <xref ref-type="bibr" rid="B9">2021</xref>). Multimodal target segmentation is challenging due to the various background, shadows, and occluded areas, and multiscale targets. The goal of multimodal imaging is to improve detection and localization of objects in complex scenes (Mart&#x000ED;-Bonmat&#x000ED; et al., <xref ref-type="bibr" rid="B6">2010</xref>). In multimodality imaging, the need to combine different information can be approached by either acquiring images at different times. In this regard, the image fusion is the process of merging data from multiple imaging modalities [e.g., red-green-blue (RGB), thermal, and depth] to obtain a fused image with a large amount of information for increasing the scene understanding applicability. Multiple target segmentation with the use of multimodal data for time-series images potentially improves scene understanding with a limited amount of labeled training data, while many target segmentation methods appear to understand single-time localization with a big training dataset. Multimodal few-shot learning can perform on unseen tasks after training a few annotated data and considers several tasks to produce a predictive function, and is an inductive transfer system whose main goal is to improve generalization ability for multiple targets. These approaches excel at learning complicated features from small set using weakly-supervised learning.</p>
<p>Deep learning has been successful in target segmentation (Dimou et al., <xref ref-type="bibr" rid="B2">2016</xref>). But the major bottleneck of deep learning in target segmentation is the need for large-scale labeled datasets for training, particularly in multimodal data (Yao et al., <xref ref-type="bibr" rid="B18">2017</xref>). Owing to the advances of deep learning networks, new insights have been presented in the field of multimodal processing for time-series images. In Shivakumar et al. (<xref ref-type="bibr" rid="B10">2020</xref>), a camera calibration method and a dual-stream CNN architecture was applied to multimodal image segmentation that is able to fuse RGB with thermal information. The quantitative assessments of this study for PST900 dataset show that the mean intersection over union (mIoU) is about 68% for RGB-Thermal mode. A squeeze-and-attention network (Zhong et al., <xref ref-type="bibr" rid="B21">2020</xref>) is proposed for RGB image segmentation based on pixel-group attention, and pixel-wise prediction, which achieves 83.2% mIoU. Moreover, an Edge-Aware Guidance Fusion Network (EGFNet) was introduces in Zhou et al. (<xref ref-type="bibr" rid="B22">2022</xref>) for multimodal scene parsing and segmentation. This study used only RGB and thermal modality for scene understanding. The quantitative results of this work show that the mean accuracies for FuseSeg-161 (Sun et al., <xref ref-type="bibr" rid="B12">2021</xref>), and the proposed method are about 62.1%, and 74.4%, respectively. FuseSeg-161 is a multimodal data fusion with RGB and thermal images to achieve superior performance of semantic segmentation in urban scenes.</p>
<p>Researchers have studied multimodal data from time-series images, with deep learning approaches as the preferred choice. Although some efforts have been devoted to the development of scene understanding with a large training dataset from thermal, and RGB images, little attention has been devoted to multimodal target detection for different specific target of interest with a small training dataset based on depth, thermal, and RGB images. Most of the current few-shot learning methods use single-modal sensory data, which are usually the RGB images produced by visible cameras. However, the target segmentation performance of these networks is prone to be degraded when lighting conditions are not satisfied, such as dim light or darkness. Moreover, we can improve the accuracy by overcoming the segmentation challenges such as dim light or darkness by thermal information. Thermal image is invariant to lighting variations and affords the ability to take advantage of spectral separation between objects. To the best of the authors&#x00027; knowledge, although the related deep learning methods are fairly powerful for image segmentation from multimodal image segmentation with a large training dataset, there is not still outstanding performance for multiple target segmentation with a small training dataset. A robust target segmentation method not only needs a strong model architecture and learning algorithms but also relies on a comprehensive large-scale training set (Zheng, <xref ref-type="bibr" rid="B20">2022</xref>). Generating and annotating such multimodal datasets can be labor-intensive and costly (Bauer et al., <xref ref-type="bibr" rid="B1">2021</xref>). In real-world applications, only a few labeled datasets may be available at model training time. As a solution, few-shot learning aims to build accurate trained models with less training data, contrary to the general experiment of using a large amount of data (Wang et al., <xref ref-type="bibr" rid="B17">2019</xref>; Feyjie et al., <xref ref-type="bibr" rid="B3">2020</xref>).</p>
<p>In this study, we propose a new multimodal few-shot learning for real-time multiple target segmentation from a small labeled multimodal dataset, including RGB, thermal, and depth images for a search and rescue scenario.</p>
</sec>
<sec id="s2">
<title>Proposed method</title>
<p>For multimodal few-shot learning, we define three datasets where each set contains <italic>N</italic> multimodal images, including: a training dataset <italic>S</italic><sub><italic>train</italic></sub> with semantic classes <italic>y</italic><sub><italic>c</italic></sub>, for training step <inline-formula><mml:math id="M1"><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, with <italic>D</italic> &#x02282;&#x0211D;<sup>3</sup> a composite RGB, depth, and thermal image (RGBDT) space, <inline-formula><mml:math id="M2"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x02192;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0211D;</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> an input RGBDT image, and <inline-formula><mml:math id="M3"><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x02192;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> its corresponding binary mask, a support dataset <inline-formula><mml:math id="M4"><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>p</mml:mi><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, and a test dataset <inline-formula><mml:math id="M5"><mml:msub><mml:mrow><mml:mi>S</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>.</p>
<p>The proposed multimodal few-shot learning method aims at training a new squeeze-and attention CNN &#x003C6;(&#x003B5;, &#x003B8;) on the time-series training set to have the capability to extract a new target <italic>tS</italic><sub><italic>train</italic></sub> on the time-series test set based on <italic>t</italic> references from <italic>S</italic><sub><italic>support</italic></sub>. The proposed squeeze-and-attention mechanism <italic>Y</italic><sub><italic>sa</italic></sub> is defined as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M6"><mml:mrow><mml:mtable columnalign='left'><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mo>=</mml:mo></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>;</mml:mo><mml:msub><mml:mi>&#x00398;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003A9;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mo>=</mml:mo></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:msubsup><mml:mrow><mml:msup><mml:mstyle mathsize='140%' displaystyle='true'><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mtext>&#x0200B;</mml:mtext></mml:msup></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>5</mml:mn></mml:msubsup><mml:msubsup><mml:mrow><mml:msup><mml:mstyle mathsize='140%' displaystyle='true'><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mtext>&#x0200B;</mml:mtext></mml:msup></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:msup><mml:mstyle mathsize='140%' displaystyle='true'><mml:mo>&#x02211;</mml:mo></mml:mstyle><mml:mtext>&#x0200B;</mml:mtext></mml:msup></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>W</mml:mi><mml:mo>&#x02212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy='false'>(</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo>&#x000D7;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>j</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msub><mml:mi>&#x00398;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x003A9;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy='false'>)</mml:mo><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr columnalign='left'><mml:mtd columnalign='left'><mml:mrow></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x000A0;</mml:mo><mml:mo>&#x000A0;</mml:mo><mml:mo>&#x000A0;</mml:mo></mml:mrow></mml:mtd><mml:mtd columnalign='left'><mml:mrow><mml:mo>&#x000D7;</mml:mo><mml:mo>&#x000A0;</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:msub><mml:mn>1</mml:mn><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy='false'>(</mml:mo><mml:msubsup><mml:mi>&#x00398;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow></mml:mrow><mml:mo>&#x02032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x003A9;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow></mml:mrow><mml:mo>&#x02032;</mml:mo></mml:msup></mml:mrow></mml:msubsup><mml:mo stretchy='false'>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M7"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mrow><mml:mi>s</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:mo>=</mml:mo></mml:mtd><mml:mtd><mml:mi>&#x02127;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x003C3;</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>&#x02297;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>A</italic><sub><italic>i</italic></sub> is a function to calculate the attention maps given the input feature maps <italic>P</italic>(<italic>x</italic><sub><italic>i</italic></sub>). <italic>P</italic>(<italic>x</italic><sub><italic>i</italic></sub>) is a median pooling layer for input feature map. <italic>f</italic><sub><italic>attn</italic></sub> is an attention function emphasizes the attention of pixel groups that belong to the same classes at different spatial scales (Zhong et al., <xref ref-type="bibr" rid="B21">2020</xref>), which is parameterized by &#x00398;<sub><italic>attn</italic></sub> and &#x003A9;<sub><italic>attn</italic></sub>. &#x00398;<sub><italic>attn</italic></sub> and &#x003A9;<sub><italic>attn</italic></sub> represent the weights and biases from two stacked convolutional layers is added to output map. Moreover, &#x02127;(.) is an upsampling function for expanding the result of the attention channel, &#x003C3; is a relu function, and <italic>x</italic><sub><italic>res</italic></sub> is a residual feature map with an element-wise multiplication &#x02297; (Khoshboresh-Masouleh and Shah-Hosseini, <xref ref-type="bibr" rid="B5">2021</xref>).</p>
<p>The pipeline illustration of the proposed method is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The proposed method has three components, including (a) learns the representation from time-series images pair based on Equation (1) from training dataset with two encoding blocks for support and query images, (b) for getting prior knowledge from a few support samples, a new class registration (CR) network for few support samples with new target class is designed based on dilation convolution layers, and (c) a multiscale decoder network by a transposed convolution with the stride of two for generating the final target detection mask from the test dataset.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Overview of our method for multimodal few-shot learning. <bold>(A)</bold> Proposed architecture for feature extraction and real-time multiple target segmentation, <bold>(B)</bold> proposed squeeze-and-attention mechanism, <bold>(C)</bold> proposed class registration (CR) network, and <bold>(D)</bold> legend.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-1062792-g0001.tif"/>
</fig>
<p>As depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>, the convolutional blocks include 3 &#x000D7; 3 kernels which are applied to the input feature maps using stride one. The proposed architecture is composed of eight squeeze-and-attention layers followed by batch normalization, and rectified linear unit functions to generate feature maps, as well as max-pooling blocks to reduce the size of feature maps. We use multiscale fusion to take advantage of its good ability for multimodal data fusion. In the proposed squeeze-and-attention mechanism (<xref ref-type="fig" rid="F1">Figure 1B</xref>), the multiscale multimodal fusion is performed by three multiscale dilated convolution layers and an element-wise summation of feature layers from the direct path for RGB, thermal, and depth data, resulting in better target boundaries.</p>
</sec>
<sec id="s3">
<title>Dataset</title>
<p>We evaluate our model on the PST900 dataset (Shivakumar et al., <xref ref-type="bibr" rid="B10">2020</xref>) includes 256 images for test set, 128 images for support dataset, and 352 for training set. PST900 is built from synchronized and calibrated RGBDT time-series images with a size of 1,280 &#x000D7; 720 pixels for real-time target segmentation, and contains five target categories. In action, three classes, including Fire-Extinguisher, Backpack, and clutter are used for training, and the remaining two categories, including Hand-Drill, and Survivor for testing. In this study, the PST900 dataset is single-fold with three training classes and two test classes. It consists of 32 time-series sets, where each class contains about 30&#x02013;80 images with their corresponding pixel-level ground truth annotations.</p>
</sec>
<sec id="s4">
<title>Implementation details</title>
<p>We train the proposed network in PyTorch with multi-class cross-entropy over the training class during 120 epochs on PST900 with batch size set to 5, and use Adam as optimizer with the initial learning rate set to 10<sup>&#x02212;4</sup>. &#x003B2;<sub>1</sub>, and &#x003B2;<sub>2</sub> are set to 0.9, and 0.999 and weight decay to 10<sup>&#x02212; 8</sup>.</p>
</sec>
<sec id="s5">
<title>Evaluation protocol</title>
<p>In our work, the evaluation protocol used in most works in few-shot semantic segmentation is employed (Wang H. et al., <xref ref-type="bibr" rid="B15">2020</xref>). The intersection-over-union (IoU) is the standard metric used in evaluating pixel-wise target segmentation. Given two ground truth (<italic>g</italic>) and predicted segment (<italic>p</italic>) masks, the IoU can be defined as <inline-formula><mml:math id="M8"><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x02229;</mml:mo><mml:mi>g</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>p</mml:mi><mml:mo>&#x0222A;</mml:mo><mml:mi>g</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:math></inline-formula>.</p>
</sec>
<sec sec-type="results" id="s6">
<title>Results</title>
<p>To investigate the behavior of the multimodal few-shot learning, we investigate the components of the proposed method PST900 where the training and test classes are required to be simultaneously identified. Although the related models are applied to few-shot learning, there is not still a good method for multimodal few-shot learning for real-time multiple target segmentation. For a fair comparison, all state-of-the-art models were trained from the beginning, using the same training set and feature extractor that was applied for the training of the proposed model. We compare our model against relevant methods HSNet (Min et al., <xref ref-type="bibr" rid="B7">2021</xref>), CANet (Zhang et al., <xref ref-type="bibr" rid="B19">2019</xref>), and PFENet (Tian et al., <xref ref-type="bibr" rid="B13">2022</xref>). Qualitative results of one-shot multiple target detection for HSNet, CANet, PFENet, and our method are shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. We provide some qualitative results on multimodal time-series images that show how our model helps refine the real-time multiple target detection. Note that the proposed method and PFENet can effectively remove irrelevant boundaries and fill the target region. <xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref> show the performance of the proposed method in comparison to the other state-of-the-art methods on the PST900 for Hand-Drill and Survivor classes.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Qualitative results (one-shot) of HSNet, CANet, PFENet, and our method on PST900&#x00027;s time-series images.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-1062792-g0002.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Hand-Drill class mIoU and inference time (in ms) results on PST900 for multimodal few-shot learning.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>Backbone</bold></th>
<th valign="top" align="center"><bold>1 shot</bold></th>
<th valign="top" align="center"><bold>5 shot</bold></th>
<th valign="top" align="center"><bold>10 shot</bold></th>
<th valign="top" align="center"><bold>ms<xref ref-type="table-fn" rid="TN2"><sup>b</sup></xref></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HSNet<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></td>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">48.6</td>
<td valign="top" align="center">53.2</td>
<td valign="top" align="center">59.4</td>
<td valign="top" align="center"><underline>40</underline></td>
</tr>
<tr>
<td valign="top" align="left">CANet<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></td>
<td/>
<td valign="top" align="center">38.5</td>
<td valign="top" align="center">43.3</td>
<td valign="top" align="center">51.6</td>
<td valign="top" align="center">51</td>
</tr>
<tr>
<td valign="top" align="left">PFENet<xref ref-type="table-fn" rid="TN1"><sup>a</sup></xref></td>
<td/>
<td valign="top" align="center"><bold>58.9</bold></td>
<td valign="top" align="center"><underline>62.1</underline></td>
<td valign="top" align="center"><underline>67.3</underline></td>
<td valign="top" align="center">64</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center"><underline>58.1</underline></td>
<td valign="top" align="center"><bold>65.4</bold></td>
<td valign="top" align="center"><bold>78.7</bold></td>
<td valign="top" align="center"><bold>35</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Best results in bold and the underlined font denotes the second-best result.</p>
<fn id="TN1"><label>a</label><p>This model was revised for multimodal training based on the proposed method for multiple target detection.</p></fn>
<fn id="TN2"><label>b</label><p>On an NVIDIA Tesla K80.</p></fn>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Survivor class mIoU and inference time (in ms) results on PST900 for multimodal few-shot learning.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Methods</bold></th>
<th valign="top" align="center"><bold>Backbone</bold></th>
<th valign="top" align="center"><bold>1 shot</bold></th>
<th valign="top" align="center"><bold>5 shot</bold></th>
<th valign="top" align="center"><bold>10 shot</bold></th>
<th valign="top" align="center"><bold>ms<xref ref-type="table-fn" rid="TN4"><sup>b</sup></xref></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">HSNet<xref ref-type="table-fn" rid="TN3"><sup>a</sup></xref></td>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">44.8</td>
<td valign="top" align="center">47.1</td>
<td valign="top" align="center">51.4</td>
<td valign="top" align="center"><underline>42</underline></td>
</tr>
<tr>
<td valign="top" align="left">CANet<xref ref-type="table-fn" rid="TN3"><sup>a</sup></xref></td>
<td/>
<td valign="top" align="center">32.2</td>
<td valign="top" align="center">33.6</td>
<td valign="top" align="center">37.5</td>
<td valign="top" align="center">54</td>
</tr>
<tr>
<td valign="top" align="left">PFENet<xref ref-type="table-fn" rid="TN3"><sup>a</sup></xref></td>
<td/>
<td valign="top" align="center"><underline>59.4</underline></td>
<td valign="top" align="center"><underline>52.3</underline></td>
<td valign="top" align="center"><underline>55.7</underline></td>
<td valign="top" align="center">65</td>
</tr>
<tr>
<td valign="top" align="left">Ours</td>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center"><bold>60.3</bold></td>
<td valign="top" align="center"><bold>61.1</bold></td>
<td valign="top" align="center"><bold>68.4</bold></td>
<td valign="top" align="center"><bold>38</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Best results in bold and the underlined font denotes the second-best result.</p>
<fn id="TN3"><label>a</label><p>This model was revised for multimodal training based on the proposed method for multiple target detection.</p></fn>
<fn id="TN4"><label>b</label><p>On an NVIDIA Tesla K80.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>According to the results, the function of PFENet was found to be better than HSNet and CANet. But, many of the target pixels were not detected in both 5-shot and 10-shot scenarios. We observe that PFENet outperforms the proposed method in the 1-shot scenario for Hand-Drill class. The proposed method achieves a mIoU for the Hand-Drill class of 58.1, 65.4, and 78.7, while the mIoUs for the Survivor class are 60.3, 61.1, and 68.4 in 1-shot, 5-shot, and 10-shot scenarios.</p>
</sec>
<sec id="s7">
<title>Ablation study</title>
<p>In this section, we present an ablation study to compare a number of different model variants, such as different modalities (e.g., RGB, thermal, and depth), loss functions [Bootstrapped Cross-entropy (BC), Dice, and multi-class cross-entropy], and backbones (e.g., ResNet50, and HRNet), and justify our design choices. <xref ref-type="table" rid="T3">Table 3</xref> shows some ablation study results to investigate the behavior of the proposed model.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Ablation study on using different modalities, loss functions, and backbones.</p></caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left"><bold>Feature extractor</bold></th>
<th valign="top" align="center"><bold>Backbone</bold></th>
<th valign="top" align="center"><bold>Loss</bold></th>
<th valign="top" align="center"><bold>Modalities</bold></th>
<th valign="top" align="center"><bold>mIoU<xref ref-type="table-fn" rid="TN5"><sup>a</sup></xref></bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Proposed layer <bold>Y</bold><sub><bold>sa</bold></sub> (10 shot)</td>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">D</td>
<td valign="top" align="center">RGB</td>
<td valign="top" align="center">49.83</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">49.01</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">BC</td>
<td/>
<td valign="top" align="center">50.46</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">48.64</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">M2CE</td>
<td/>
<td valign="top" align="center"><bold>52.34</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">49.72</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">D</td>
<td valign="top" align="center">RGB&#x0002B;Thermal</td>
<td valign="top" align="center">59.63</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">57.24</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">BC</td>
<td/>
<td valign="top" align="center">59.43</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">58.90</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">M2CE</td>
<td/>
<td valign="top" align="center"><bold>60.04</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">59.74</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">D</td>
<td valign="top" align="center">RGB&#x0002B;Depth&#x0002B; Thermal</td>
<td valign="top" align="center">69.17</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">67.12</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">BC</td>
<td/>
<td valign="top" align="center">70.20</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">68.34</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">ResNet50</td>
<td valign="top" align="center">M2CE</td>
<td/>
<td valign="top" align="center"><bold>73.55</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">HRNet</td>
<td/>
<td/>
<td valign="top" align="center">68.46</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Best results in bold.</p>
<fn id="TN5"><label>a</label><p>Mean intersection over union for Hand-Drill, and Survivor classes.</p></fn>
</table-wrap-foot>
</table-wrap>
<p>Visualization of ablation study results of the example set for the highest accuracy for each modality among the compared methods is shown in <xref ref-type="fig" rid="F3">Figure 3</xref>. For each test result with RGB, RGB-Thermal, and RGB&#x0002B;Depth&#x0002B;Thermal, the proposed method is effective in target segmentation with RGB&#x0002B;Depth&#x0002B;Thermal modality. Moreover, the proposed method obtained better target segmentation results than other models with RGB or RGB-Thermal modalities.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The visualization of real-time target segmentation output through each modality with ResNet50 backbone and M2CE function.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcomp-04-1062792-g0003.tif"/>
</fig>
<sec>
<title>Modalities</title>
<p>The proposed model for real-time multiple target segmentation was tested with different data modalities, such as RGB, thermal, and depth images.</p>
</sec>
<sec>
<title>Loss functions</title>
<p>All networks were additionally evaluated with different loss functions. Although the proposed model with multi-class cross-entropy (M2CE) function delivers good results, the proposed model was tested with another two loss functions, consist of Dice (D) (Sudre et al., <xref ref-type="bibr" rid="B11">2017</xref>) and Bootstrapped Cross-entropy (BC) (Gaj et al., <xref ref-type="bibr" rid="B4">2021</xref>). In this regard, we train our model using the absolute error between the ground truth map and the model&#x00027;s predicted.</p>
</sec>
<sec>
<title>Backbones</title>
<p>The proposed model can be set up with different backbones for real-time multiple target segmentation. We selected two backbones for ablation study. The experiments were carried out with the ResNet50, and HRNet (Wang J. et al., <xref ref-type="bibr" rid="B16">2020</xref>) backbones.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s8">
<title>Conclusion</title>
<p>We have presented a new multimodal few-shot learning for real-time multiple target detection in a real-world application. Different from the previous few-shot learning methods, the proposed method aims at accurate and fast identifying new targets from time-series images. Our method has outstanding results on a challenging dataset PST900 with 32 time-series sets, compared with recent prominent models in few-shot learning.</p>
</sec>
<sec sec-type="data-availability" id="s9">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="s10">
<title>Author contributions</title>
<p>MK-M and RS-H conceived the original idea for this manuscript and discussed the results. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s11">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bauer</surname> <given-names>D. F.</given-names></name> <name><surname>Russ</surname> <given-names>T.</given-names></name> <name><surname>Waldkirch</surname> <given-names>B. I.</given-names></name> <name><surname>T&#x000F6;nnes</surname> <given-names>C.</given-names></name> <name><surname>Segars</surname> <given-names>W. P.</given-names></name> <name><surname>Schad</surname> <given-names>L. R.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Generation of annotated multimodal ground truth datasets for abdominal medical image registration</article-title>. <source>Int. J. Comput. Assist. Radiol. Surg.</source> <volume>16</volume>, <fpage>1277</fpage>&#x02013;<lpage>1285</lpage>. <pub-id pub-id-type="doi">10.1007/s11548-021-02372-7</pub-id><pub-id pub-id-type="pmid">33934313</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Dimou</surname> <given-names>A.</given-names></name> <name><surname>Medentzidou</surname> <given-names>P.</given-names></name> <name><surname>Garc&#x000ED;a</surname> <given-names>F. &#x000C1;.</given-names></name> <name><surname>Daras</surname> <given-names>P.</given-names></name></person-group> (<year>2016</year>). <article-title>Multi-target detection in CCTV footage for tracking applications using deep learning techniques,</article-title> in <source>2016 IEEE International Conference on Image Processing (ICIP)</source> (<publisher-loc>Phoenix, AZ</publisher-loc>), <fpage>928</fpage>&#x02013;<lpage>932</lpage>. <pub-id pub-id-type="doi">10.1109/ICIP.2016.7532493</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Feyjie</surname> <given-names>A. R.</given-names></name> <name><surname>Azad</surname> <given-names>R.</given-names></name> <name><surname>Pedersoli</surname> <given-names>M.</given-names></name> <name><surname>Kauffman</surname> <given-names>C.</given-names></name> <name><surname>Ayed</surname> <given-names>I. B.</given-names></name> <name><surname>Dolz</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Semi-supervised few-shot learning for medical image segmentation</article-title>. <source>ArXiv200308462 Cs</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/2003.08462">http://arxiv.org/abs/2003.08462</ext-link> (accessed March 24, 2022).</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gaj</surname> <given-names>S.</given-names></name> <name><surname>Ontaneda</surname> <given-names>D.</given-names></name> <name><surname>Nakamura</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>Automatic segmentation of gadolinium-enhancing lesions in multiple sclerosis using deep learning from clinical MRI</article-title>. <source>PLoS ONE</source> <volume>16</volume>:<fpage>e0255939</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0255939</pub-id><pub-id pub-id-type="pmid">34469432</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khoshboresh-Masouleh</surname> <given-names>M.</given-names></name> <name><surname>Shah-Hosseini</surname> <given-names>R.</given-names></name></person-group> (<year>2021</year>). <article-title>Building panoptic change segmentation with the use of uncertainty estimation in squeeze-and-attention CNN and remote sensing observations</article-title>. <source>Int. J. Remote Sens.</source> <volume>42</volume>, <fpage>7798</fpage>&#x02013;<lpage>7820</lpage>. <pub-id pub-id-type="doi">10.1080/01431161.2021.1966853</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mart&#x000ED;-Bonmat&#x000ED;</surname> <given-names>L.</given-names></name> <name><surname>Sopena</surname> <given-names>R.</given-names></name> <name><surname>Bartumeus</surname> <given-names>P.</given-names></name> <name><surname>Sopena</surname> <given-names>P.</given-names></name></person-group> (<year>2010</year>). <article-title>Multimodality imaging techniques</article-title>. <source>Contrast Media Mol. Imaging</source> <volume>5</volume>, <fpage>180</fpage>&#x02013;<lpage>189</lpage>. <pub-id pub-id-type="doi">10.1002/cmmi.393</pub-id><pub-id pub-id-type="pmid">20812286</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Min</surname> <given-names>J.</given-names></name> <name><surname>Kang</surname> <given-names>D.</given-names></name> <name><surname>Cho</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Hypercorrelation squeeze for few-shot segmentation,</article-title> in <source>2021 IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Montreal, QC</publisher-loc>), <fpage>6941</fpage>&#x02013;<lpage>6952</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content/ICCV2021/html/Min_Hypercorrelation_Squeeze_for_Few-Shot_Segmentation_ICCV_2021_paper.html">https://openaccess.thecvf.com/content/ICCV2021/html/Min_Hypercorrelation_Squeeze_for_Few-Shot_Segmentation_ICCV_2021_paper.html</ext-link> (accessed March 24, 2022).</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morelande</surname> <given-names>M. R.</given-names></name> <name><surname>Kreucher</surname> <given-names>C. M.</given-names></name> <name><surname>Kastella</surname> <given-names>K.</given-names></name></person-group> (<year>2007</year>). <article-title>A Bayesian approach to multiple target detection and tracking</article-title>. <source>IEEE Trans. Signal Process.</source> <volume>55</volume>, <fpage>1589</fpage>&#x02013;<lpage>1604</lpage>. <pub-id pub-id-type="doi">10.1109/TSP.2006.889470</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Rahman</surname> <given-names>M. M.</given-names></name> <name><surname>Rahman</surname> <given-names>T.</given-names></name> <name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Alam</surname> <given-names>M. A. U.</given-names></name></person-group> (<year>2021</year>). <article-title>Knowledge transfer across imaging modalities via simultaneous learning of adaptive autoencoders for high-fidelity mobile robot vision,</article-title> in <source>2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source> (<publisher-loc>Prague</publisher-loc>), <fpage>1267</fpage>&#x02013;<lpage>1273</lpage>. <pub-id pub-id-type="doi">10.1109/IROS51168.2021.9636360</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Shivakumar</surname> <given-names>S. S.</given-names></name> <name><surname>Rodrigues</surname> <given-names>N.</given-names></name> <name><surname>Zhou</surname> <given-names>A.</given-names></name> <name><surname>Miller</surname> <given-names>I. D.</given-names></name> <name><surname>Kumar</surname> <given-names>V.</given-names></name> <name><surname>Taylor</surname> <given-names>C. J.</given-names></name></person-group> (<year>2020</year>). <article-title>PST900: RGB-thermal calibration, dataset and segmentation network,</article-title> in <source>2020 IEEE International Conference on Robotics and Automation (ICRA)</source> (<publisher-loc>Paris</publisher-loc>), <fpage>9441</fpage>&#x02013;<lpage>9447</lpage>. <pub-id pub-id-type="doi">10.1109/ICRA40945.2020.9196831</pub-id><pub-id pub-id-type="pmid">27295638</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sudre</surname> <given-names>C. H.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Vercauteren</surname> <given-names>T.</given-names></name> <name><surname>Ourselin</surname> <given-names>S.</given-names></name> <name><surname>Cardoso</surname> <given-names>M. J.</given-names></name></person-group> (<year>2017</year>). <article-title>Generalised dice overlap as a deep learning loss function for highly unbalanced segmentations</article-title>. <source>ArXiv:170703237</source> Cs <volume>10553</volume>, <fpage>240</fpage>&#x02013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-67558-9_28</pub-id><pub-id pub-id-type="pmid">34104926</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name> <name><surname>Yun</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>FuseSeg: semantic segmentation of urban scenes based on RGB and thermal data fusion</article-title>. <source>IEEE Trans. Autom. Sci. Eng.</source> <volume>18</volume>, <fpage>1000</fpage>&#x02013;<lpage>1011</lpage>. <pub-id pub-id-type="doi">10.1109/TASE.2020.2993143</pub-id></citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>Z.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Shu</surname> <given-names>M.</given-names></name> <name><surname>Yang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>R.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2022</year>). <article-title>Prior guided feature enrichment network for few-shot segmentation</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>44</volume>, <fpage>1050</fpage>&#x02013;<lpage>1065</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2020.3013717</pub-id><pub-id pub-id-type="pmid">32750843</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wagner</surname> <given-names>D.</given-names></name> <name><surname>Schmalstieg</surname> <given-names>D.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name></person-group> (<year>2009</year>). <article-title>Multiple target detection and tracking with guaranteed framerates on mobile phones,</article-title> in <source>2009 8th IEEE International Symposium on Mixed and Augmented Reality</source> (<publisher-loc>Washington, DC</publisher-loc>), <fpage>57</fpage>&#x02013;<lpage>64</lpage>. <pub-id pub-id-type="doi">10.1109/ISMAR.2009.5336497</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>X.</given-names></name> <name><surname>Zhen</surname> <given-names>X.</given-names></name></person-group> (<year>2020</year>). <article-title>Few-shot semantic segmentation with democratic attention networks,</article-title> in <source>Computer Vision &#x02013; ECCV 2020 Lecture Notes in Computer Science</source>, eds <person-group person-group-type="editor"><name><surname>Vedaldi</surname> <given-names>A.</given-names></name> <name><surname>Bischof</surname> <given-names>H.</given-names></name> <name><surname>Brox</surname> <given-names>T.</given-names></name> <name><surname>Frahm</surname> <given-names>J.-M.</given-names></name></person-group> (<publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>), <fpage>730</fpage>&#x02013;<lpage>746</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-58601-0_43</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Sun</surname> <given-names>K.</given-names></name> <name><surname>Cheng</surname> <given-names>T.</given-names></name> <name><surname>Jiang</surname> <given-names>B.</given-names></name> <name><surname>Deng</surname> <given-names>C.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Deep high-resolution representation learning for visual recognition</article-title>. <source>ArXiv190807919 Cs</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://arxiv.org/abs/1908.07919">http://arxiv.org/abs/1908.07919</ext-link> (accessed May 3, 2022).<pub-id pub-id-type="pmid">32248092</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Liew</surname> <given-names>J. H.</given-names></name> <name><surname>Zou</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>D.</given-names></name> <name><surname>Feng</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>PANet: few-shot image semantic segmentation with prototype alignment,</article-title> in <source>2019 IEEE/CVF International Conference on Computer Vision (ICCV)</source> (<publisher-loc>Seoul</publisher-loc>), <fpage>9197</fpage>&#x02013;<lpage>9206</lpage>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content_ICCV_2019/html/Wang_PANet_Few-Shot_Image_Semantic_Segmentation_With_Prototype_Alignment_ICCV_2019_paper.html">https://openaccess.thecvf.com/content_ICCV_2019/html/Wang_PANet_Few-Shot_Image_Semantic_Segmentation_With_Prototype_Alignment_ICCV_2019_paper.html</ext-link> [Accessed March 24, 2022].</citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Yao</surname> <given-names>H.</given-names></name> <name><surname>Yu</surname> <given-names>Q.</given-names></name> <name><surname>Xing</surname> <given-names>X.</given-names></name> <name><surname>He</surname> <given-names>F.</given-names></name> <name><surname>Ma</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Deep-learning-based moving target detection for unmanned air vehicles,</article-title> in <source>2017 36th Chinese Control Conference (CCC)</source> (<publisher-loc>Dalian</publisher-loc>), <fpage>11459</fpage>&#x02013;<lpage>11463</lpage>. <pub-id pub-id-type="doi">10.23919/ChiCC.2017.8029186</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.</given-names></name> <name><surname>Lin</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Yao</surname> <given-names>R.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name></person-group> (<year>2019</year>). <article-title>CANet: class-agnostic segmentation networks with iterative refinement and attentive few-shot learning,</article-title> in <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Long Beach, CA</publisher-loc>), <fpage>5217</fpage>&#x02013;<lpage>5226</lpage>. Available at: <ext-link ext-link-type="uri" xlink:href="https://openaccess.thecvf.com/content_CVPR_2019/html/Zhang_CANet_Class-Agnostic_Segmentation_Networks_With_Iterative_Refinement_and_Attentive_Few-Shot_CVPR_2019_paper.html">https://openaccess.thecvf.com/content_CVPR_2019/html/Zhang_CANet_Class-Agnostic_Segmentation_Networks_With_Iterative_Refinement_and_Attentive_Few-Shot_CVPR_2019_paper.html</ext-link> (accessed March 24, 2022).</citation>
</ref>
<ref id="B20">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>L.</given-names></name></person-group> (<year>2022</year>). <source>The 1st Workshop on Vision Datasets Understanding - CVPR 2022</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://sites.google.com/view/vdu-cvpr22">https://sites.google.com/view/vdu-cvpr22</ext-link> (accessed March 24, 2022).</citation>
</ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhong</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>Z. Q.</given-names></name> <name><surname>Bidart</surname> <given-names>R.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Daya</surname> <given-names>I. B.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Squeeze-and-attention networks for semantic segmentation,</article-title> in <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source> (<publisher-loc>Seattle, WA</publisher-loc>), <fpage>13062</fpage>&#x02013;<lpage>13071</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01308</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>W.</given-names></name> <name><surname>Dong</surname> <given-names>S.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Qian</surname> <given-names>Y.</given-names></name></person-group> (<year>2022</year>). <article-title>Edge-aware guidance fusion network for RGB&#x02013;thermal scene parsing</article-title>. <source>Proc. AAAI Conf. Artif. Intell.</source> <volume>36</volume>, <fpage>3571</fpage>&#x02013;<lpage>3579</lpage>. <pub-id pub-id-type="doi">10.1609/aaai.v36i3.20269</pub-id></citation>
</ref>
</ref-list> 
</back>
</article>
