<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Phys.</journal-id>
<journal-title>Frontiers in Physics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Phys.</abbrev-journal-title>
<issn pub-type="epub">2296-424X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1206944</article-id>
<article-id pub-id-type="doi">10.3389/fphy.2023.1206944</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Physics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A boundary enhancement and pixel alignment based smoke segmentation network</article-title>
<alt-title alt-title-type="left-running-head">Zhou et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fphy.2023.1206944">10.3389/fphy.2023.1206944</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhou</surname>
<given-names>Fangrong</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wen</surname>
<given-names>Gang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Pan</surname>
<given-names>Hao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Yifan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wang</surname>
<given-names>Guiqian</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2275145/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yuan</surname>
<given-names>Feiniu</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1458019/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Joint Laboratory of Power Remote Sensing Technology</institution>, <institution>Electric Power Research Institute of Yunnan Electric Power Company</institution>, <addr-line>Kunming</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Mathematics and Science College</institution>, <institution>Shanghai Normal University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>College of Information, Mechanical and Electrical Engineering</institution>, <institution>Shanghai Normal University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Key Innovation Group of Digital Humanities Resource and Research</institution>, <institution>Shanghai Normal University</institution>, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1504166/overview">Anna Cimmino</ext-link>, ELI Beamlines, Czechia</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/661492/overview">Jakub Nalepa</ext-link>, Silesian University of Technology, Poland</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1664211/overview">Guanqiu Qi</ext-link>, Buffalo State College, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Guiqian Wang, <email>guiqianw@163.com</email>
</corresp>
</author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>11</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>11</volume>
<elocation-id>1206944</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Zhou, Wen, Pan, Wang, Wang and Yuan.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Zhou, Wen, Pan, Wang, Wang and Yuan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Image segmentation methods usually fuse shallow and deep features to locate object boundaries, but it is difficult to improve the accuracy of smoke segmentation by conventional fusion methods. It is a very difficult vision task to perform semantic segmentation of smoke images, because the translucency and irregular shapes of smoke lead to extremely complicated mixtures with background that are adverse to segmentation. To improve the segmentation accuracy of smoke scenes, we propose a Boundary Enhancement and Pixel Alignment based smoke segmentation Network for fire alarms. For the shallow features of the network, an attention mechanism is adopted to capture spatially details of smoke for improving boundary precision. For the deep layers, the Pyramid Pooling Module is used to extract local features and abstract semantic ones simultaneously. Finally, to efficiently merge shallow and deep features, a Pixel Alignment Module is adopted to model the relationship between pixel locations. The experimental results show that the mean Intersection over Union of the proposed method on the three synthetic smoke test datasets is 78.61%, 77.63% and 77.30%, respectively, and it outperforms most of the existing methods. In addition, our method obtains satisfying results on inconspicuous smoke and smoke-like images.</p>
</abstract>
<kwd-group>
<kwd>smoke segmentation</kwd>
<kwd>deep neural network</kwd>
<kwd>boundary enhancement module</kwd>
<kwd>pixel alignment module</kwd>
<kwd>pyramid pooling module</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Radiation Detectors and Imaging</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>The frequent occurrence of fires not only causes significant economic losses for society, but more importantly, it will jeopardize social public security and have extremely bad effects on the natural and ecological environment. It is too late to detect the naked flame, so the recognition of smoke early in the fire can control this disaster to a certain extent.</p>
<p>To solve fire detection in open or large spaces, a number of deep models [<xref ref-type="bibr" rid="B1">1</xref>&#x2013;<xref ref-type="bibr" rid="B6">6</xref>] have been proposed for fire detection. These deep fire detection methods differ from traditional ones. Deep learning models segment smoke areas at a fine granularity for separating smoke targets from cluttered backgrounds, which are more accurate than those obtained by traditional methods. According to segmented maps, staffs can analyze safe areas and predict fire trends for reducing damages. Detecting smoke from fires at the pixel level is important, but the task is very challenging due to the variability introduced by the visual characteristics of smoke.</p>
<p>Smoke is visually semi-transparent, so it becomes more challenging to separate pixels in smoke edge regions. The problem that needs to be focused on is how to collect semantic information for categorization and localization. Some methods try to increase the depth of network for capturing more semantic features, but this technique also loses local details. The solutions to the above-mentioned conflicts can be divided into three main categories. The first category is to adopt a skip-connection structure [<xref ref-type="bibr" rid="B7">7</xref>] between deep and shallow levels. U-Net [<xref ref-type="bibr" rid="B7">7</xref>] adopts gradual upsampling to avoid the loss of features. The second one is to use the atrous convolution [<xref ref-type="bibr" rid="B8">8</xref>] to capture information at large scales. The third one is to use the multi-scale fusion [<xref ref-type="bibr" rid="B9">9</xref>] approach to integrate information from different scales. Spatial attention [<xref ref-type="bibr" rid="B10">10</xref>] is able to capture the most informative region of feature maps by globally modeling the relevance of all pixels, thus it effectively solves the misclassification problem of isolated regions that are far away from the main smoke area and strengthens the boundaries of smoke.</p>
<p>Based on the analysis of existing methods, we propose a Boundary Enhancement and Pixel Alignment based smoke segmentation Network (BEPA-Net). For the shallow features, we design a Boundary Enhancement Module (BEM) to model the long-range ability of attention mechanism for obtaining clear target boundaries. For the deep features, we adopt the Pyramid Pooling Module (PPM) [<xref ref-type="bibr" rid="B11">11</xref>] to extract global and local contextual information for enhancing semantics. By fusing different resolution feature maps, we propose a Pixel Alignment Module (PAM), which can construct the pixel correspondence relationship between feature maps for better information fusion.</p>
<p>This paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes related work on image semantic segmentation and smoke segmentation. In <xref ref-type="sec" rid="s3">Section 3</xref>, we describe the main idea of this paper. <xref ref-type="sec" rid="s4">Section 4</xref> presents experiments and analysis. At last, we conclude this paper in <xref ref-type="sec" rid="s5">Section 5</xref>.</p>
</sec>
<sec id="s2">
<title>2 Related works</title>
<sec id="s2-1">
<title>2.1 Semantic segmentation</title>
<p>Semantic segmentation is an image classification task at the pixel level [<xref ref-type="bibr" rid="B12">12</xref>]. proposed a Full Convolution Network (FCN), which is a basic paradigm of semantic segmentation methods. FCN pioneers an end-to-end approach to achieve pixel-by-pixel classifications. In the encoder of FCN, a series of convolutional layers and successive down-sampling ones are often used to extract deep features with large receptive fields, and then the decoder upsamples the extracted deep features to the same resolution as the input. To lessen the loss of spatial information caused by down-sampling, skip connections are used to fuse low-level features with high-level ones from the different scales of encoders and decoders. Two fundamental paradigms have emerged as a result of further researches, including symmetric codec architectures [<xref ref-type="bibr" rid="B7">7</xref>, <xref ref-type="bibr" rid="B13">13</xref>] and asymmetric codec ones [<xref ref-type="bibr" rid="B8">8</xref>, <xref ref-type="bibr" rid="B14">14</xref>&#x2013;<xref ref-type="bibr" rid="B17">17</xref>]. These symmetric codec structures mainly focus on minimizing the information loss caused by frequent down-sampling for expanding receptive fields. Meanwhile, asymmetric codec structures revolve the problem of extracting the most abstract semantics without reducing the feature map resolutions too much. Thus, a balance of spatial and semantic information can be achieved.</p>
</sec>
<sec id="s2-2">
<title>2.2 Smoke segmentation</title>
<p>Traditional smoke segmentation methods make an effort to separate smoke targets from images by extracting the color features of images in different color spaces [<xref ref-type="bibr" rid="B18">18</xref>]. Combined color features and shape features to present a fast smoke detection method for video surveillance [<xref ref-type="bibr" rid="B19">19</xref>]. Used background removal methods and color features to filter non-smoke pixels [<xref ref-type="bibr" rid="B20">20</xref>]. Used the rough set theory to extract candidate smoke regions for video fire detection.</p>
<p>With the rapid development of deep learning in recent years, there are many deep neural networks that have also been proposed for smoke semantic segmentation. Without the need for designing complex hand-crafted features, deep learning based semantic segmentation approaches combine feature extraction and classification for implementing an end-to-end manner [<xref ref-type="bibr" rid="B21">21</xref>]. directly used the AlexNet network [<xref ref-type="bibr" rid="B22">22</xref>] for smoke recognition [<xref ref-type="bibr" rid="B23">23</xref>]. proposed a 3D parallel FCN model to segment smoke regions from videos [<xref ref-type="bibr" rid="B24">24</xref>]. proposed a coding and decoding network by designing a dual-path structure to obtain detailed information and semantic information for smoke segmentation [<xref ref-type="bibr" rid="B25">25</xref>]. proposed a Wave-shaped deep neural Network (W-Net) for smoke density estimation, which is factually a regression over each pixel [<xref ref-type="bibr" rid="B26">26</xref>]. proposed a global smoke attention network that makes a full use of the modeling capabilities of attention mechanism.</p>
<p>Traditional methods rely on manual features, while recent networks only focus on the performance of CNN itself and cannot perform smoke segmentation well. We focus more on the correlation between smoke boundaries and pixels, and complete smoke segmentation by improving the segmentation ability of boundaries and strengthening global modeling capabilities.</p>
</sec>
</sec>
<sec id="s3">
<title>3 The proposed method</title>
<sec id="s3-1">
<title>3.1 The network architecture</title>
<p>We choose the ResNet50 network [<xref ref-type="bibr" rid="B27">27</xref>] as our backbone network because it can obtain rich information. As shown in <xref ref-type="fig" rid="F1">Figure 1</xref>, the ResNet50 [<xref ref-type="bibr" rid="B27">27</xref>] is used as the backbone of our method to extract features, and the backbone network is divided into four stages. Atrous convolutions [<xref ref-type="bibr" rid="B30">30</xref>] increase the receptive field and extract the abundant features of images without significantly reducing spatial resolutions, so we adopt atrous convolutions to compensate for the reduction of spatial resolutions due to down-sampling. The outputs of Stages 2, 3, and 4 greatly improve the semantic representation of deep features after feature fusion and multi-scale context extraction. The shallow features from Stage 1 are first delivered to the proposed Boundary Enhancement Module (BEM) for pixel alignment between different feature maps of different layers. Then the proposed Pixel Alignment Module (PAM) accepts the warping information of the BEM module for information fusion. Thus, we implement pixel alignment for different features. Finally, the fused structure map is upsampled to the original map size for generating the final segmentation map.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Deep and shallow features based Smoke Segmentation Network. BEM, PAM and PPM denotes Boundary Enhancement Module, Pixel Alignment Module and Pyramid Pooling Module, respectively. Conv 1 &#xd7; 1 is a convolutional kernel with the kernel size equal to 1.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g001.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>3.2 Pyramid pooling module</title>
<p>In deep neural networks, the size of receptive fields can roughly represent the degree of using contextual information. High-level contextual information can be captured by explicitly fusing features of objects at different scales, and it can effectively solve the problem of pixel inconsistency within the objects in the segmentation task and reinforce deep semantics.</p>
<p>To obtain more contextual information, we use the Pyramid Pooling Module (PPM) [<xref ref-type="bibr" rid="B11">11</xref>] to obtain the context information of objects, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The PPM module contains global and local information using the adaptive global pooling with the different output sizes of 1, 2, 3, and 6. Then, the pooled feature maps are filtered using a 1 &#xd7; 1 convolution. Next, bilinear interpolation is used to up-sample these filtered feature maps to the same size as the original input of PPM. The resized feature maps are concatenated together. Three layers of convolution (Conv), batch normalization (BN), and activation (ReLU) are used to obtain a feature map with a large amount of contextual information.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Pyramid Pooling Module (PPM). AdaptiveAvgPool is average pooling that uses different pooling kernel. Conv-BN-ReLu 3 &#xd7; 3 denotes a convolution kernel with batch normalization and ReLU activation function.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g002.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>3.3 Boundary enhancement module</title>
<p>The embedding of spatial contexts can emphasize the most informative parts and enable the network to selectively focus on more important features. The proposal of attention mechanism offers a new direction for extracting powerful spatial and contextual information. Most spatial attention methods adopt matrix operations to capture the relationship between any two pixels in the global scope. To obtain spatial attention maps and improve object boundary localization accuracy, we propose a Boundary Enhancement Module (BEM). As shown in <xref ref-type="fig" rid="F3">Figure 3</xref>, the input feature map of BEM is fed into the three branches for extracting features. <italic>H</italic>, <italic>W</italic> and <italic>C</italic> denote the height, width and channel numbers of the feature map, respectively. Average pooling captures the low frequency components of features, maximum pooling is able to extract the high frequency signals, and a 1 &#xd7; 1 convolution learns the features about objects. As a result, we obtain three two-dimensional attentional feature maps, which are fused by point-wise summation. The fused feature map is activated by the <italic>sigmoid</italic> function for generating a spatial attention map. Finally, the input to our BEM is weighted with the spatial attention feature map to produce the feature map of the adaptive refinement boundary. Conv1x1 is a convolutional kernel with a kernel of 1.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Boundary enhancement module (BEM).</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g003.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Pixel alignment module</title>
<p>To aggregate multi-scale features, feature fusion is often performed in different levels of feature maps, but pixel positions corresponding to their objects are different. To solve this problem, we propose a Pixel Alignment Module (PAM). This module first calculates the pixel offsets for pixel alignment, and the feature maps are pixel-aligned for effective fusion. <xref ref-type="fig" rid="F4">Figure 4</xref> shows the proposed PAM, which accomplishes pixel alignment by obtaining the pixel offset field. The value of each pixel in the offset field can be viewed as a moving distance for the pixel. In other words, the offset field can also be called a motion field. The relationship between a feature map and another one obtained by convolutions can be reconstructed by the pixel motion field. Thus, we actually obtain translation invariance by convolutions.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Pixel Alignment Module (PAM). Concat means concatenation along channel, DW-Conv is a deeply separable convolutional kernel with a size of 1, BN is a batch normalization operation, ReLU is an activation function, and &#x201c;3 &#xd7; 3, s &#x3d; <italic>k</italic>&#x201d; represents a convolution with a step size of <italic>k</italic> and a kernel size of 3 &#xd7; 3.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g004.tif"/>
</fig>
<p>A bilinear interpolation layer is used to up-sample the feature map <italic>F</italic>
<sub>1</sub> to produce a feature map <italic>F</italic>
<sub>7</sub>, which has the same size as the feature map <italic>F</italic>
<sub>2</sub>, and feature maps <italic>F</italic>
<sub>7</sub> and <italic>F</italic>
<sub>2</sub> are concatenated to generate a feature map <italic>F</italic>
<sub>3</sub> for channel fusion. A set of depth-wise separable convolutions (DW-Conv) are used to establish the positional relationships between pixels on different feature maps. The pixel motion field <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is then generated using a 3 &#xd7; 3 convolution in a similar way to DCN [<xref ref-type="bibr" rid="B28">28</xref>]. The field map <italic>F</italic>
<sub>4</sub> contains the spatial translation offsets along <italic>x</italic> and <italic>y</italic>-axes, and the feature values at each pixel position <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> on <italic>F</italic>
<sub>4</sub> are used to move the pixel of <italic>F</italic>
<sub>1</sub> to a new position, resulting in a warped feature map <inline-formula id="inf3">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, formulated as:<disp-formula id="e1">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>&#x3c1;</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:munder>
</mml:mstyle>
<mml:msub>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>&#x3c1;</mml:mi>
</mml:msub>
</mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>&#x3c1;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the weight of the bilinear kernel on the curved space grid, which is calculated by <italic>F</italic>
<sub>4</sub>, and <inline-formula id="inf5">
<mml:math id="m6">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes the adjacent position of pixel position <inline-formula id="inf6">
<mml:math id="m7">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c1;</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The warped feature map <italic>F</italic>
<sub>5</sub> is generated by <italic>F</italic>
<sub>1</sub> and <italic>F</italic>
<sub>4</sub>, and <italic>F</italic>
<sub>4</sub> is produced from <italic>F</italic>
<sub>1</sub> and <italic>F</italic>
<sub>2</sub>, so <italic>F</italic>
<sub>5</sub> certainly has a strong relationship with <italic>F</italic>
<sub>2</sub>. Therefore, we concatenate <italic>F</italic>
<sub>5</sub> and <italic>F</italic>
<sub>2</sub> along the channel axis for further feature enhancement. Finally, the concatenated feature map is processed by a 3 &#xd7; 3 convolution without BN and ReLU layers for feature fusion and dimensionality control to obtain the final output <italic>F</italic>
<sub>6</sub>.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experiments and analysis</title>
<sec id="s4-1">
<title>4.1 Datasets and implementation details</title>
<p>In this paper, the datasets used for training and testing are the same as those in [<xref ref-type="bibr" rid="B25">25</xref>], including a virtual synthetic smoke training set with 70,632 images and three virtual synthetic smoke test sets. The three test sets are respectively named DS01, DS02 and DS03, and each set has 1,000 images. The synthetic dataset was created from 8,162 pure smoke images [<xref ref-type="bibr" rid="B25">25</xref>]. adopted computer graphics to generate these pure smoke images with a variety of transparency, texture and fluid properties. Each pure smoke sample is an RGBA image with a spatial resolution of 256 &#xd7; 256, containing RGB color channels <bold>
<italic>S</italic>
</bold> and a opacity channel <italic>&#x3b1;</italic>, respectively. The opacity <italic>&#x3b1;</italic> is limited in the range [0, 1]. According to the rules of linear composition, the combination of a pure smoke image <bold>
<italic>S</italic>
</bold> and a background <bold>
<italic>B</italic>
</bold> generates an observation image <bold>
<italic>I</italic>
</bold>. This procedure can be mathematically defined as:<disp-formula id="e2">
<mml:math id="m8">
<mml:mrow>
<mml:mi mathvariant="normal">I</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="normal">&#x3b1;</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3b1;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x22c5;</mml:mo>
<mml:mi>B</mml:mi>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>Using the above method, we are able to construct a large number of training sets without the tedious labeling. Each virtual smoke image is randomly and linearly combined with a real background image for generating a virtual smoke training dataset. The generated virtual dataset is diverse in terms of colors, sizes, textures of smoke and background for simulating most real smoke scenes.</p>
<p>All the experiments were performed on a windows 10&#xa0;PC with an NVIDIA RTX3090 GPU, and the programming environment is the python 3.7 and pytorch 1.7 framework. The Stochastic Gradient Descent (SGD) optimizer was used for training. The learning rate is set to 0.0001, the momentum is set to 0.9, and the learning rate decay is set to 0.95 using step decay. The mean Intersection over Union (mIoU) is used as the segmentation metric. mIoU is widely used to evaluate the overall performance of semantic segmentation algorithms and reflects the degree of overlap between the predicted results and their corresponding true labels. mIoU is formulated as:<disp-formula id="e3">
<mml:math id="m9">
<mml:mrow>
<mml:mtext>mIoU</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
</mml:mstyle>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x222a;</mml:mo>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where <italic>P</italic>
<sub>
<italic>k</italic>
</sub> and <italic>G</italic>
<sub>
<italic>k</italic>
</sub> are the predicted result of the <italic>k</italic>th image and the corresponding true label, respectively.</p>
</sec>
<sec id="s4-2">
<title>4.2 Comparison experiments</title>
<p>To evaluate the effectiveness of the proposed network, we tested it on three synthetic test datasets and one real smoke dataset, and compared it with several state-of-the-art semantic segmentation methods based on deep learning, including FCN-8S [<xref ref-type="bibr" rid="B12">12</xref>], SegNet [<xref ref-type="bibr" rid="B13">13</xref>], SMD [<xref ref-type="bibr" rid="B29">29</xref>], TBFCN [<xref ref-type="bibr" rid="B9">9</xref>], Deeplab v1 [<xref ref-type="bibr" rid="B15">15</xref>], ESPNet [<xref ref-type="bibr" rid="B31">31</xref>], LRN [<xref ref-type="bibr" rid="B32">32</xref>], DSS [<xref ref-type="bibr" rid="B24">24</xref>],HG-Net [<xref ref-type="bibr" rid="B33">33</xref>], MS-Net [<xref ref-type="bibr" rid="B34">34</xref>], W-Net [<xref ref-type="bibr" rid="B25">25</xref>], and GSANet [<xref ref-type="bibr" rid="B26">26</xref>]. In addition, we also compared it with some Transformer structures ViT [<xref ref-type="bibr" rid="B35">35</xref>], Swin-Transformer [<xref ref-type="bibr" rid="B36">36</xref>] and SegFormer [<xref ref-type="bibr" rid="B37">37</xref>]. To objectively and fairly evaluate the performance of each method, we used the same dataset and experimental configurations to train all the compared methods, and the results are shown in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison results of existing algorithms.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Methods</th>
<th colspan="3" align="center">mIoU (%)</th>
</tr>
<tr>
<th align="left">DS01</th>
<th align="left">DS02</th>
<th align="left">DS03</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">FCN-8S</td>
<td align="left">64.03</td>
<td align="left">63.28</td>
<td align="left">64.38</td>
</tr>
<tr>
<td align="left">SegNet</td>
<td align="left">56.94</td>
<td align="left">56.77</td>
<td align="left">57.18</td>
</tr>
<tr>
<td align="left">SMD</td>
<td align="left">62.88</td>
<td align="left">61.50</td>
<td align="left">62.09</td>
</tr>
<tr>
<td align="left">TBFCN</td>
<td align="left">66.67</td>
<td align="left">65.85</td>
<td align="left">66.20</td>
</tr>
<tr>
<td align="left">Deeplabv1</td>
<td align="left">68.41</td>
<td align="left">68.97</td>
<td align="left">68.71</td>
</tr>
<tr>
<td align="left">ESPNet</td>
<td align="left">61.85</td>
<td align="left">61.90</td>
<td align="left">62.77</td>
</tr>
<tr>
<td align="left">LRN</td>
<td align="left">66.43</td>
<td align="left">67.71</td>
<td align="left">67.46</td>
</tr>
<tr>
<td align="left">DSS</td>
<td align="left">71.04</td>
<td align="left">70.01</td>
<td align="left">69.81</td>
</tr>
<tr>
<td align="left">HG-Net2</td>
<td align="left">63.58</td>
<td align="left">62.40</td>
<td align="left">63.61</td>
</tr>
<tr>
<td align="left">HG-Net8</td>
<td align="left">63.85</td>
<td align="left">63.27</td>
<td align="left">64.46</td>
</tr>
<tr>
<td align="left">W-Net</td>
<td align="left">73.06</td>
<td align="left">73.97</td>
<td align="left">73.36</td>
</tr>
<tr>
<td align="left">GSANet</td>
<td align="left">73.13</td>
<td align="left">73.81</td>
<td align="left">74.25</td>
</tr>
<tr>
<td align="left">ViT</td>
<td align="left">75.20</td>
<td align="left">75.29</td>
<td align="left">74.10</td>
</tr>
<tr>
<td align="left">Swin-Transformer</td>
<td align="left">76.49</td>
<td align="left">75.55</td>
<td align="left">75.80</td>
</tr>
<tr>
<td align="left">SegFormer</td>
<td align="left">78.76</td>
<td align="left">78.50</td>
<td align="left">78.03</td>
</tr>
<tr>
<td align="left">Our</td>
<td align="left">78.61</td>
<td align="left">77.63</td>
<td align="left">77.30</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The mIoU metrics achieved by our method on the three virtual smoke test datasets are 78.61%, 77.63% and 77.30%, respectively. Our method achieves the good mIoUs among all the existing methods second only to the SegFormer.</p>
<p>The visualized segmentation results on virtual smoke images are shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, where the first and second columns are the original and labeled images, respectively. According to <xref ref-type="table" rid="T1">Table 1</xref> and <xref ref-type="fig" rid="F5">Figure 5</xref>, the models with the mIoU below 70 obtain poor performance, while DSS, W-Net, and GSANet, as specially designed smoke segmentation models, have good performance. However, compared to our method, they are still slightly worse, both in terms of evaluation metrics and visual image quality. Our network has more distinct boundaries that are basically consistent with the original image. Compared with the latest transformer structures in recent years, our model also has good performance, only slightly worse than Segformer.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Segmentation results for the virtual smoke test dataset. <bold>(A)</bold> Virtual smoke images, <bold>(B)</bold> labeled maps, <bold>(C)</bold> FCN, <bold>(D)</bold> SegNet, <bold>(E)</bold> SMD, <bold>(F)</bold> TBFCN, <bold>(G)</bold> DeepLab v1, <bold>(H)</bold> ESPNet, <bold>(I)</bold> HG-Net 2, <bold>(J)</bold> HG-Net 8, <bold>(K)</bold> W-net, <bold>(L)</bold> GSANet, <bold>(M)</bold> ViT, <bold>(N)</bold> Swin-Transformer, <bold>(O)</bold> SegFormer, <bold>(P)</bold> the proposed method.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g005.tif"/>
</fig>
<p>The segmentation results on the real smoke images are essentially as good as those on the synthetic smoke images. As shown in <xref ref-type="fig" rid="F6">Figure 6</xref>, the predicted results by our BEPA-Net are visually similar to their corresponding real images. To accurately locate smoke edges, feature maps require spatial details, local and global semantic abstractions for delineating smoke. Our BEPA-Net model is proposed to solve these problems. The reasons may be that fusing multi-scale features can easily extract global and local information for better smoke representations.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The segmentation results on the real dataset. <bold>(A)</bold> Realistic smoke images, <bold>(B)</bold> FCN, <bold>(C)</bold> SegNet, <bold>(D)</bold> SMD, <bold>(E)</bold> TBFCN, <bold>(F)</bold> DeepLab v1, <bold>(G)</bold> ESPNet, <bold>(H)</bold> HG-Net 2, <bold>(I)</bold> HG-Net 8, <bold>(J)</bold> W-net, <bold>(K)</bold> GSANet, <bold>(L)</bold> ViT, <bold>(M)</bold> Swin-Transformer, <bold>(N)</bold> SegFormer, <bold>(O)</bold> the proposed method.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g006.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="F6">Figures 6H, I</xref>, the generalization of HG-Net is poor. Although it has achieved certain results on virtual datasets, the performance on real images is poor. The reason may be the lack of skip connections to complete the fusion of deep and shallow features. The segmentation area obtained by our method is basically consistent with the real smoke area. In addition, by comparing the visualized results of virtual smoke datasets and real data, we find that Transformers obtain good results on virtual data, but the results on real data were very poor, as shown in <xref ref-type="fig" rid="F6">Figures 6L&#x2013;N</xref>. Hence, it may be overfitting.</p>
<p>Compared with DSS, W-Net, and GSANet, our method uses multi-scale fusion and skip connections, resulting in better performance than them. This is because the pixel alignment is performed during feature fusion. This technique greatly improves model performance. We also compared the fusion methods in ablation experiments.</p>
</sec>
<sec id="s4-3">
<title>4.3 Ablation experiments</title>
<p>Commonly used methods for feature fusion are pixel-wise addition (Addition) and channel concatenation (Concat). In this paper, we use the technique of pixel alignment followed by feature fusion to bridge the semantic gap between different feature maps during channel fusion. Further comparison results of feature fusion are shown in <xref ref-type="table" rid="T2">Table 2</xref>. The experiments show that performing pixel alignment first and then fusing features can achieve better results than other configurations. It proves that the proposed modules are powerful for feature representations.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Ablation experiments for feature-fused.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Network architecture</th>
<th colspan="3" align="center">mIOU(%)</th>
<th colspan="3" align="center">Statistical analyses</th>
</tr>
<tr>
<th align="left">DS01</th>
<th align="left">DS02</th>
<th align="left">DS03</th>
<th align="left">t score</th>
<th align="left">
<italic>p</italic>-value</th>
<th align="left">Significant</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">ResNet &#x2b; Addition</td>
<td align="left">70.68</td>
<td align="left">66.22</td>
<td align="left">67.15</td>
<td align="left">1.8555</td>
<td align="left">0.1371202</td>
<td align="left">no (<italic>p</italic>&#x3e;5%)</td>
</tr>
<tr>
<td align="left">ResNet &#x2b; Concat</td>
<td align="left">71.84</td>
<td align="left">67.18</td>
<td align="left">68.49</td>
<td align="left">1.1796</td>
<td align="left">0.3035128</td>
<td align="left">no (<italic>p</italic>&#x3e;5%)</td>
</tr>
<tr>
<td align="left">ResNet &#x2b; Concat &#x2b; PAM</td>
<td align="left">73.35</td>
<td align="left">69.64</td>
<td align="left">70.78</td>
<td align="left">&#x2014;</td>
<td align="left">&#x2014;</td>
<td align="left">&#x2014;</td>
</tr>
<tr>
<td align="left">ResNet &#x2b; Addition &#x2b; PAM</td>
<td align="left">72.81</td>
<td align="left">68.38</td>
<td align="left">69.46</td>
<td align="left">0.6022</td>
<td align="left">0.5795012</td>
<td align="left">no (<italic>p</italic>&#x3e;5%)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To evaluate the performance of the proposed modules, several ablation experiments were performed on the data set for the different combinations of the proposed modules in our network. The results are shown in <xref ref-type="table" rid="T3">Table 3</xref>. We adopt ResNet-50 as the backbone network of all the variants of our method for ablation experiments. In <xref ref-type="table" rid="T3">Table 3</xref>, Atrous-Conv indicates the improved Atrous convolution, PPM is the Pyramid Pooling Module, BEM stands for the Boundary Enhancement Module, and PAM denotes the pixel alignment module.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Ablation experiments of different modules.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="center">ResNet50</th>
<th rowspan="2" align="center">Atrous-conv</th>
<th rowspan="2" align="center">PPM</th>
<th rowspan="2" align="center">BEM</th>
<th rowspan="2" align="center">PAM</th>
<th colspan="3" align="center">mIOU(%)</th>
<th colspan="3" align="center">Statistical analyses</th>
</tr>
<tr>
<th align="center">DS01</th>
<th align="center">DS02</th>
<th align="center">DS03</th>
<th align="left">t score</th>
<th align="left">
<italic>p</italic>-value</th>
<th align="left">Significant</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="center">&#x2713;</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="center">61.67</td>
<td align="center">60.23</td>
<td align="center">62.09</td>
<td align="left">24.04</td>
<td align="left">0.0000178</td>
<td align="left">yes (<italic>p</italic>&#x3c;5%)</td>
</tr>
<tr>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="left"/>
<td align="left"/>
<td align="left"/>
<td align="center">64.03</td>
<td align="center">63.28</td>
<td align="center">64.38</td>
<td align="left">27.36</td>
<td align="left">0.0000106</td>
<td align="left">yes (<italic>p</italic>&#x3c;5%)</td>
</tr>
<tr>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="left"/>
<td align="left"/>
<td align="center">71.83</td>
<td align="center">67.68</td>
<td align="center">69.47</td>
<td align="left">6.474</td>
<td align="left">0.0029330</td>
<td align="left">yes (<italic>p</italic>&#x3c;5%)</td>
</tr>
<tr>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="left"/>
<td align="center">73.53</td>
<td align="center">69.70</td>
<td align="center">70.16</td>
<td align="left">5.290</td>
<td align="left">0.0061303</td>
<td align="left">yes (<italic>p</italic>&#x3c;5%)</td>
</tr>
<tr>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">&#x2713;</td>
<td align="center">78.61</td>
<td align="center">77.63</td>
<td align="center">77.30</td>
<td align="left">&#x2014;</td>
<td align="left">&#x2014;</td>
<td align="left">&#x2014;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to <xref ref-type="table" rid="T3">Table 3</xref>, we find that the performance of the proposed network can be obviously enhanced by employing Atrous convolutions in the backbone network. After the pyramid pooling module (PPM) is enabled, we achieve the mIoUs of 71.83%, 67.68% and 69.47% on the test datasets of DS01, DS02, and DS03, respectively. The mIoUs are improved by about 2% after using the Boundary Enhancement Module (BEM). Although the boundary pixels often occupy a relatively small portion of the whole image, boundary information plays a key role in improving the accuracy of segmentation. The mIoUs are greatly improved by 7%&#x2013;8% after the pixel alignment module (PAM) is used.</p>
<p>In addition, we compute the <italic>p</italic>-values of our results in <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T3">3</xref> for analyzing the statistical significance of ablation experiments. Given two random sets <italic>X</italic> and <italic>Y</italic>, the <italic>t</italic>-score of the two sets is computed as follows:<disp-formula id="e4">
<mml:math id="m10">
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>&#x3bc;</mml:mi>
<mml:mi>y</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>y</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>y</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>y</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>x</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>y</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>&#x3bc;</italic>
<sub>
<italic>x</italic>
</sub>, <italic>&#x3c3;</italic>
<sub>
<italic>x</italic>
</sub> and <italic>n</italic>
<sub>
<italic>x</italic>
</sub> are respectively the mean, the standard deviation and the sample number of the best set <italic>X</italic>, and <italic>&#x3bc;</italic>
<sub>
<italic>y</italic>
</sub>, <italic>&#x3c3;</italic>
<sub>
<italic>y</italic>
</sub> and <italic>n</italic>
<sub>
<italic>y</italic>
</sub> are the mean, the standard deviation and the sample number of the tested set <italic>Y</italic>, respectively.</p>
<p>According to the <italic>t</italic>-scores and degree of freedom <italic>d</italic>
<sub>
<italic>f</italic>
</sub> &#x3d; <italic>n</italic>
<sub>
<italic>x</italic>
</sub> &#x2b; <italic>n</italic>
<sub>
<italic>y</italic>
</sub>-2, we can find the ranges of the <italic>p</italic>-values in the <italic>t</italic>-score lookup table. For the sake of convenience, we use the <italic>Excel</italic> software to automatically compute the <italic>p</italic>-values. By observing <xref ref-type="table" rid="T2">Table 2</xref>, we find that all the <italic>p</italic>-values are greater than 5%, so it is not significant in statistics. In vision fields, although existing methods achieving only 1% increase of mIoUs are not significant statistically, they are often considered as excellent algorithms, and they do not perform <italic>p</italic>-value analyses. In <xref ref-type="table" rid="T3">Table 3</xref>, we find that all the <italic>p</italic>-values are far less than 5%, so the best variant is statistically significant.</p>
</sec>
<sec id="s4-4">
<title>4.4 Testing on wild scenes with electric power transmission lines</title>
<p>The proposed method also achieves good results for real fires in wild scenes, as shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. The proposed method can accurately segment smoke regions. Although there are smoke-like objects, such as clouds, our model easily discriminates between clouds and smokes.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Smoke segmentation results in wilderness scenes.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g007.tif"/>
</fig>
<p>In order to ensure the safety of electric power transmission lines, it is necessary to check fire safety around the electric lines. We tested the proposed method on several images captured from iron towers of electric power transmission lines, as shown in <xref ref-type="fig" rid="F8">Figure 8</xref>. From the experimental results, we can see that the proposed method not only detects smoke successfully, but also obtain a relatively accurate smoke contours. Segmented smoke contours can allow the relevant personnel in the electric power department to have a more accurate judgment of the possible fire spread trend, and take corresponding countermeasures in advance to determine the safety of power lines.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Experiments on images of electric power transmission lines. Reproduced from the Yunnan Electric Power Company, with permission from the Company.</p>
</caption>
<graphic xlink:href="fphy-11-1206944-g008.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s5">
<title>5 Conclusion</title>
<p>In this paper, a deep neural network is proposed to improve the performance of smoke semantic segmentation. To learn the spatial details and contextual information about objects, we design a spatial attention mechanism to enhance the localization accuracy of object boundaries for improving the representation ability of the network. To improve the segmentation performance of blurry smoke objects, we use Atrous convolutions with different rates and the Pyramid Pooling Module (PPM) to obtain contextual and abstract information. To effectively aggregate features, we propose a Pixel Alignment Module (PAM) to recalibrate the position of features and produce more powerful features. Compared with other excellent semantic segmentation algorithms, the proposed method consistently outperforms existing algorithms on the three synthetic smoke datasets and real smoke images. In addition, our method also achieves very good segmentation results on images captured from wild scenes with electric power transmission lines. However, our method is still not lightweight enough and requires a lot of computational resources. Compared to the existing transformer structure, its performance cannot achieve optimal results. In future work, we will focus on lightweight and Transformer structures to further improve accuracy.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>FZ was responsible for scheme designs and method requirements, GaW and YW collected and annotated training data, HP completed the testing experiments of the proposed model, GuW designed, trained the network and drafted the paper, and FY created the test dataset and revised the paper. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by the Major Scientific and Technological Projects of Yunnan Province (202202AD080010).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>Authors FZ, GW, HP, and YW were employed by Electric Power Research Institute of Yunnan Electric Power Company.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S</given-names>
</name>
</person-group>. <article-title>Encoding pairwise Hamming distances of Local Binary Patterns for visual smoke recognition</article-title>. <source>Computer Vis Image Understanding</source> (<year>2019</year>) <volume>178</volume>:<fpage>43</fpage>&#x2013;<lpage>53</lpage>. <pub-id pub-id-type="doi">10.1016/j.cviu.2018.10.008</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Mei</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>High-order local ternary patterns with locality preserving projection for smoke detection and image classification</article-title>. <source>Inf Sci</source> (<year>2016</year>) <volume>372</volume>:<fpage>225</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2016.08.040</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Learning-based smoke detection for unmanned aerial vehicles applied to forest fire surveillance</article-title>. <source>J Intell Robotic Syst</source> (<year>2019</year>) <volume>93</volume>(<issue>1</issue>):<fpage>337</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1007/s10846-018-0803-y</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mahmoud</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Forest fire detection and identification using image processing and SVM</article-title>. <source>J Inf Process Syst</source> (<year>2019</year>) <volume>15</volume>(<issue>1</issue>):<fpage>159</fpage>&#x2013;<lpage>68</lpage>. <pub-id pub-id-type="doi">10.3745/JIPS.01.0038</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tian</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Ogunbona</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Detection and separation of smoke from single image frames</article-title>. <source>IEEE Trans Image Process</source> (<year>2017</year>) <volume>27</volume>(<issue>3</issue>):<fpage>1164</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2017.2771499</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>A gated recurrent network with dual classification assistance for smoke semantic segmentation</article-title>. <source>IEEE Trans Image Process</source> (<year>2021</year>) <volume>30</volume>:<fpage>4409</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1109/tip.2021.3069318</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ronneberger</surname>
<given-names>O</given-names>
</name>
<name>
<surname>Fischer</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Brox</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>U-net: convolutional networks for biomedical image segmentation</article-title>. <source>Int Conf Med image Comput computer-assisted intervention</source> (<year>2015</year>) <volume>9351</volume>:<fpage>234</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Papandreou</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Schroff</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Adam</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Rethinking atrous convolution for semantic image segmentation</article-title> (<year>2017</year>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1706.05587">https://arxiv.org/abs/1706.05587</ext-link>.</comment>
</citation>
</ref>
<ref id="B9">
<label>9.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>Multi-oriented text detection with fully convolutional networks</article-title>. In: <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>; <conf-date>June 2016</conf-date>; <conf-loc>Las Vegas, NV, USA</conf-loc> (<year>2016</year>). p. <fpage>4159</fpage>&#x2013;<lpage>67</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.451</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Girshick</surname>
<given-names>R</given-names>
</name>
<name>
<surname>Gupta</surname>
<given-names>A</given-names>
</name>
<name>
<surname>He</surname>
<given-names>K</given-names>
</name>
</person-group>. <article-title>Non-local neural networks</article-title>. In: <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>; <conf-date>June 2018</conf-date>; <conf-loc>Salt Lake City, UT, USA</conf-loc> (<year>2018</year>). p. <fpage>7794</fpage>&#x2013;<lpage>803</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2018.00813</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Pyramid scene parsing network</article-title>. In: <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>; <conf-date>July 2017</conf-date>; <conf-loc>Honolulu, HI, USA</conf-loc> (<year>2017</year>). p. <fpage>6230</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2017.660</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Long</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Shelhamer</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Darrell</surname>
<given-names>T</given-names>
</name>
</person-group>. <article-title>Fully convolutional networks for semantic segmentation</article-title>. <source>IEEE Trans Pattern Anal Machine Intelligence</source> (<year>2017</year>) <volume>39</volume>(<issue>4</issue>):<fpage>640</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2572683</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Badrinarayanan</surname>
<given-names>V</given-names>
</name>
<name>
<surname>Kendall</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Cipolla</surname>
<given-names>R</given-names>
</name>
</person-group>. <article-title>Segnet: a deep convolutional encoder-decoder architecture for image segmentation</article-title>. <source>
<italic>IEEE Tra</italic>ns pattern Anal machine intelligence</source> (<year>2017</year>) <volume>39</volume>(<issue>12</issue>):<fpage>2481</fpage>&#x2013;<lpage>95</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2016.2644615</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Shan</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Dual super-resolution learning for semantic segmentation</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>; <conf-date>June 2020</conf-date>; <conf-loc>Seattle, WA, USA</conf-loc> (<year>2020</year>). p. <fpage>3774</fpage>&#x2013;<lpage>83</lpage>.</citation>
</ref>
<ref id="B15">
<label>15.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Papandreou</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Kokkinos</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Yuile</surname>
<given-names>A</given-names>
</name>
</person-group>. <article-title>Semantic image segmentation with deep convolutional nets and fully connected crfs</article-title> (<year>2014</year>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1412.7062">https://arxiv.org/abs/1412.7062</ext-link>.</comment>
</citation>
</ref>
<ref id="B16">
<label>16.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Papandreou</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Kokkinos</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Yuile</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Yuille</surname>
<given-names>AL</given-names>
</name>
</person-group>. <article-title>Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs</article-title>. <source>IEEE Trans pattern Anal machine intelligence</source> (<year>2017</year>) <volume>40</volume>(<issue>4</issue>):<fpage>834</fpage>&#x2013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2699184</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Papandreou</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Schroff</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Adam</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</article-title>. <source>Proc Eur Conf Comput Vis</source> (<year>2018</year>) <volume>11211</volume>:<fpage>833</fpage>&#x2013;<lpage>51</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_49</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Filonenko</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Hern&#xe1;ndez</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Jo</surname>
<given-names>K</given-names>
</name>
</person-group>. <article-title>Fast smoke detection for video surveillance using CUDA</article-title>. <source>IEEE Trans Ind Inform</source> (<year>2017</year>) <volume>14</volume>(<issue>2</issue>):<fpage>725</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1109/TII.2017.2757457</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dimitropoulos</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Barmpoutis</surname>
<given-names>P</given-names>
</name>
<name>
<surname>Grammalidis</surname>
<given-names>N</given-names>
</name>
</person-group>. <article-title>Higher order linear dynamical systems for smoke detection in video surveillance applications</article-title>. <source>IEEE Trans Circuits Syst Video Tech</source> (<year>2017</year>) <volume>27</volume>(<issue>5</issue>):<fpage>1143</fpage>&#x2013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2016.2527340</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Candidate smoke region segmentation of fire video based on rough set theory</article-title>. <source>J Electr Comp Eng</source> (<year>2015</year>) <volume>2015</volume>:<fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1155/2015/280415</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tao</surname>
<given-names>C</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>Smoke detection based on deep convolutional neural networks</article-title>. In: <conf-name>Proceedings of the 2016 International Conference on Industrial Informatics - Computing Technology, Intelligent Technology, Industrial Information Integration (ICIICII)</conf-name>; <conf-date>December 2016</conf-date>; <conf-loc>Wuhan, China</conf-loc> (<year>2016</year>). p. <fpage>150</fpage>&#x2013;<lpage>3</lpage>. <pub-id pub-id-type="doi">10.1109/ICIICII.2016.0045</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G</given-names>
</name>
</person-group>. <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Assoc Comput Machinery</source> (<year>2017</year>) <volume>60</volume>(<issue>6</issue>):<fpage>84</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1145/3065386</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>C</given-names>
</name>
</person-group>. <article-title>3D parallel fully convolutional networks for real-time video wildfire smoke detection</article-title>. <source>IEEE Trans Circuits Syst Video Tech</source> (<year>2020</year>) <volume>30</volume>(<issue>1</issue>):<fpage>89</fpage>&#x2013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1109/TCSVT.2018.2889193</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>Deep smoke segmentation</article-title>. <source>Neurocomputing</source> (<year>2019</year>) <volume>357</volume>:<fpage>248</fpage>&#x2013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.05.011</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>Q</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>A wave-shaped deep neural network for smoke density estimation</article-title>. <source>IEEE Trans Image Process</source> (<year>2019</year>) <volume>29</volume>:<fpage>2301</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2019.2946126</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dong</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Xue</surname>
<given-names>X</given-names>
</name>
</person-group>. <article-title>Improved spatial and channel information based global smoke attention network</article-title>. <source>J Beijing Univ Aeronautics Astronautics</source> (<year>2022</year>) <volume>48</volume>(<issue>8</issue>):<fpage>1471</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.13700/j.bh.1001-5965.2021.0549</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>; <conf-date>June 2016</conf-date>; <conf-loc>Las Vegas, NV, USA</conf-loc> (<year>2016</year>). p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dai</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Qi</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>H</given-names>
</name>
<etal/>
</person-group> <article-title>Deformable convolutional networks</article-title>. In: <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>; <conf-date>October 2017</conf-date>; <conf-loc>Venice, Italy</conf-loc> (<year>2017</year>). p.<fpage>764</fpage>&#x2013;<lpage>73</lpage>. <pub-id pub-id-type="doi">10.1109/ICCV.2017.89</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Shao</surname>
<given-names>L</given-names>
</name>
</person-group>. <article-title>Video salient object detection via fully convolutional networks</article-title>. <source>IEEE Trans Image Process</source> (<year>2017</year>) <volume>27</volume>(<issue>1</issue>):<fpage>38</fpage>&#x2013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1109/TIP.2017.2754941</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>LC</given-names>
</name>
<name>
<surname>Papandreou</surname>
<given-names>G</given-names>
</name>
<name>
<surname>Kokkinos</surname>
<given-names>I</given-names>
</name>
<name>
<surname>Murphy</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Yuille</surname>
<given-names>AL</given-names>
</name>
</person-group>. <article-title>Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs</article-title>. <source>IEEE Trans pattern Anal machine intelligence</source> (<year>2017</year>) <volume>40</volume>(<issue>4</issue>):<fpage>834</fpage>&#x2013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1109/tpami.2017.2699184</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mehta</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Rastegari</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Caspi</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Shapiro</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Hajishirzi</surname>
<given-names>H</given-names>
</name>
</person-group>. <article-title>Espnet: efficient spatial pyramid of dilated convolutions for semantic segmentation</article-title>. <source>Eur Conf Comput Vis</source> (<year>2018</year>) <volume>11214</volume>:<fpage>561</fpage>&#x2013;<lpage>80</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-01249-6_34</pub-id>
</citation>
</ref>
<ref id="B32">
<label>32.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Islam</surname>
<given-names>MA</given-names>
</name>
<name>
<surname>Naha</surname>
<given-names>S</given-names>
</name>
<name>
<surname>Rochan</surname>
<given-names>M</given-names>
</name>
<name>
<surname>Bruce</surname>
<given-names>N</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y</given-names>
</name>
</person-group>. <article-title>Label refinement network for coarse-to-fine semantic segmentation</article-title> (<year>2017</year>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1703.00551">https://arxiv.org/abs/1703.00551</ext-link>.</comment>
</citation>
</ref>
<ref id="B33">
<label>33.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Newell</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>K</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Stacked hourglass networks for human pose estimation</article-title>. <source>Eur Conf Comput Vis</source> (<year>2016</year>) <volume>9912</volume>:<fpage>483</fpage>&#x2013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-319-46484-8_29</pub-id>
</citation>
</ref>
<ref id="B34">
<label>34.</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>F</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>B</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>J</given-names>
</name>
</person-group>. <article-title>Convolutional neural networks based on multi-scale additive merging layers for visual smoke recognition</article-title>. <source>Machine Vis Appl</source> (<year>2019</year>) <volume>30</volume>:<fpage>345</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1007/s00138-018-0990-3</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Dosovitskiy</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Beyer</surname>
<given-names>L</given-names>
</name>
<name>
<surname>Kolesnikov</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Weissenborn</surname>
<given-names>D</given-names>
</name>
<name>
<surname>Zhai</surname>
<given-names>X</given-names>
</name>
<name>
<surname>Unterthiner</surname>
<given-names>T</given-names>
</name>
<etal/>
</person-group> <article-title>An image is worth 16x16 words: transformers for image recognition at scale</article-title>. In: <conf-name>Proceedings of the International Conference on Learning Representations</conf-name>; <conf-date>May 2021</conf-date>; <conf-loc>Vienna, Austria</conf-loc> (<year>2021</year>).</citation>
</ref>
<ref id="B36">
<label>36.</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>H</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>Y</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z</given-names>
</name>
<etal/>
</person-group> <article-title>Swin transformer: hierarchical vision transformer using shifted windows</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <conf-date>October 2021</conf-date>; <conf-loc>Montreal, QC, Canada</conf-loc> (<year>2021</year>). p. <fpage>9992</fpage>&#x2013;<lpage>10002</lpage>.</citation>
</ref>
<ref id="B37">
<label>37.</label>
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Xie</surname>
<given-names>E</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z</given-names>
</name>
<name>
<surname>Anandkumar</surname>
<given-names>A</given-names>
</name>
<name>
<surname>Alvarez</surname>
<given-names>J</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>P</given-names>
</name>
</person-group>. <article-title>SegFormer: simple and efficient design for semantic segmentation with transformers</article-title> (<year>2021</year>). <comment>Available at: <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/2105.15203">https://arxiv.org/abs/2105.15203</ext-link>.</comment>
</citation>
</ref>
</ref-list>
</back>
</article>