<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Oncol.</journal-id>
<journal-title>Frontiers in Oncology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Oncol.</abbrev-journal-title>
<issn pub-type="epub">2234-943X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fonc.2022.1087438</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Oncology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Automatic polyp image segmentation and cancer prediction based on deep learning</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Shen</surname>
<given-names>Tongping</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2080301"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Xueguang</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Information Engineering, Anhui University of Chinese Medicine</institution>, <addr-line>Hefei</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Graduate School, Angeles University Foundation</institution>, <addr-line>Angeles</addr-line>, <country>Philippines</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>School of Computer Science and Technology, Henan Institute of Technology</institution>, <addr-line>Xinxiang</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Jialiang Yang, Geneis (Beijing) Co. Ltd, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Palash Ghosal, Sikkim Manipal University, India; Luisa F. S&#xe1;nchez-Peralta, Jesus Uson Minimally Invasive Surgery Centre, Spain</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Tongping Shen, <email xlink:href="mailto:shentp2010@ahtcm.edu.cn">shentp2010@ahtcm.edu.cn</email>
</p>
</fn>
<fn fn-type="other" id="fn002">
<p>This article was submitted to Cancer Imaging and Image-directed Interventions, a section of the journal Frontiers in Oncology</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>12</volume>
<elocation-id>1087438</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>11</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Shen and Li</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Shen and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The similar shape and texture of colonic polyps and normal mucosal tissues lead to low accuracy of medical image segmentation algorithms. To solve these problems, we proposed a polyp image segmentation algorithm based on deep learning technology, which combines a HarDNet module, attention module, and multi-scale coding module with the U-Net network as the basic framework, including two stages of coding and decoding. In the encoder stage, HarDNet68 is used as the main backbone network to extract features using four null space convolutional pooling pyramids while improving the inference speed and computational efficiency; the attention mechanism module is added to the encoding and decoding network; then the model can learn the global and local feature information of the polyp image, thus having the ability to process information in both spatial and channel dimensions, to solve the problem of information loss in the encoding stage of the network and improving the performance of the segmentation network. Through comparative analysis with other algorithms, we can find that the network of this paper has a certain degree of improvement in segmentation accuracy and operation speed, which can effectively assist physicians in removing abnormal colorectal tissues and thus reduce the probability of polyp cancer, and improve the survival rate and quality of life of patients. Also, it has good generalization ability, which can provide technical support and prevention for colon cancer.</p>
</abstract>
<kwd-group>
<kwd>colon polyps</kwd>
<kwd>attention mechanism</kwd>
<kwd>HarDNet</kwd>
<kwd>image segmentation</kwd>
<kwd>deep learning</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="3"/>
<equation-count count="20"/>
<ref-count count="45"/>
<page-count count="12"/>
<word-count count="7013"/>
</counts>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Colon cancer is a malignant tumor on the colonic mucosa, mostly formed by adenomatous polyps, characterized by high incidence and high lethality, and is currently one of the three most prevalent malignancies in the world (<xref ref-type="bibr" rid="B1">1</xref>). Colon polyps are convexities that grow in the colonic mucosa, and the gold standard for early screening of colorectal cancer is the use of colonoscopy to detect polyps larger than 5&#xa0;mm in diameter in the intestine (<xref ref-type="bibr" rid="B2">2</xref>). There is a link between the accurate detection rate of polyps and the incidence of colon cancer. Dougla et&#xa0;al. (<xref ref-type="bibr" rid="B3">3</xref>) showed that for every 1% increase in polyp detection, the prevalence of colorectal cancer would decrease by 3%. Prevention and diagnosis of colorectal cancer through early screening are very important and can improve patient survival rates. Polyps can usually be effectively detected by colonoscopic screening, but even if the polyps are of the same type, it varies in size, color, and texture. Secondly, in colonoscopic images of intestinal polyps, the contrast between the polyps and the surrounding mucosa is not strong enough and the border is blurred and not clear enough due to the intestinal mucus and the reflection of the intestinal polyps under colonoscopy. Therefore, it may cause the physician to miss the polyps and segment them inaccurately. Therefore, how to segment colon polyps quickly and accurately is important for the early prevention of colorectal cancer coming (<xref ref-type="bibr" rid="B4">4</xref>).</p>
<p>For the segmentation of colon polyp images, a large number of methods exist at home and abroad. They can be mainly included two types: early traditional algorithms and deep learning-based algorithms. The traditional segmentation methods mainly extract features such as color, shape, and texture, and then use classifiers to distinguish polyps from their surrounding non-polyp regions.</p>
<p>Saul et&#xa0;al. in 2009 proposed the use of similarity measures to discriminate colon polyps as a way to reduce the workload of physicians (<xref ref-type="bibr" rid="B5">5</xref>). Bashar et&#xa0;al. in 2010 proposed the segmentation of the capsule endoscopic images from the possible presence of polyps based on the energy distribution of the images (<xref ref-type="bibr" rid="B6">6</xref>). Segui et&#xa0;al. in 2012 used the texture and color information as the basis for screening (<xref ref-type="bibr" rid="B7">7</xref>). Turcza et&#xa0;al. in 2013 proposed to use an entropy encoder to initially screen capsule endoscopic images and filter out some highly similar images (<xref ref-type="bibr" rid="B8">8</xref>). Hassan et&#xa0;al. in 2015 proposed to perform gray value statistics on capsule endoscopic images and observe their spectral features to screen out images of polyps (<xref ref-type="bibr" rid="B9">9</xref>). In 2016, Qiao et&#xa0;al. attempted to convert capsule endoscopy images from red, green, and blue (RGB) color gamut to columnar color gamut, which enhanced the contrast between the polyp part and normal bowel part (<xref ref-type="bibr" rid="B10">10</xref>). In 2014, Mamonov et&#xa0;al. designed a binary classifier to label each frame of a colonoscopy video based on the geometric analysis and texture content on each frame as containing or not containing polyps (<xref ref-type="bibr" rid="B11">11</xref>). In 2015, Tajbakhsh et&#xa0;al. used an automatic method to detect colonoscopy video polyps based on shape and contextual information (<xref ref-type="bibr" rid="B12">12</xref>). In 2015, Bernal et&#xa0;al. designed a model based on the appearance of polyps using the median depth of the valley accumulation window. The algorithm associated with the probability of polyp presence WM-DOVA energy map to obtain the specific range of distinction between polyps and surrounding tissue and thus determine the specific location of polyps (<xref ref-type="bibr" rid="B13">13</xref>).</p>
<p>In these early algorithms for image segmentation of colon polyps, more reliance was placed on manual efforts to extract feature information from the images, followed by classifiers to segment polyps and tissues.</p>
<p>Although the traditional algorithms are relatively simple in implementation, they cannot combine the effective features of the polyp region by considering them simultaneously. At the same time, since the designed classifiers usually produce effects only on the specified polyp dataset, the traditional algorithms have poor generalization and the segmentation effects on other polyp datasets are significantly reduced.</p>
<p>Deep neural network techniques are rapidly developing and are increasingly using in polyp image segmentation.</p>
<p>Brandao et&#xa0;al. proposed Fully Convolutional Networks (FCN) for recognition and segmentation of polyp images in 2017 and achieved good polyp segmentation results (<xref ref-type="bibr" rid="B14">14</xref>). Wang used Dynamic Convolution Neural Network (DCNN) model in 2018 to validate and evaluate two publicly available polyp datasets (<xref ref-type="bibr" rid="B15">15</xref>). In 2018, Zhou et&#xa0;al. designed an encoder-decoder network structure UNet ++, which employs a deeply supervised training approach that allows supervised learning of the model&#x2019;s multi-branch output. The model applies more jump connections between high-dimensional and low-dimensional information, reduces the feature error between semantic information, and better improves the segmentation accuracy of colon polyp images (<xref ref-type="bibr" rid="B16">16</xref>). In 2019, Jha et&#xa0;al. improved the segmentation accuracy of colon polyp images by using residual blocks, compressed excitation blocks, null spatial pyramid pooling and attention mechanism to design the ResUNet++ network, which significantly improved the segmentation of colon polyps (<xref ref-type="bibr" rid="B17">17</xref>). Fang et&#xa0;al. in 2019 designed the Selective Feature Aggregation Network (SFANet) to predict and segment regions and boundaries of polyp images (<xref ref-type="bibr" rid="B18">18</xref>). In 2020, to capture more effective semantic information, Jha et&#xa0;al. further designed the DoubleU-Net model, which effectively bridges two U-shaped structures through void space pyramidal pooling, and validated the performance of the model on a colon polyp image segmentation dataset, but the model&#x2019;s parameter number was large (<xref ref-type="bibr" rid="B19">19</xref>).</p>
<p>Nadimi et&#xa0;al. proposed an improved AlexNet compounded with migration learning, data preprocessing, and data enhancement to detect capsule endoscopic polyps in 2020 and achieved 98.0% accuracy and 98.1% sensitivity (<xref ref-type="bibr" rid="B20">20</xref>). Owais et&#xa0;al. proposed a method in 2020 and achieved better results in 52471 endoscopic capsule images to achieve better polyp detection (<xref ref-type="bibr" rid="B21">21</xref>). Yanada used a novel deep learning automatic detection method in 2020 and produced a dataset to confirm the method&#x2019;s effectiveness for polyp images, thus improving the early detection rate of intestinal tumors (<xref ref-type="bibr" rid="B22">22</xref>). Lee et&#xa0;al. proposed in 2021 the use of a variable depth of CNN for cancer risk level assessment of various intestinal diseases such as polyps (<xref ref-type="bibr" rid="B23">23</xref>). Lai et&#xa0;al. proposed in 2021 that the raw capsule endoscopic images were multi-channel separated and then fed into a deep CNN for training, which resulted in more accurate polyp identification (<xref ref-type="bibr" rid="B24">24</xref>).</p>
<p>In 2020, Fan designed the PraNet to improve segmentation accuracy on five colon polyp datasets, allowing real-time prediction of polyps in colonoscopy detection videos (<xref ref-type="bibr" rid="B25">25</xref>). In 2021, Zhao et&#xa0;al. designed the Multi-scale Subtraction Network (MSNet). The different levels of perceptual fields are then set in a pyramidal fashion to obtain rich multi-scale difference information (<xref ref-type="bibr" rid="B26">26</xref>). In 2021, Huang et&#xa0;al. improved on PraNet by designing a simple encoder-decoder model using Harmonic DenseNet (HarDNet) (<xref ref-type="bibr" rid="B27">27</xref>) HarDNet-MSEG. The original Res2Net50 (<xref ref-type="bibr" rid="B23">23</xref>) backbone network was replaced using the HarDNet68 backbone network, and the attention mechanism was removed to achieve more accurate polyp image segmentation. However, the method still does not work well for the diversity of polyp size, shape, and texture (<xref ref-type="bibr" rid="B28">28</xref>).</p>
<p>Instead of relying on the manual acquisition of features, deep learning algorithms use colon polyp datasets to continuously train on the established neural network model and finally optimize the model with the highest segmentation accuracy. However, many polyp segmentation networks based on deep learning algorithms focus on developing complex network structures to achieve better polyp segmentation performance, resulting in increased network model parameters and larger computation, which directly affect the computational efficiency of polyp segmentation networks. As shown in reference (<xref ref-type="bibr" rid="B29">29</xref>), the model parameter size of UACANet-L is 69.15M, HRNetV2-W48 model parameter size is 65.84M, while the parameter size of the proposed model in this paper is 23.11M.</p>
<p>To solve these problems, we propose a multi-scale coded colon polyp segmentation network combining HarDNet and attention mechanism; the segmentation network uses the U-Net network as the basic framework, including two stages of encoding and decoding. In the encoder stage, HarDNet68 is used as the backbone network to extract features using four null space convolutional pooling pyramids while improving the inference speed and computational efficiency. The attention mechanism module is added to the encoding and decoding networks so that the segmentation network can learn the global and local feature information of polyp images, thus having the information processing capability in both spatial and channel dimensions, solving the problem of encoder part of the information loss and the difficulty of small lesion segmentation.</p>
<p>The main contributions of this paper are as follows:</p>
<list list-type="simple">
<list-item>
<p>(1) Considering the impact of computation and memory access on the model design, this paper adopts HarDNet68 as the backbone network. HarDNet68 network can both learn the global feature information of polyp images and reduce the computation of the model, thus improving the operation speed of the model and the segmentation effect and accuracy of polyp images.</p>
</list-item>
<list-item>
<p>(2) An integrated spatial and channel attention (SCA) module is proposed, which can assign different attention weights from both spatial and channel dimensions, enabling the model to focus more on the image segmentation task. The model can be integrated into mainstream neural network segmentation tasks.</p>
</list-item>
<list-item>
<p>(3) The DenseASPP module is used in the segmentation network, which constitutes a dense feature pyramid, and the field obtains a larger perceptual field to improve the model&#x2019;s ability to obtain information about image features.</p>
</list-item>
<list-item>
<p>(4) To address the problem that the targets segmented in this paper have a small proportion in the image, the combined algorithm of Dice loss and Focal loss is proposed as a loss function to reduce the weight of simple samples and improve the segmentation accuracy of small target samples.</p>
</list-item>
<list-item>
<p>(5) We perform analytical analysis on five polyps datasets with different sizes and performances, as well as generalization experiments, etc., to verify the effectiveness and excellence of the algorithms in this paper.</p>
</list-item>
</list>
</sec>
<sec id="s2">
<label>2</label>
<title>The proposed architecture</title>
<p>The proposed polyp image segmentation network is based on the traditional U-Net network structure, including the encoder-decoder structure. Incorporating the HarDNet68 structure, the attention mechanism module, and the multi-scale null convolution module into the network, as shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. The input image is subjected to a 3&#xd7;3 convolution operation, followed by five consecutive HarDNet integration blocks. HarDNet68 improves the global dense connection of the original DenseNet into a sparse connection with the convolution, BN, and ReLU activation functions to achieve repetition of batch normalization, reducing the number of parameters to obtain shorter inference speed while maintaining accuracy. The output of the encoder is used as the input of the multi-scale cavity convolution module, thus capturing more visual information of different scale features while capturing feature information of different scales. The decoder consists of three SCA modules, which obtains the spatial and attentional feature outputs on the channels by dot-product operations from the output of the previous stage and the corresponding edge outputs in HarDNet68. The Sigmoid function obtains the final segmentation result at the output of the decoder.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>The proposed architecture.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g001.tif"/>
</fig>
<p>The segmentation network proposed in this paper forms an iterative interaction mechanism between the encoder and the decoder, which can effectively correct the conflicting regions such as boundaries in the prediction results thus improving the segmentation accuracy, as well as enhancing the inference speed and computational efficiency of the model.</p>
<sec id="s2_1">
<label>2.1</label>
<title>U-net</title>
<p>The U-net network structure (<xref ref-type="bibr" rid="B30">30</xref>) is all convolutional layers without fully connected layers, as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. In the encoder part, the feature of the image is continuously extracted through the convolution layer, and the size of the feature image is also reduced. In the decoder part, the image is restored to the original size by de-convolution, and in the decoding part, the same size of the feature map in the encoding process is connected through cross-links, so that the image information features are lost as little as possible.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>U-net architecture (<xref ref-type="bibr" rid="B29">29</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g002.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>HardNet</title>
<p>HarDNet structure, a harmonic densely connected network proposed by Chao et&#xa0;al. in 2019 (<xref ref-type="bibr" rid="B27">27</xref>), is optimized and improved based on the DenseNet network structure, optimization and improvement are made to improve the model running speed, as shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>HarDNet network structure (<xref ref-type="bibr" rid="B27">27</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g003.tif"/>
</fig>
<p>HarDNet uses a sparse connection, assuming that layer k is connected to layer , 2<sup>n</sup> which can divide k integer, where n&#x2265;0, k&#x2013;2<sup>n</sup>&#x2265;0. With this connection, if 2<sup>n</sup> is processed, layer 1 to layer 2<sup>n</sup>&#x2013;1 can be cleared from memory, reducing the amount of model parameter computation. At the same time, HarDNet balances the channel ratio between the input and output of key layer layers by increasing the number of channels of some key layers to reduce the amount of model memory access.</p>
<p>According to the above design ideas, six HarDNet structures with different parameter settings are proposed, and the HarDNet68 structure is used in this paper. <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref> shows the parameter settings of HarDNet68.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>HarDNet68 parameters.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">m</th>
<th valign="middle" align="center">Stride 2</th>
<th valign="middle" align="center">Stride 4</th>
<th valign="middle" align="center">Stride 8</th>
<th valign="middle" align="center">Stride 16</th>
<th valign="middle" align="center">Stride32</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="2" align="left">1.7</td>
<td valign="middle" align="center">3&#xd7;3,32,<break/>Stride=2</td>
<td valign="middle" rowspan="2" align="center">8(HDB),<break/>k=14,t=128</td>
<td valign="middle" align="center">16(HDB),<break/>k=16,t=256</td>
<td valign="middle" rowspan="2" align="center">16(HDB),<break/>k=40,t=640</td>
<td valign="middle" rowspan="2" align="center">4(HDB),<break/>k=160,t=1024</td>
</tr>
<tr>
<td valign="middle" align="center">3&#xd7;3,64</td>
<td valign="middle" align="center">16(HDB),<break/>k=20,t=320</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Multiplier m is the low-dimensional compression factor, &#x201c;3&#xd7;3, 32&#x201d; means 32 convolutional layers with 3&#xd7;3, HDB is the number of HarDNet modules, k is the growth rate, and t is the number of output channels of 1 &#xd7; 1 convolutional transition.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>DenseASPP</title>
<p>In the deep learning network model, the size of the image is generally uniform by stretching or cropping, but this causes information loss, distortion, and many other problems. He et&#xa0;al. proposed the Spatia Pyramid Pooling (SPP) structure, which is to use multiple pooling layers of different scales for feature extraction and fusion into an n-dimensional vector input to a fully connected layer (<xref ref-type="bibr" rid="B31">31</xref>). The Google team proposed the Atrous Spatia Pyramid Pooling (ASPP) structure based on the characteristics of multi-scale information and extension convolution in the DeepLab (<xref ref-type="bibr" rid="B32">32</xref>) series of work. ASPP introduces the concept of void convolution on the basis of SPP and further improves the ability of image feature information extraction.</p>
<p>However, the polyps have various shapes and textures, and the picture background is complicated, which easily causes poor segmentation effect in the process of polyp image segmentation. We introduces the DenseASPP module (<xref ref-type="bibr" rid="B33">33</xref>) to replace the ASPP module, which has the network structure shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. DenseASPP further increases the denseness of the null convolution and expands the range of the network model, improving the ability of the network model to extract image feature information without significantly increasing the model size. Since using a null convolution with too large an expansion rate can lead to convolutional degradation and cause degradation in feature extraction performance, only null convolutional layers with expansion rates of 3, 6, 12, and 18 are used in this paper.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>DenseASPP network structure (<xref ref-type="bibr" rid="B32">32</xref>).</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g004.tif"/>
</fig>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Attention mechanism</title>
<p>When human beings observe things, their attention will focus on the areas of interest and ignore other areas. Through the attention mechanism, the input information is selectively distinguished, located, and analyzed. In the process of neural network learning, the attention mechanism will also be applied in the fields of image segmentation, target tracking, and behavior detection. We also can obtain the image feature information of different spaces and latitudes by giving different weight information to the input image, to improve the adaptability and effect of the neural network model (<xref ref-type="bibr" rid="B34">34</xref>, <xref ref-type="bibr" rid="B35">35</xref>). How to build an attention mechanism model and integrate it into the mainstream neural network structure, so that the simple neural network can achieve complex and high-precision image segmentation tasks, is one of the problems that need to be solved.</p>
<sec id="s2_4_1">
<label>2.4.1</label>
<title>Spatial attention</title>
<p>After analyzing the input image, the neural network assigns more weights to the regions which are closely related to the segmentation task, which makes the target segmentation region more prominent. At the same time, the image region feature information which has nothing to do with the segmentation task is suppressed. After the maximum and average pooling operations, the two pooling results are stitched together to achieve the input image feature information fusion. Then, the convolution kernel of 1&#xd7;1 is multiplied by the fused feature information, the sigmoid activation function is used for nonlinear transformation, and the spatial attention weight map is obtained. Finally, the input feature information of the image is multiplied by the spatial attention weight map to get the final output result, the formula is shown in Eq. (1) to Eq. (4).</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mtext>&#xa0;x</mml:mtext>
</mml:mrow>
<mml:mi>a</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>A</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>w</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>x<italic>
<sub>m</sub>
</italic> and x<italic>
<sub>a</sub>
</italic>, are obtained using maximum pooling and average pooling operations on the feature map. <italic>x<sub>input</sub>
</italic> represents the input feature maps. x<italic>
<sub>graphs</sub>
</italic> represents that the feature map obtained after nonlinear transformation of the fused image features with the sigmoid activation function.</p>
<p>x<italic>
<sub>outputs</sub>
</italic> represents the output of the input image multiplied with the spatial attention weight information.</p>
</sec>
<sec id="s2_4_2">
<label>2.4.2</label>
<title>Channel attention</title>
<p>In the neural network, the corresponding channel represents the image feature information. Convolution kernels of different scales process the input image features to generate different image feature channel information. The feature information on each channel is different, and the importance of the global feature information of the whole image is also different. Through the analysis of the segmented image, we can assign weight information to each channel information, indicating the importance of the channel information to the global feature description. For the input image feature information, firstly, the dimensionality reduction operation is carried out by maximum pooling and average pooling, and the two feature infor &#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;&#xa0;<italic>x</italic>
<sub>
<italic>m</italic>
</sub>=<italic>Maxpool</italic>(<italic>x</italic>
<sub>
<italic>input</italic>
</sub>) mation is input into a shared network structure for processing. The convolution kernel of 1&#xd7;1 is multiplied by the fused feature information, the Sigmoid activation function is used for nonlinear transformation, and the channel attention weight map is obtained. Finally, the input feature information of the image is multiplied by the channel attention weight map to get the final output result, the formula is shown in Eq. (5) to Eq. (10).</p>
<disp-formula>
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>A</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mtext>w</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>w</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>w</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mo>*</mml:mo>
<mml:mtext>x</mml:mtext>
</mml:mrow>
<mml:mtext>m</mml:mtext>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mrow>
<mml:mtext>fc</mml:mtext>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mi>c</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>d</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mi>f</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>*</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>x<italic>
<sub>m</sub>
</italic> and x<italic>
<sub>a</sub>
</italic>, are obtained using maximum pooling and average pooling operations on the feature map. <italic>x<sup>input</sup>
</italic> represents the input feature maps. x<italic>
<sub>graphs</sub>
</italic> represents that the feature map obtained after nonlinear transformation of the fused image features with the sigmoid activation function. <italic>w</italic>
<sub>
<italic>fc</italic>1</sub>&#x2208;<italic>R</italic>
<sup>
<italic>C</italic>/8&#xd7;1&#xd7;1</sup> , &#xa0;<italic>w</italic>
<sub>
<italic>fc</italic>2</sub>&#x2208;<italic>R</italic>
<sup>
<italic>C</italic>/8&#xd7;1&#xd7;1</sup> , &#xa0;<italic>w</italic>
<sub>
<italic>fc</italic>3</sub>&#x2208;<italic>R</italic>
<sup>
<italic>C</italic>&#xd7;1&#xd7;1</sup> .x<italic>
<sub>outputs</sub>
</italic> represents the output of the input image multiplied with the spatial attention weight information.</p>
</sec>
<sec id="s2_4_3">
<label>2.4.3</label>
<title>Spatial channel attention</title>
<p>In the paper, the channel attention and spatial attention mechanism are combined and given different weights. Finally, the output image feature information processed by the SCA module is as Eq. (11):</p>
<disp-formula>
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>x</mml:mtext>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>c</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>x<italic>
<sub>outputs</sub>
</italic> represents the output of the input image multiplied with the spatial attention weight information.</p>
<p>x<italic>
<sub>outputc</sub>
</italic> represents the output of the input image multiplied with the channel attention weight information.</p>
<p>Our proposed SCA module (<xref ref-type="bibr" rid="B36">36</xref>), with a general design idea similar to the architecture proposed by Fu et&#xa0;al. (<xref ref-type="bibr" rid="B37">37</xref>), integrates spatial and channel attention integration modules into an improved U-net network structure. The SCA module combines spatial and channel attention mechanisms to get comprehensive attention mechanism information. This module enhances the significant features of the up-sampling process by applying attention weights to high-dimensional and low-dimensional image feature information.</p>
<p>The feature map x&#x2208;R<sup>C&#xd7;H&#xd7;W</sup> as input, the attention weights &#xa0;M<sub>c</sub>&#xa0;(F)&#x2208;R<sup>C&#xd7;1&#xd7;1</sup>&#xa0; and weights M<sub>s</sub>&#xa0;(F)&#x2208;R<sup>1&#xd7;H&#xd7;W</sup>&#xa0; in the channel and space are obtained respectively through the SCA module. Finally, the results of the two modules are operated by concatenation, as shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>The SCA architecture.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g005.tif"/>
</fig>
<p>The SCA module proposed in the paper can address the feature information of medical images, highlight more of the key feature information of medical images, and suppress the interference of noise factors in medical images.</p>
</sec>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Loss function</title>
<p>The loss function is the function of a neural network to measure the degree of loss and error, and it is the index of a neural network to find the optimal weight parameters (<xref ref-type="bibr" rid="B38">38</xref>). Through the loss function, the difference between the segmentation result of the model and the actual result can be reflected. There are many kinds of loss functions, including single loss function and mixed loss function. The loss function adopted in this paper is the combined loss function of Dice loss (<xref ref-type="bibr" rid="B39">39</xref>) and Focal loss (<xref ref-type="bibr" rid="B40">40</xref>). This function can combine the advantages of the two functions to make the network better find the optimal parameters for optimization learning <italic>&#x2112;</italic>
<sub>Diceloss</sub> is the loss function of Dice loss. It is mainly used to measure the degree of loss of similarity between the segmented image predicted by the model and the real segmented image, and the value range is [0,1]. The calculation formula of the function is shown in Eq. (12). <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:msup>
<mml:mo>&#x2229;</mml:mo>
<mml:mo>&#x200b;</mml:mo>
</mml:msup>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> represents the number of intersections between a real segmented image and a model predicted image, |<italic>X</italic>| and |<italic>Y</italic>| represent the number of real segmented images and model predicted images respectively.</p>
<disp-formula>
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:msub>
<mml:mi>&#x2112;</mml:mi>
<mml:mrow>
<mml:mtext>Diceloss</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:msup>
<mml:mo>&#x2229;</mml:mo>
<mml:mo>&#x200b;</mml:mo>
</mml:msup>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<italic>&#x2112;</italic>
<sub>Focalloss</sub> is a loss function to deal with the unbalanced classification of samples. According to the difficulty of sample resolution, different weight coefficients &#x3b1; are added to the samples to reduce the adverse effects on training loss caused by the imbalance of sample classification. The calculation formula of the function is shown in Eq. (13).</p>
<disp-formula>
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:msub>
<mml:mi>&#x2112;</mml:mi>
<mml:mrow>
<mml:mtext>Focalloss</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>&#x3b1;</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>p</mml:mtext>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:msup>
<mml:mo>&#xd7;</mml:mo>
<mml:mtext>log</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>p&#x2208;[0,1] is the model&#x2019;s probability of predicting the positive sample. &#xa0;&#x3b1;&#x2208;[0,1] is used to balance the uneven proportions of positive and negative samples themselves. is used to adjust the rate of weight reduction for simple samples. <italic>&#x2112;</italic>
<sub>Focalloss</sub> is a modification of the cross entropy loss function. The total loss function proposed in this paper is <italic>&#x2112;</italic>
<sub>Total</sub> , the formula is shown in Eq. (14).</p>
<disp-formula>
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:msub>
<mml:mi>&#x2112;</mml:mi>
<mml:mrow>
<mml:mtext>Total</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>&#x2112;</mml:mi>
<mml:mrow>
<mml:mtext>Diceloss</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>&#x2112;</mml:mi>
<mml:mrow>
<mml:mtext>Focalloss</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Experimental results and analysis</title>
<sec id="s3_1">
<label>3.1</label>
<title>Datasets and preprocessing</title>
<p>To evaluate the model performance and verify the effectiveness of the algorithm, the model in this paper was subjected to relevant experiments on the Kvasir dataset (<xref ref-type="bibr" rid="B41">41</xref>), CVC-ClinicDB dataset (<xref ref-type="bibr" rid="B13">13</xref>), ETIS-Larib dataset (<xref ref-type="bibr" rid="B42">42</xref>), CVC-ColonDB (<xref ref-type="bibr" rid="B13">13</xref>) and CVC-300 (<xref ref-type="bibr" rid="B43">43</xref>) datasets.</p>
<p>We increase the number of training samples by the data augmentation method to improve the training effect of the model.</p>
<p>Firstly, the data of Kvasir and CVC-ClinicDB datasets are combined respectively, and then the training set, validation set and test set are divided according to the ratio of 80%, 10%, and 10%. CVC-ColonDB, ETIS-Larib, and CVC-300, are only used as the test set and do not participate in the dataset division. We expand the number of images in the training set by operations such as data enhancement, including center cropping, random rotation, Gaussian blurring etc. Among, center cropping size is (160,160), random rotation is 30 degrees, and Gaussian blurring kernel size is (3,3). A single image can be augmented into 20 different images, the original number of images in the Kvasir dataset is 800, and 16,000 after data enhancement; the original number of images in the CVC-ClinicDB dataset is 490, and 9,800 after data enhancement.</p>
<p>In the above datasets, polyps are highly variable in shape, size, structure, and orientation, and the boundaries between them and the background are very blurred and difficult to distinguish, which poses a great challenge to accurate polyp segmentation. The model resizes the resolution of colonoscopy images with different resolutions at the time of input and resizes all images uniformly to 256&#xd7;256 size.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Training and evaluation metrics</title>
<p>In the experimental part, the settings of all models are kept the same. The hardware device parameters of the algorithm running environment in this paper are: the processor is Intel i5-12400F, the graphics card is NVIDIA GeForce RTX2080T, and the deep learning framework is PyTorch1.6 framework. The network training process uses small batch training iterations, the batch size is set to 4, and the total number of training rounds is 300. The training is terminated early when the validation set accuracy no longer gets better for 50 consecutive rounds. The weighted sum of Dice loss and Focal loss function is used as the loss function, and the early stop method is triggered. We use the Adam optimization algorithm, and the learning rate set to 1e<sup>-4</sup>.</p>
<p>X is the set of pixels of the polyp region in the predicted segmentation result, and Y is the set of pixels of the gold standard polyp region in the original polyp image. Segmentation Results (SR) is the set of pixels predicted by the model for polyp segmentation, Ground Truth (GT) is the set of pixels for actual polyp segmentation, TP is the pixels correctly segmented in the polyp segmentation results, TN is the pixels incorrectly segmented in the polyp segmentation results, FP is the background pixels incorrectly treated as polyp pixels in the polyp segmentation results, and FN is the polyp pixels incorrectly treated as background pixels in the segmentation results. n is the number of images in the test set.</p>
<p>Several evaluation metrics of segmentation are used in the experiments, and the specific definitions of these metrics are given below.</p>
<p>Dice Similarity Coefficient (DSC): Calculates the similarity between the predicted target region and the actual target region. In this paper, the sum of the similarity coefficients of all test results in the test set is averaged and denoted as mDice. The relevant formula for the similarity coefficient is calculated as follows:</p>
<disp-formula>
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mtext>mDice</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>&#x2229;</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>|</mml:mo>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Intersection over Union (IoU): Calculates the ratio of the intersection of the two sets of predicted and actual values to the concurrent set. In this paper, the sum of the Intersection-over-Union coefficients of all the test results in the test set is averaged and denoted as mIoU. The relevant formula for the Intersection-over-Union coefficient is calculated as follows:</p>
<disp-formula>
<label>(16)</label>
<mml:math display="block" id="M16">
<mml:mrow>
<mml:mtext>mIou</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mtext>X</mml:mtext>
<mml:mo>&#x2229;</mml:mo>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>|</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mtext>X</mml:mtext>
<mml:mo>&#x2229;</mml:mo>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Sensitivity (sens): which indicates the ratio of the number of pixels correctly segmented to the number of all pixels in the image segmentation result, is calculated by the relevant formula as follows:</p>
<disp-formula>
<label>(17)</label>
<mml:math display="block" id="M17">
<mml:mrow>
<mml:mtext>mIou</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mtext>X</mml:mtext>
<mml:mo>&#x2229;</mml:mo>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>Y</mml:mi>
<mml:mo>|</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mtext>X</mml:mtext>
<mml:mo>&#x2229;</mml:mo>
<mml:mtext>Y</mml:mtext>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Specificity (spe): which indicates the ratio of the number of pixels correctly identified as incorrectly segmented to the number of all incorrectly segmented pixels in the segmentation result, is calculated by the relevant formula as follows:</p>
<disp-formula>
<label>(18)</label>
<mml:math display="block" id="M18">
<mml:mrow>
<mml:mtext>spe</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mtext>TN</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>TN</mml:mtext>
<mml:mo>+</mml:mo>
<mml:mtext>FP</mml:mtext>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The structural similarity measure Sm is used to assess the similarity between predicted and manually labeled real graphs (<xref ref-type="bibr" rid="B44">44</xref>).</p>
<disp-formula>
<label>(19)</label>
<mml:math display="block" id="M19">
<mml:mrow>
<mml:mtext>Sm</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mtext>&#x3b1;</mml:mtext>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mtext>S</mml:mtext>
<mml:mtext>O</mml:mtext>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>&#x3b1;</mml:mtext>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mtext>S</mml:mtext>
<mml:mtext>R</mml:mtext>
</mml:msub>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x3b1; is taken as 0.5; S<sub>R</sub> and S<sub>O</sub> are calculated by the structural similarity metrics in the field of image quality evaluation (<xref ref-type="bibr" rid="B44">44</xref>) denoting region-oriented and object-oriented structural similarity, respectively.</p>
<p>Mean Absolute Error (MAE): compare the pixel-by-pixel absolute difference between the predicted value y and the actual value y.</p>
<disp-formula>
<label>(20)</label>
<mml:math display="block" id="M20">
<mml:mrow>
<mml:mtext>MAE</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mtext>n</mml:mtext>
</mml:mfrac>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mi>&#xa0;</mml:mi>
<mml:mtext>&#xa0;</mml:mtext>
</mml:munderover>
<mml:munderover>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mtext>i</mml:mtext>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mtext>n</mml:mtext>
</mml:munderover>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mrow>
<mml:mtext>GT</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mtext>SR</mml:mtext>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Analysis of experimental results</title>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Comparative analysis of different algorithm models</title>
<p>To verify the performance and segmentation effect of the algorithm, the segmentation results on the Kvasir and CVC-ClinicDB polyp image datasets were compared with those of Unet (<xref ref-type="bibr" rid="B29">29</xref>), UNet++ (<xref ref-type="bibr" rid="B16">16</xref>), SFA (<xref ref-type="bibr" rid="B18">18</xref>), PraNet (<xref ref-type="bibr" rid="B25">25</xref>), UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>), and UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>) networks, respectively. <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> shows the final obtained methods of each of segmentation performance metrics, where the bolded values indicate the optimal metrics.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Comparison of segmentation effects of different methods on Kvasir and CVC-ClinicDB datasets.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Datasets</th>
<th valign="middle" align="center">Methods</th>
<th valign="middle" align="center">mDice</th>
<th valign="middle" align="center">mIou</th>
<th valign="middle" align="center">sens</th>
<th valign="middle" align="center">spe</th>
<th valign="middle" align="center">Sm</th>
<th valign="middle" align="center">MAE</th>
<th valign="top" align="center">Time</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="7" align="left">Kvasir</td>
<td valign="middle" align="left">UNet (<xref ref-type="bibr" rid="B29">29</xref>)</td>
<td valign="middle" align="center">0.818 &#xb1; 0.058</td>
<td valign="middle" align="center">0.746 &#xb1; 0.069</td>
<td valign="middle" align="center">0.887 &#xb1; 0.061</td>
<td valign="middle" align="center">0.941 &#xb1; 0.057</td>
<td valign="middle" align="center">0.858 &#xb1; 0.062</td>
<td valign="middle" align="center">0.055 &#xb1; 0.003</td>
<td valign="top" align="center">~8.3h</td>
</tr>
<tr>
<td valign="middle" align="left">UNet++ (<xref ref-type="bibr" rid="B16">16</xref>)</td>
<td valign="middle" align="center">0.821 &#xb1; 0.061</td>
<td valign="middle" align="center">0.763 &#xb1; 0.073</td>
<td valign="middle" align="center">0.900 &#xb1; 0.059</td>
<td valign="middle" align="center">0.956 &#xb1; 0.072</td>
<td valign="middle" align="center">0.862 &#xb1; 0.058</td>
<td valign="middle" align="center">0.048 &#xb1; 0.002</td>
<td valign="top" align="center">~9.2h</td>
</tr>
<tr>
<td valign="middle" align="left">PraNet (<xref ref-type="bibr" rid="B25">25</xref>)</td>
<td valign="middle" align="center">0.898 &#xb1; 0.072</td>
<td valign="middle" align="center">0.840 &#xb1; 0.120</td>
<td valign="middle" align="center">0.899 &#xb1; 0.056</td>
<td valign="middle" align="center">0.969 &#xb1; 0.047</td>
<td valign="middle" align="center">0.915 &#xb1; 0.082</td>
<td valign="middle" align="center">0.030 &#xb1; 0.003</td>
<td valign="top" align="center">~8.5h</td>
</tr>
<tr>
<td valign="middle" align="left">SFA (<xref ref-type="bibr" rid="B18">18</xref>)</td>
<td valign="middle" align="center">0.723 &#xb1; 0.124</td>
<td valign="middle" align="center">0.611 &#xb1; 0.097</td>
<td valign="middle" align="center">0.874 &#xb1; 0.106</td>
<td valign="middle" align="center">0.932 &#xb1; 0.053</td>
<td valign="middle" align="center">0.782 &#xb1; 0.069</td>
<td valign="middle" align="center">0.075 &#xb1; 0.004</td>
<td valign="top" align="center">~13.1h</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.912 &#xb1; 0.071</td>
<td valign="middle" align="center">0.859 &#xb1; 0.083</td>
<td valign="middle" align="center">0.907 &#xb1; 0.096</td>
<td valign="middle" align="center">0.958 &#xb1; 0.084</td>
<td valign="middle" align="center">0.917 &#xb1; 0.058</td>
<td valign="middle" align="center">0.025 &#xb1; 0.002</td>
<td valign="top" align="center">~11.4h</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.905 &#xb1; 0.069</td>
<td valign="middle" align="center">0.852 &#xb1; 0.086</td>
<td valign="middle" align="center">0.909 &#xb1; 0.078</td>
<td valign="middle" align="center">0.959 &#xb1; 0.063</td>
<td valign="middle" align="center">0.914 &#xb1; 0.083</td>
<td valign="middle" align="center">0.026 &#xb1; 0.002</td>
<td valign="top" align="center">~9.7h</td>
</tr>
<tr>
<td valign="middle" align="left">Ours</td>
<td valign="middle" align="center">
<bold>0.915 &#xb1; 0.085</bold>
</td>
<td valign="middle" align="center">
<bold>0.862 &#xb1; 0.120</bold>
</td>
<td valign="middle" align="center">
<bold>0.911 &#xb1; 0.086</bold>
</td>
<td valign="middle" align="center">
<bold>0.968 &#xb1; 0.081</bold>
</td>
<td valign="middle" align="center">
<bold>0.922 &#xb1; 0.074</bold>
</td>
<td valign="middle" align="center">
<bold>0.024 &#xb1; 0.001</bold>
</td>
<td valign="top" align="center">~7.5h</td>
</tr>
<tr>
<td valign="middle" rowspan="7" align="left">CVC-ClinicDB</td>
<td valign="middle" align="left">UNet (<xref ref-type="bibr" rid="B29">29</xref>)</td>
<td valign="middle" align="center">0.823 &#xb1; 0.654</td>
<td valign="middle" align="center">0.755 &#xb1; 0.073</td>
<td valign="middle" align="center">0.886 &#xb1; 0.061</td>
<td valign="middle" align="center">0.943 &#xb1; 0.059</td>
<td valign="middle" align="center">0.889 &#xb1; 0.072</td>
<td valign="middle" align="center">0.019 &#xb1; 0.003</td>
<td valign="top" align="center">~6.2h</td>
</tr>
<tr>
<td valign="middle" align="left">UNet++ (<xref ref-type="bibr" rid="B16">16</xref>)</td>
<td valign="middle" align="center">0.794 &#xb1; 0.058</td>
<td valign="middle" align="center">0.729 &#xb1; 0.075</td>
<td valign="middle" align="center">0.893 &#xb1; 0.082</td>
<td valign="middle" align="center">0.961 &#xb1; 0.079</td>
<td valign="middle" align="center">0.873 &#xb1; 0.083</td>
<td valign="middle" align="center">0.022 &#xb1; 0.002</td>
<td valign="top" align="center">~7.3h</td>
</tr>
<tr>
<td valign="middle" align="left">PraNet (<xref ref-type="bibr" rid="B25">25</xref>)</td>
<td valign="middle" align="center">0.899 &#xb1; 0.103</td>
<td valign="middle" align="center">0.849 &#xb1; 0.093</td>
<td valign="middle" align="center">0.935 &#xb1; 0.062</td>
<td valign="middle" align="center">0.974 &#xb1; 0.072</td>
<td valign="middle" align="center">0.936 &#xb1; 0.063</td>
<td valign="middle" align="center">0.009 &#xb1; 0.001</td>
<td valign="top" align="center">~6.5h</td>
</tr>
<tr>
<td valign="middle" align="left">SFA (<xref ref-type="bibr" rid="B18">18</xref>)</td>
<td valign="middle" align="center">0.700 &#xb1; 0.120</td>
<td valign="middle" align="center">0.607 &#xb1; 0.104</td>
<td valign="middle" align="center">0.879 &#xb1; 0.093</td>
<td valign="middle" align="center">0.937 &#xb1; 0.038</td>
<td valign="middle" align="center">0.793 &#xb1; 0.105</td>
<td valign="middle" align="center">0.042 &#xb1; 0.002</td>
<td valign="top" align="center">~11.1h</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.926 &#xb1; 0.073</td>
<td valign="middle" align="center">0.880 &#xb1; 0.071</td>
<td valign="middle" align="center">
<bold>0.941 &#xb1; 0.052</bold>
</td>
<td valign="middle" align="center">0.985 &#xb1; 0.013</td>
<td valign="middle" align="center">0.943 &#xb1; 0.032</td>
<td valign="middle" align="center">0.006 &#xb1; 0.003</td>
<td valign="top" align="center">~9.6h</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.916 &#xb1; 0.580</td>
<td valign="middle" align="center">0.870 &#xb1; 0.063</td>
<td valign="middle" align="center">0.927 &#xb1; 0.061</td>
<td valign="middle" align="center">
<bold>0.989 &#xb1; 0.011</bold>
</td>
<td valign="middle" align="center">0.940 &#xb1; 0.043</td>
<td valign="middle" align="center">0.008 &#xb1; 0.004</td>
<td valign="top" align="center">~7.3h</td>
</tr>
<tr>
<td valign="middle" align="left">Ours</td>
<td valign="middle" align="center">
<bold>0.931 &#xb1; 0.046</bold>
</td>
<td valign="middle" align="center">
<bold>0.892 &#xb1; 0.092</bold>
</td>
<td valign="middle" align="center">0.933 &#xb1; 0.047</td>
<td valign="middle" align="center">0.984 &#xb1; 0.092</td>
<td valign="middle" align="center">
<bold>0.945 &#xb1; 0.051</bold>
</td>
<td valign="middle" align="center">
<bold>0.005 &#xb1; 0.002</bold>
</td>
<td valign="top" align="center">~5.5h</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>From <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>, our proposed algorithm achieved good results on all six evaluation metrics on the Kvasir dataset, mDice, mIou, sens, spe, Sm and MAE were 0.915, 0.862, 0.911, 0.968, 0.922 and 0.024, respectively. Compared with the SFA network structure, mDice, mIou, sens, spe and Sm improve by 0.192, 0.251, 0.037, 0.036 and 0.140, respectively, and MAE decreases by 0.051. On the CVC-ClinicDB dataset, mDice, mIou, sens, spe, Sm and MAE were 0.931, 0.892, 0.933, 0.984, 0.945, and 0.005, respectively. Compared to the SFA network structure, mDice, mIou, sens, spe, and Sm improved by 0.231, 0.285, 0.054, 0.047, and 0.152, respectively, and MAE decreased by 0.037. The UACANet-L and UACANet-S network structures improve the metrics sens and spe metrics by 0.008 and 0.005, respectively. On the two datasets, the running time of the model in this paper is the least.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Generalization performance</title>
<p>To verify the generalization performance of the algorithm, the unseen datasets (CVC-ColonDB, CVC-300, ETIS-Larib) are used to test the generalization ability of the model (the model training data are only from Kvasir and CVC-ClinicDB). From <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>, we can found that the generalization ability of UNet, UNet++, and SFA is poor on the three datasets, especially the evaluation metrics of SFA decreases sharply, while the algorithm proposed in this paper shows superior levels of all indexes on the three test sets and achieves good results. Where the bolded values indicate the optimal metrics.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>Comparison results of different methods on an unseen dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="left">Datasets</th>
<th valign="middle" align="center">Methods</th>
<th valign="middle" align="center">mDice</th>
<th valign="middle" align="center">mIou</th>
<th valign="middle" align="center">sens</th>
<th valign="middle" align="center">spe</th>
<th valign="middle" align="center">Sm</th>
<th valign="middle" align="center">MAE</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="7" align="left">CVC-ColonDB</td>
<td valign="middle" align="left">UNet (<xref ref-type="bibr" rid="B29">29</xref>)</td>
<td valign="middle" align="center">0.512</td>
<td valign="middle" align="center">0.444</td>
<td valign="middle" align="center">0.754</td>
<td valign="middle" align="center">0.853</td>
<td valign="middle" align="center">0.712</td>
<td valign="middle" align="center">0.061</td>
</tr>
<tr>
<td valign="middle" align="left">UNet++ (<xref ref-type="bibr" rid="B16">16</xref>)</td>
<td valign="middle" align="center">0.483</td>
<td valign="middle" align="center">0.410</td>
<td valign="middle" align="center">0.735</td>
<td valign="middle" align="center">0.846</td>
<td valign="middle" align="center">0.691</td>
<td valign="middle" align="center">0.064</td>
</tr>
<tr>
<td valign="middle" align="left">PraNet (<xref ref-type="bibr" rid="B25">25</xref>)</td>
<td valign="middle" align="center">0.709</td>
<td valign="middle" align="center">0.640</td>
<td valign="middle" align="center">0.821</td>
<td valign="middle" align="center">0.914</td>
<td valign="middle" align="center">0.819</td>
<td valign="middle" align="center">0.045</td>
</tr>
<tr>
<td valign="middle" align="left">SFA (<xref ref-type="bibr" rid="B18">18</xref>)</td>
<td valign="middle" align="center">0.469</td>
<td valign="middle" align="center">0.347</td>
<td valign="middle" align="center">0.716</td>
<td valign="middle" align="center">0.842</td>
<td valign="middle" align="center">0.634</td>
<td valign="middle" align="center">0.094</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.751</td>
<td valign="middle" align="center">0.678</td>
<td valign="middle" align="center">0.837</td>
<td valign="middle" align="center">0.927</td>
<td valign="middle" align="center">0.835</td>
<td valign="middle" align="center">0.039</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.783</td>
<td valign="middle" align="center">0.704</td>
<td valign="middle" align="center">0.841</td>
<td valign="middle" align="center">0.936</td>
<td valign="middle" align="center">0.848</td>
<td valign="middle" align="center">0.034</td>
</tr>
<tr>
<td valign="middle" align="left">Ours</td>
<td valign="middle" align="center">
<bold>0.805</bold>
</td>
<td valign="middle" align="center">
<bold>0.722</bold>
</td>
<td valign="middle" align="center">
<bold>0.849</bold>
</td>
<td valign="middle" align="center">
<bold>0.938</bold>
</td>
<td valign="middle" align="center">
<bold>0.858</bold>
</td>
<td valign="middle" align="center">
<bold>0.031</bold>
</td>
</tr>
<tr>
<td valign="middle" rowspan="7" align="left">CVC-300</td>
<td valign="middle" align="left">UNet (<xref ref-type="bibr" rid="B29">29</xref>)</td>
<td valign="middle" align="center">0.710</td>
<td valign="middle" align="center">0.627</td>
<td valign="middle" align="center">0.897</td>
<td valign="middle" align="center">0.923</td>
<td valign="middle" align="center">0.843</td>
<td valign="middle" align="center">0.022</td>
</tr>
<tr>
<td valign="middle" align="left">UNet++ (<xref ref-type="bibr" rid="B16">16</xref>)</td>
<td valign="middle" align="center">0.707</td>
<td valign="middle" align="center">0.624</td>
<td valign="middle" align="center">0.915</td>
<td valign="middle" align="center">0.931</td>
<td valign="middle" align="center">0.839</td>
<td valign="middle" align="center">0.018</td>
</tr>
<tr>
<td valign="middle" align="left">PraNet (<xref ref-type="bibr" rid="B25">25</xref>)</td>
<td valign="middle" align="center">0.871</td>
<td valign="middle" align="center">0.797</td>
<td valign="middle" align="center">0.938</td>
<td valign="middle" align="center">0.966</td>
<td valign="middle" align="center">0.925</td>
<td valign="middle" align="center">0.010</td>
</tr>
<tr>
<td valign="middle" align="left">SFA (<xref ref-type="bibr" rid="B18">18</xref>)</td>
<td valign="middle" align="center">0.467</td>
<td valign="middle" align="center">0.329</td>
<td valign="middle" align="center">0.887</td>
<td valign="middle" align="center">0.925</td>
<td valign="middle" align="center">0.640</td>
<td valign="middle" align="center">0.065</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.910</td>
<td valign="middle" align="center">0.849</td>
<td valign="middle" align="center">0.928</td>
<td valign="middle" align="center">0.970</td>
<td valign="middle" align="center">0.937</td>
<td valign="middle" align="center">0.005</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.902</td>
<td valign="middle" align="center">0.837</td>
<td valign="middle" align="center">0.931</td>
<td valign="middle" align="center">0.975</td>
<td valign="middle" align="center">0.934</td>
<td valign="middle" align="center">0.006</td>
</tr>
<tr>
<td valign="middle" align="left">Ours</td>
<td valign="middle" align="center">
<bold>0.923</bold>
</td>
<td valign="middle" align="center">
<bold>0.857</bold>
</td>
<td valign="middle" align="center">
<bold>0.945</bold>
</td>
<td valign="middle" align="center">
<bold>0.981</bold>
</td>
<td valign="middle" align="center">
<bold>0.946</bold>
</td>
<td valign="middle" align="center">
<bold>0.003</bold>
</td>
</tr>
<tr>
<td valign="middle" rowspan="7" align="left">ETIS-Larib</td>
<td valign="middle" align="left">UNet (<xref ref-type="bibr" rid="B29">29</xref>)</td>
<td valign="middle" align="center">0.398</td>
<td valign="middle" align="center">0.335</td>
<td valign="middle" align="center">0.673</td>
<td valign="middle" align="center">0.782</td>
<td valign="middle" align="center">0.684</td>
<td valign="middle" align="center">0.036</td>
</tr>
<tr>
<td valign="middle" align="left">UNet++ (<xref ref-type="bibr" rid="B16">16</xref>)</td>
<td valign="middle" align="center">0.401</td>
<td valign="middle" align="center">0.344</td>
<td valign="middle" align="center">0.687</td>
<td valign="middle" align="center">0.779</td>
<td valign="middle" align="center">0.683</td>
<td valign="middle" align="center">0.035</td>
</tr>
<tr>
<td valign="middle" align="left">PraNet (<xref ref-type="bibr" rid="B25">25</xref>)</td>
<td valign="middle" align="center">0.628</td>
<td valign="middle" align="center">0.567</td>
<td valign="middle" align="center">0.765</td>
<td valign="middle" align="center">0.807</td>
<td valign="middle" align="center">0.794</td>
<td valign="middle" align="center">0.031</td>
</tr>
<tr>
<td valign="middle" align="left">SFA (<xref ref-type="bibr" rid="B18">18</xref>)</td>
<td valign="middle" align="center">0.297</td>
<td valign="middle" align="center">0.217</td>
<td valign="middle" align="center">0.524</td>
<td valign="middle" align="center">0.723</td>
<td valign="middle" align="center">0.557</td>
<td valign="middle" align="center">0.109</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-L (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.766</td>
<td valign="middle" align="center">0.689</td>
<td valign="middle" align="center">0.797</td>
<td valign="middle" align="center">0.827</td>
<td valign="middle" align="center">0.859</td>
<td valign="middle" align="center">0.012</td>
</tr>
<tr>
<td valign="middle" align="left">UACANet-S (<xref ref-type="bibr" rid="B45">45</xref>)</td>
<td valign="middle" align="center">0.694</td>
<td valign="middle" align="center">0.615</td>
<td valign="middle" align="center">0.801</td>
<td valign="middle" align="center">0.789</td>
<td valign="middle" align="center">0.815</td>
<td valign="middle" align="center">0.023</td>
</tr>
<tr>
<td valign="middle" align="left">Ours</td>
<td valign="middle" align="center">
<bold>0.774</bold>
</td>
<td valign="middle" align="center">
<bold>0.691</bold>
</td>
<td valign="middle" align="center">
<bold>0.812</bold>
</td>
<td valign="middle" align="center">
<bold>0.834</bold>
</td>
<td valign="middle" align="center">
<bold>0.864</bold>
</td>
<td valign="middle" align="center">
<bold>0.009</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Results visualization analysis</title>
<p>The model can clearly distinguish polyps from other tissues in the polyp segmentation task, keep the polyp edge segmentation intact while reducing the miss segmentation inside the polyp, and the segmentation results are shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, and the red areas in the figure are the pixel points missed segmentation by the network. Some of them are segmentation networks that incorrectly mark background pixels as polyps; some of them are segmentation networks that incorrectly mark polyps pixels as background pixels.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Polyp segmentation results.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fonc-12-1087438-g006.tif"/>
</fig>
<p>From <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, we can find that the segmentation model proposed in this paper can achieve better segmentation results. However, for images with smooth polyp edges that are easy to distinguish, such as the third row in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, all models are able to segment accurately. When the edges of the lesion region are similar to the background, all segmentation models have certain challenges and are prone to multi-segmentation or omission of pixel points. From the segmentation results in the first and second rows of <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>, it is found that our models have poor segmentation results; however, for some polyps images with complex backgrounds, our models can basically distinguish the lesion regions with blurred borders completely, as shown in the fourth and fifth rows of <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref>.</p>
<p>From the visualized results, we can conclude that our proposed model can generally overcome the problem of polyps with similar colors and backgrounds well, and detect polyps with different shapes and sizes and colors of tissues, and the delineated areas and boundaries are more clear and accurate.</p>
</sec>
</sec>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>We propose a network model for polyp image segmentation based on deep neural network techniques by combining HarDNet and a multiscale coding module of attention mechanism and perform comparison and generalization experiments on five polyp image datasets. The <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> shows the segmentation results of UNet, UNet++, SFA, PraNet, UACANet-L, and UACANet-L networks on the Kvasir and CVCClinicDB polyp image datasets. <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref> shows the networks such as UNet, UNet++, SFA, PraNet, UACANet-L, and UACANet-L tested on the Kvasir and CVCClinicDB network models on the CVC-ColonDB, CVC-300, and ETIS-Larib datasets to verify the generalization performance of the models. <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref> shows the segmentation results of each model on the Kvasir, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS-Larib polyp image datasets, respectively.</p>
<p>For the data analysis in <xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="table" rid="T3">
<bold>3</bold>
</xref>, and <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6</bold>
</xref> above, we found that the polyp image segmentation model based on deep learning proposed in this paper can well overcome the problem of similar a color of polyps and backgrounds, detect polyp tissues with different shapes and sizes and colors, and achieve excellent results with clearer and more accurate delineation of regions and boundaries.</p>
<p>Through the analysis of the above results, we know that the U-Net network causes the semantic gap phenomenon due to the information difference between the output features in the encoding and decoding stages, which affects the segmentation results. Our proposed polyp segmentation network is able to capture the path of contextual information and effectively reduce the semantic discrepancy, which well overcomes the diversity of polyp shape, size, color and texture as well as the unclear boundary between the polyp and its surrounding mucosa to achieve more accurate segmentation results.</p>
<p>Our proposed polyp image segmentation network has achieved a certain degree of improvement, but the model still has room for further enhancement.</p>
<p>First, to improve the training speed of the network, we used parameters and weights based on ImageNet pre-training, but there are huge differences in features, textures, and other information between ordinary images and polyp images, which impose certain limitations on the effectiveness of medical image segmentation. Secondly, our network has not been attempted to be validated on a 3D medical image segmentation dataset. Finally, the size of polyp images in the polyp image dataset studied in this paper varies greatly, but we set a uniform size of 256&#xd7;256 during image preprocessing and did not validate the effect of image size on model segmentation accuracy. In addition, we just merged the Kvasir and CVC-ClinicDB datasets and could not guarantee the independence between the subsets. We will further investigate these issues in our future work.</p>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusions</title>
<p>In the paper, a multi-scale coded colon polyp image segmentation network combining HarDNet and attention mechanism was proposed for automatic polyp segmentation of colonoscopy images, reducing the effects of orientation, shape, texture, and size on the results. The proposed segmentation network was evaluated on five polyp image datasets, just as the Kvasir, CVC-ClinicDB, CVC-ColonDB, CVC-300, and ETIS-Larib, and analyzed and compared with other existing representative methods. Through comparative analysis with other models, we can found that the accuracy of the segmentation algorithm proposed in this paper is better than other methods, and for images with very low contrast between polyps and surrounding mucosa. Through experimental comparison and analysis, the segmentation algorithm proposed has better accuracy than other methods. It can accurately segment the boundary of polyps even for images with very low contrast between polyps and surrounding mucosa. There are no image artifacts outside the boundary with good image coherence. The polyp segmentation network proposed has excellent performance and good generalization ability, which can assist physicians in the diagnosis of colorectal polyps and reduce the leakage and misdiagnosis in clinical time, and is of reference for the processing and analysis of colorectal polyp images.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>TS did the conceptualisation, analysis and writing. XL provided conceptualisation, aided analysis and revised writing. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="funding-information">
<title>Funding</title>
<p>This research was funded by Excellent Young Talents in Anhui Universities Project (Granted No. gxyq2022026), Anhui Province Quality Engineering Project (Granted No. 2021jyxm0801), and Natural Science Foundation of Anhui University of Chinese Medicine (Granted No. 2020zrzd18).</p>
</sec>
<sec id="s9" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sung</surname> <given-names>H</given-names>
</name>
<name>
<surname>Ferlay</surname> <given-names>J</given-names>
</name>
<name>
<surname>Siegel</surname> <given-names>RL</given-names>
</name>
<name>
<surname>Laversanne</surname> <given-names>M</given-names>
</name>
<name>
<surname>Soerjomataram</surname> <given-names>I</given-names>
</name>
<name>
<surname>Jemal</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries</article-title>. <source>Ca: Cancer J Clin</source> (<year>2021</year>) <volume>171</volume>:<page-range>209&#x2013;49</page-range>. doi: <pub-id pub-id-type="doi">10.3322/caac.21660</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lieberman</surname> <given-names>DA</given-names>
</name>
<name>
<surname>Rex</surname> <given-names>DK</given-names>
</name>
<name>
<surname>Winawer</surname> <given-names>SJ</given-names>
</name>
<name>
<surname>Giardiello</surname> <given-names>FM</given-names>
</name>
<name>
<surname>Johnson</surname> <given-names>DA</given-names>
</name>
<name>
<surname>Levin</surname> <given-names>TR</given-names>
</name>
</person-group>. <article-title>Guidelines for colonoscopy surveillance after screening and polypectomy: A consensus update by the US multi-society task force on colorectal cancer</article-title>. <source>Gastroenterology</source> (<year>2012</year>) <volume>143</volume>(<issue>3</issue>):<page-range>844&#x2013;57</page-range>. doi: <pub-id pub-id-type="doi">10.1053/j.gastro.2012.06.001</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Corley</surname> <given-names>DA</given-names>
</name>
<name>
<surname>Jensen</surname> <given-names>CD</given-names>
</name>
<name>
<surname>Marks</surname> <given-names>AR</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>WK</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>JK</given-names>
</name>
<name>
<surname>Doubeni</surname> <given-names>CA</given-names>
</name>
<etal/>
</person-group>. <article-title>Adenoma detection rate and risk of colorectal cancer and death</article-title>. <source>New Engl J Med</source> (<year>2014</year>) <volume>370</volume>:<page-range>1298&#x2013;306</page-range>. doi: <pub-id pub-id-type="doi">10.1056/NEJMoa1309086</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jia</surname> <given-names>X</given-names>
</name>
<name>
<surname>Xing</surname> <given-names>X</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Xing</surname> <given-names>L</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>MQ</given-names>
</name>
</person-group>. (<year>2019</year>). <article-title>Wireless capsule endoscopy: A new tool for cancer screening in the colon with deep-learning-based polyp recognition</article-title>. <source>Proceedings of the IEEE</source>,. <publisher-loc>United States</publisher-loc>, <publisher-loc>IEEE</publisher-loc>,.Vol. <volume>108</volume>. pp. <fpage>178</fpage>&#x2013;<lpage>197</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JPROC.2019.2950506</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saul</surname> <given-names>C</given-names>
</name>
<name>
<surname>Canali</surname> <given-names>C</given-names>
</name>
<name>
<surname>Teixeira</surname> <given-names>CR</given-names>
</name>
<name>
<surname>Parada</surname> <given-names>AA</given-names>
</name>
<name>
<surname>Prolla</surname> <given-names>JC</given-names>
</name>
<name>
<surname>da Silva</surname> <given-names>VD</given-names>
</name>
</person-group>. <article-title>Digital morphometric characterization of mucosal surface lesion patterns under magnification colonoscopy</article-title>. <source>Analytical Quantitative Cytology Histol</source> (<year>2009</year>) <volume>31</volume>:<page-range>375&#x2013;9</page-range>.</citation>
</ref>
<ref id="B6">
<label>6</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bashar</surname> <given-names>MK</given-names>
</name>
<name>
<surname>Kitasaka</surname> <given-names>T</given-names>
</name>
<name>
<surname>Suenaga</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Mekada</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Mori</surname> <given-names>K</given-names>
</name>
</person-group>. <article-title>Automatic detection of informative frames from wireless capsule endoscopy images</article-title>. <source>Med Image Anal</source> (<year>2010</year>) <volume>14</volume>:<page-range>449&#x2013;70</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.media.2009.12.001</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Segui</surname> <given-names>S</given-names>
</name>
<name>
<surname>Drozdzal</surname> <given-names>M</given-names>
</name>
<name>
<surname>Vilarino</surname> <given-names>F</given-names>
</name>
<name>
<surname>Malagelada</surname> <given-names>C</given-names>
</name>
<name>
<surname>Azpiroz</surname> <given-names>F</given-names>
</name>
<name>
<surname>Radeva</surname> <given-names>P</given-names>
</name>
<etal/>
</person-group>. <article-title>Categorization and segmentation of intestinal content frames for wireless capsule endoscopy</article-title>. <source>IEEE Trans Inf Technol Biomedicine</source> (<year>2012</year>) <volume>16</volume>:<page-range>1341&#x2013;52</page-range>. doi: <pub-id pub-id-type="doi">10.1109/TITB.2012.2221472</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Turcza</surname> <given-names>P</given-names>
</name>
<name>
<surname>Duplaga.</surname> <given-names>M</given-names>
</name>
</person-group>. <article-title>Hardware-efficient low-power image processing system for wireless capsule endoscopy</article-title>. <source>IEEE J Biomed Health Inf</source> (<year>2013</year>) <volume>13</volume>:<page-range>1046&#x2013;56</page-range>. doi: <pub-id pub-id-type="doi">10.1109/JBHI.2013.2266101</pub-id>
</citation>
</ref>
<ref id="B9">
<label>9</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hassan</surname> <given-names>AR</given-names>
</name>
<name>
<surname>Haque</surname> <given-names>MA</given-names>
</name>
</person-group>. <article-title>Computer-aided gastrointestinal hemorrhage detection in wireless capsule endoscopy videos</article-title>. <source>Comput Methods Programs Biomedicine</source> (<year>2015</year>) <volume>122</volume>:<page-range>341&#x2013;53</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2015.09.005</pub-id>
</citation>
</ref>
<ref id="B10">
<label>10</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qiao</surname> <given-names>P</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>H</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>X</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Pi</surname> <given-names>X</given-names>
</name>
</person-group>. <article-title>A smart capsule system for automated detection of intestinal bleeding using hsl color recognition</article-title>. <source>PloS One</source> (<year>2016</year>) <volume>11</volume>:<fpage>14</fpage>. doi: <pub-id pub-id-type="doi">10.1371/journal.pone.0166488</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mamonov</surname> <given-names>AV</given-names>
</name>
<name>
<surname>Figueiredo</surname> <given-names>IN</given-names>
</name>
<name>
<surname>Figueiredo</surname> <given-names>PN</given-names>
</name>
<name>
<surname>Tsai</surname> <given-names>YH</given-names>
</name>
</person-group>. <article-title>Automated polyp detection in colon capsule endoscopy</article-title>. <source>IEEE Trans Med Imaging</source> (<year>2014</year>) <volume>33</volume>:<page-range>1488&#x2013;502</page-range>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2014.2314959</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tajbakhsh</surname> <given-names>N</given-names>
</name>
<name>
<surname>Gurudu</surname> <given-names>S</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>Automated polyp detection in colonoscopy videos using shape and context information</article-title>. <source>IEEE Trans Med Imaging</source> (<year>2015</year>) <volume>35</volume>:<page-range>630&#x2013;44</page-range>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2015.2487997</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernal</surname> <given-names>J</given-names>
</name>
<name>
<surname>S&#xe1;nchez</surname> <given-names>FJ</given-names>
</name>
<name>
<surname>Fern&#xe1;ndez-Esparrach</surname> <given-names>G</given-names>
</name>
<name>
<surname>Gil</surname> <given-names>D</given-names>
</name>
<name>
<surname>Rodr&#xed;guez</surname> <given-names>C</given-names>
</name>
<name>
<surname>Vilari&#xf1;o</surname> <given-names>F</given-names>
</name>
</person-group>. <article-title>WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians</article-title>. <source>Computerized Med Imaging Graphics</source> (<year>2015</year>) <volume>43</volume>:<fpage>99</fpage>&#x2013;<lpage>111</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.compmedimag.2015.02.007</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Brandao</surname> <given-names>P</given-names>
</name>
<name>
<surname>Mazomenos</surname> <given-names>E</given-names>
</name>
<name>
<surname>Ciuti</surname> <given-names>G</given-names>
</name>
<name>
<surname>Cali&#xf2;</surname> <given-names>R</given-names>
</name>
<name>
<surname>Bianchi</surname> <given-names>F</given-names>
</name>
<name>
<surname>Menciassi</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>Fully convolutional neural networks for polyp segmentation in colonoscopy</article-title>. In: <source>Medical imaging 2017:Computer-aided diagnosis</source>. <publisher-loc>Orlando, Florida, United States</publisher-loc>: <publisher-name>SPIE</publisher-name> (<year>2017</year>). p. <page-range>101&#x2013;7</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1117/12.2254361</pub-id>
</citation>
</ref>
<ref id="B15">
<label>15</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>L</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>Y</given-names>
</name>
</person-group>. <article-title>IDDF2018-ABS-0259 segmentation of intestinal polyps via a deep learning algorithm</article-title>. <source>Int Digestive Dis Forum (IDDF)</source> (<year>2018</year>), <page-range>. 83&#x2013;84</page-range>. doi: <pub-id pub-id-type="doi">10.1136/gutjnl-2018-IDDFabstracts.180</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Siddiquee</surname> <given-names>MMR</given-names>
</name>
<name>
<surname>Tajbakhsh</surname> <given-names>N</given-names>
</name>
<name>
<surname>Liang</surname> <given-names>JM</given-names>
</name>
</person-group>. <article-title>UNet++: A nested U-net architecture for medical image segmentation</article-title>. In: <source>Deep learning in medical image analysis and multimodal learning for clinical decision support</source>. <publisher-loc>Granada, Spain</publisher-loc>: <publisher-name>Springer</publisher-name>. pp.<fpage>3</fpage>&#x2013;<lpage>11</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-00889-5_1</pub-id>
</citation>
</ref>
<ref id="B17">
<label>17</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jha</surname> <given-names>D</given-names>
</name>
<name>
<surname>Smedsrud</surname> <given-names>PH</given-names>
</name>
<name>
<surname>Riegler</surname> <given-names>MA</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>D</given-names>
</name>
<name>
<surname>De Lange</surname> <given-names>T</given-names>
</name>
<name>
<surname>Halvorsen</surname> <given-names>P</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Resunet++: An advanced architecture for medical image segmentation</article-title>, in: <conf-name>2019 IEEE International Symposium on Multimedia (ISM)</conf-name>. <publisher-loc>San Diego, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <fpage>225</fpage>&#x2013;<lpage>2255</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ISM46123.2019.00049</pub-id>
</citation>
</ref>
<ref id="B18">
<label>18</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fang</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>C</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Tong</surname> <given-names>KY</given-names>
</name>
</person-group>. (<year>2019</year>). <article-title>Selective feature aggregation network with area-boundary constraints for polyp segmentation</article-title>, in: <conf-name>International Conference on Medical Image Computing and Computer-Assisted Intervention</conf-name>. <publisher-loc>Shenzhen, China</publisher-loc>: <publisher-name>Springer</publisher-name>. pp. <page-range>302&#x2013;10</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-32239-7_34</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jha</surname> <given-names>D</given-names>
</name>
<name>
<surname>Riegler</surname> <given-names>MA</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>D</given-names>
</name>
<name>
<surname>Halvorsen</surname> <given-names>P</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>HD</given-names>
</name>
</person-group>. (<year>2020</year>). <article-title>Doubleu-net: A deep convolutional neural network for medical image segmentation</article-title>, in: <conf-name>2020 IEEE 33rd International symposium on computer based medical systems (CBMS)</conf-name>. <publisher-loc>Rochester, MN, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>558&#x2013;64</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CBMS49503.2020.00111</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nadimi</surname> <given-names>ES</given-names>
</name>
<name>
<surname>Buijs</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Herp</surname> <given-names>J</given-names>
</name>
<name>
<surname>Kroijer</surname> <given-names>R</given-names>
</name>
<name>
<surname>Kobaek-Larsen</surname> <given-names>M</given-names>
</name>
<name>
<surname>Nielsen</surname> <given-names>E</given-names>
</name>
<etal/>
</person-group>. <article-title>Application of deep learning for autonomous detection and localization of colorectal polyps in wireless colon capsule endoscopy</article-title>. <source>Comput Electrical Eng</source> (<year>2020</year>) <volume>81</volume>:<fpage>1</fpage>&#x2013;<lpage>16</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.compeleceng.2019.106531</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Owais</surname> <given-names>M</given-names>
</name>
<name>
<surname>Arsalan</surname> <given-names>M</given-names>
</name>
<name>
<surname>Mahmood</surname> <given-names>T</given-names>
</name>
<name>
<surname>Kang</surname> <given-names>JK</given-names>
</name>
<name>
<surname>Park</surname> <given-names>KR</given-names>
</name>
</person-group>. <article-title>Automated diagnosis of various gastrointestinal lesions using a deep learning-based classification and retrieval framework with a large endoscopic database: Model development and validation</article-title>. <source>J Med Internet Res</source> (<year>2020</year>) <volume>22</volume>:<fpage>1</fpage>&#x2013;<lpage>21</lpage>. doi: <pub-id pub-id-type="doi">10.2196/18563</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yamada</surname> <given-names>A</given-names>
</name>
<name>
<surname>Niikura</surname> <given-names>R</given-names>
</name>
<name>
<surname>Otani</surname> <given-names>K</given-names>
</name>
<name>
<surname>Aoki</surname> <given-names>T</given-names>
</name>
<name>
<surname>Koike</surname> <given-names>K</given-names>
</name>
</person-group>. <article-title>Automatic detection of colorectal neoplasia in wireless colon capsule endoscopic images using a deep convolutional neural network</article-title>. <source>Endoscopy</source> (<year>2020</year>) <volume>17</volume>:<fpage>1</fpage>&#x2013;<lpage>19</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1055/a-1266-1066</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>SA</given-names>
</name>
<name>
<surname>Cho</surname> <given-names>HC</given-names>
</name>
</person-group>. <article-title>A novel approach for increased convolutional neural network performance in gastric-cancer classification using endoscopic images</article-title>. <source>IEEE Access</source> (<year>2021</year>) <volume>9</volume>:<page-range>51847&#x2013;54</page-range>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3069747</pub-id>
</citation>
</ref>
<ref id="B24">
<label>24</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lai</surname> <given-names>LL</given-names>
</name>
<name>
<surname>Blakely</surname> <given-names>A</given-names>
</name>
<name>
<surname>Invernizzi</surname> <given-names>M</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>J</given-names>
</name>
<name>
<surname>Kidambi</surname> <given-names>T</given-names>
</name>
<name>
<surname>Melstrom</surname> <given-names>KA</given-names>
</name>
<etal/>
</person-group>. <article-title>Separation of color channels from conventional colonoscopy images improves deep neural network detection of polyps</article-title>. <source>J Biomed Optics</source> (<year>2021</year>) <volume>26</volume>:<fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi: <pub-id pub-id-type="doi">10.1117/1.JBO.26.1.015001</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fan</surname> <given-names>DP</given-names>
</name>
<name>
<surname>Ji</surname> <given-names>GP</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>T</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>G</given-names>
</name>
<name>
<surname>Fu</surname> <given-names>H</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>J</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Pranet: Parallel reverse attention network for polyp segmentation</article-title>, in: <conf-name>International conference on medical image computing and computer-assisted intervention</conf-name>. <publisher-loc>Lima, Peru</publisher-loc>: <publisher-name>Springer</publisher-name>. pp. <page-range>263&#x2013;73</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-59725-2_26</pub-id>
</citation>
</ref>
<ref id="B26">
<label>26</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>X</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>H</given-names>
</name>
</person-group>. (<year>2021</year>). <article-title>Automatic polyp segmentation via multi-scale subtraction network</article-title>, in: <conf-name>International Conference on Medical Image Computing and Computer-Assisted Intervention</conf-name>. <conf-loc>Strasbourg, France</conf-loc>: <publisher-name>Springer</publisher-name>. pp. <page-range>120&#x2013;30</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-030-87193-2_12</pub-id>
</citation>
</ref>
<ref id="B27">
<label>27</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Chao</surname> <given-names>P</given-names>
</name>
<name>
<surname>Kao</surname> <given-names>CY</given-names>
</name>
<name>
<surname>Ruan</surname> <given-names>YS</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>CH</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>YL</given-names>
</name>
</person-group>. (<year>2019</year>). <article-title>Hardnet: A low memory traffic network</article-title>, in: <conf-name>Proceedings of the IEEE/CVF international conference on computer vision</conf-name>. <publisher-loc>Seoul, Korea</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>3552&#x2013;61</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV.2019.00365</pub-id>
</citation>
</ref>
<ref id="B28">
<label>28</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>CH</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>HY</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>YL</given-names>
</name>
</person-group>. <article-title>Hardnet-mseg: A simple encoder-decoder polyp segmentation neural network that achieves over 0.9 mean dice and 86 fps</article-title>. <source>arXiv preprint ar</source> (<year>2021</year>) <volume>Xiv</volume>:<fpage>2101.07172</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.48550/arXiv.2101.07172</pub-id>
</citation>
</ref>
<ref id="B29">
<label>29</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Srivastava</surname> <given-names>A</given-names>
</name>
<name>
<surname>Jha</surname> <given-names>D</given-names>
</name>
<name>
<surname>Chanda</surname> <given-names>S</given-names>
</name>
<name>
<surname>Pal</surname> <given-names>U</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>HD</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>D</given-names>
</name>
<etal/>
</person-group>. <article-title>Msrf-net: A multi-scale residual fusion network for biomedical image segmentation</article-title>. <source>IEEE J Biomed Health Inf</source> (<year>2021</year>) <volume>26</volume>(<issue>5</issue>):<page-range>2252&#x2013;63</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/JBHI.2021.3138024</pub-id>
</citation>
</ref>
<ref id="B30">
<label>30</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Ronneberger</surname> <given-names>O</given-names>
</name>
<name>
<surname>Fischer</surname> <given-names>P</given-names>
</name>
<name>
<surname>Brox</surname> <given-names>T</given-names>
</name>
</person-group>. (<year>2015</year>). <article-title>U-net: Convolutional networks for biomedical image segmentation</article-title>, in: <conf-name>International Conference on Medical image computing and computer-assisted intervention</conf-name>. <publisher-loc>Munich, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>. pp. <page-range>234&#x2013;41</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id>
</citation>
</ref>
<ref id="B31">
<label>31</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>S</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
</person-group>. <article-title>Spatial pyramid pooling in deep convolutional networks for visual recognition</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source> (<year>2014</year>) <volume>37</volume>:<page-range>1904&#x2013;16</page-range>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2015.2389824</pub-id>
</citation>
</ref>
<ref id="B32">
<label>32</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>LC</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Papandreou</surname> <given-names>G</given-names>
</name>
<name>
<surname>Schroff</surname> <given-names>F</given-names>
</name>
<name>
<surname>Adam</surname> <given-names>H</given-names>
</name>
</person-group>. <article-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</article-title>. <source>Comput Vision-ECCV 2018</source> (<year>2018</year>), <page-range>801&#x2013;18</page-range>. doi: <pub-id pub-id-type="doi">10.1007/978-3-030-01234-2_49</pub-id>
</citation>
</ref>
<ref id="B33">
<label>33</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Yang</surname> <given-names>M</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>C</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>K</given-names>
</name>
</person-group>. (<year>2018</year>). <article-title>DenseASPP for semantic segmentation in street scenes</article-title>, in: <conf-name>2018 IEEE/CVF Conference on ComputerVision and Pattern Recognition</conf-name>. <publisher-loc>Salt Lake City, UT, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>3684&#x2013;92</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2018.00388</pub-id>
</citation>
</ref>
<ref id="B34">
<label>34</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>QH</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>YH</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>YZ</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>FN</given-names>
</name>
<name>
<surname>Li</surname> <given-names>XL</given-names>
</name>
</person-group>. <article-title>Segmentation of breast ultrasound image with semantic classification of superpixels</article-title>. <source>Med Image Anal</source> (<year>2020</year>) <volume>61</volume>:<fpage>101657</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.media.2020.101657</pub-id>
</citation>
</ref>
<ref id="B35">
<label>35</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname> <given-names>QH</given-names>
</name>
<name>
<surname>Miao</surname> <given-names>ZJ</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>SC</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>C</given-names>
</name>
<name>
<surname>Li</surname> <given-names>XL</given-names>
</name>
</person-group>. <article-title>Dense prediction and local fusion of superpixels: A framework for breast anatomy segmentation in ultrasound image with scarce data</article-title>. <source>IEEE Trans Instrumentation Measurement</source> (<year>2021</year>) <volume>70</volume>:<fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIM.2021.3088421</pub-id>
</citation>
</ref>
<ref id="B36">
<label>36</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shen</surname> <given-names>TP</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>HQ</given-names>
</name>
</person-group>. <article-title>Facial expression recognition based on multi-channel attention residual network</article-title>. <source>CMES-Computer Modeling Eng Sci</source> (<year>2023</year>) <volume>135</volume>:<page-range>539&#x2013;60</page-range>. doi: <pub-id pub-id-type="doi">10.32604/cmes.2022.022312</pub-id>
</citation>
</ref>
<ref id="B37">
<label>37</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fu</surname> <given-names>J</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>J</given-names>
</name>
<name>
<surname>Tian</surname> <given-names>H</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Bao</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Fang</surname> <given-names>Z</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Dual attention network for scene segmentation</article-title>, in: <conf-name>Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</conf-name>. <publisher-loc>Long Beach, Canada</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>3146&#x2013;54</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/CVPR.2019.00326</pub-id>
</citation>
</ref>
<ref id="B38">
<label>38</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jun</surname> <given-names>M</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>JN</given-names>
</name>
<name>
<surname>Ng</surname> <given-names>M</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>R</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Li</surname> <given-names>C</given-names>
</name>
<etal/>
</person-group>. <article-title>Loss odyssey in medical image segmentation</article-title>. <source>Med Image Anal</source> (<year>2021</year>) <volume>71</volume>:<fpage>102035</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.media.2021.102035</pub-id>
</citation>
</ref>
<ref id="B39">
<label>39</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>R</given-names>
</name>
<name>
<surname>Lei</surname> <given-names>T</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>R</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>B</given-names>
</name>
<name>
<surname>Meng</surname> <given-names>H</given-names>
</name>
<name>
<surname>Nandi</surname> <given-names>AK</given-names>
</name>
</person-group>. <article-title>Medical image segmentation using deep learning: A survey</article-title>. <source>IET Image Process</source> (<year>2022</year>) <volume>162</volume>:<page-range>1243&#x2013;67</page-range>. doi: <pub-id pub-id-type="doi">10.1049/ipr2.12419</pub-id>
</citation>
</ref>
<ref id="B40">
<label>40</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>TY</given-names>
</name>
<name>
<surname>Goyal</surname> <given-names>P</given-names>
</name>
<name>
<surname>Girshick</surname> <given-names>R</given-names>
</name>
<name>
<surname>He</surname> <given-names>K</given-names>
</name>
<name>
<surname>Doll&#xe1;r</surname> <given-names>P</given-names>
</name>
</person-group>. (<year>2017</year>). <article-title>Focal loss for dense object detection</article-title>, in: <conf-name>Proceedings of the IEEE international conference on computer vision</conf-name>. <publisher-loc>Venice, Italy</publisher-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>2980&#x2013;8</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TPAMI.2018.2858826</pub-id>
</citation>
</ref>
<ref id="B41">
<label>41</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Jha</surname> <given-names>D</given-names>
</name>
<name>
<surname>Smedsrud</surname> <given-names>PH</given-names>
</name>
<name>
<surname>Riegler</surname> <given-names>MA</given-names>
</name>
<name>
<surname>Halvorsen</surname> <given-names>P</given-names>
</name>
<name>
<surname>Lange</surname> <given-names>TD</given-names>
</name>
<name>
<surname>Johansen</surname> <given-names>D</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Kvasir-seg : A segmented polyp dataset</article-title>, in: <conf-name>International Conference on Multimedia Modeling</conf-name>, <conf-loc>Daejeon, South Korea</conf-loc>. pp. <page-range>451&#x2013;62</page-range>.</citation>
</ref>
<ref id="B42">
<label>42</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Silva</surname> <given-names>J</given-names>
</name>
<name>
<surname>Histace</surname> <given-names>A</given-names>
</name>
<name>
<surname>Romain</surname> <given-names>O</given-names>
</name>
<name>
<surname>Dray</surname> <given-names>X</given-names>
</name>
<name>
<surname>Granado</surname> <given-names>B</given-names>
</name>
</person-group>. <article-title>Toward embedded detection of polyps in WCE images for early diagnosis of colorectal cancer</article-title>. <source>Int J Comput Assisted Radiol Surg</source> (<year>2014</year>) <volume>9</volume>:<page-range>283&#x2013;93</page-range>. doi: <pub-id pub-id-type="doi">10.1007/s11548-013-0926-3</pub-id>
</citation>
</ref>
<ref id="B43">
<label>43</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>V&#xe1;zquez</surname> <given-names>D</given-names>
</name>
<name>
<surname>Bernal</surname> <given-names>J</given-names>
</name>
<name>
<surname>S&#xe1;nchez</surname> <given-names>FJ</given-names>
</name>
<name>
<surname>Fern&#xe1;ndez-Esparrach</surname> <given-names>G</given-names>
</name>
<name>
<surname>L&#xf3;pez</surname> <given-names>AM</given-names>
</name>
<name>
<surname>Romero</surname> <given-names>A</given-names>
</name>
<etal/>
</person-group>. <article-title>A benchmark for endoluminal scene segmentation of colonoscopy images</article-title>. <source>J healthcare Eng</source> (<year>2017</year>) <volume>2017</volume>:<fpage>4037190</fpage>. doi: <pub-id pub-id-type="doi">10.1155/2017/4037190</pub-id>
</citation>
</ref>
<ref id="B44">
<label>44</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Fan</surname> <given-names>DP</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Li</surname> <given-names>T</given-names>
</name>
<name>
<surname>Borji</surname> <given-names>A</given-names>
</name>
</person-group>. (<year>2017</year>). <article-title>Structure-measure: A new way to evaluate foreground maps</article-title>, in: <conf-name>Proceedings of 2017 IEEE International Conference on Computer Vision</conf-name>, <conf-loc>Venice, Italy</conf-loc>: <publisher-name>IEEE</publisher-name>. pp. <page-range>4558&#x2013;67</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/ICCV.2017.487</pub-id>
</citation>
</ref>
<ref id="B45">
<label>45</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>T</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>H</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>D</given-names>
</name>
</person-group>. (<year>2021</year>). <article-title>UACANet: Uncertainty augmented context attention for polyp segmentation</article-title>, in: <conf-name>Proceedings of the 29th ACM International Conference on Multimedia</conf-name>, <conf-loc>New York, USA</conf-loc>: <publisher-name>ACM</publisher-name>. pp. <page-range>2167&#x2013;75</page-range>. doi:&#xa0;<pub-id pub-id-type="doi">10.1145/3474085.3475375</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>