<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Neurosci.</journal-id>
<journal-title>Frontiers in Neuroscience</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Neurosci.</abbrev-journal-title>
<issn pub-type="epub">1662-453X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fnins.2023.1207149</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Neuroscience</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Accurate segmentation algorithm of acoustic neuroma in the cerebellopontine angle based on ACP-TransUNet</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Zhuo</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<uri xlink:href="https://loop.frontiersin.org/people/1839901/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Xiaochen</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>Yong</given-names>
</name>
<xref rid="aff3" ref-type="aff"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Liu</surname>
<given-names>Jieyu</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zheng</surname>
<given-names>Chenzi</given-names>
</name>
<xref rid="aff4" ref-type="aff"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Bai</surname>
<given-names>Hua</given-names>
</name>
<xref rid="aff1" ref-type="aff"><sup>1</sup></xref>
<xref rid="c001" ref-type="corresp"><sup>&#x002A;</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Ma</surname>
<given-names>Quanfeng</given-names>
</name>
<xref rid="aff2" ref-type="aff"><sup>2</sup></xref>
<xref rid="c002" ref-type="corresp"><sup>&#x002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Tianjin Key Laboratory of Optoelectronic Detection Technology and Systems, School of Electronic and Information Engineering, Tiangong University</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Tianjin Cerebral Vascular and Neural Degenerative Disease Key Laboratory, Tianjin Huanhu Hospital</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Computer Science and Technology, Tiangong University</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>College of Foreign Languages, Nankai University</institution>, <addr-line>Tianjin</addr-line>, <country>China</country></aff>
<author-notes>
<fn id="fn0001" fn-type="edited-by"><p>Edited by: Jiajia Li, Xi'an University of Architecture and Technology, China</p></fn>
<fn id="fn0002" fn-type="edited-by"><p>Reviewed by: Song Tong, Tsinghua University, China; Hyo Jong Lee, Jeonbuk National University, Republic of Korea</p></fn>
<corresp id="c001">&#x002A;Correspondence: Hua Bai, <email>baihua@tiangong.edu.cn</email></corresp>
<corresp id="c002">Quanfeng Ma, <email>zhang20220614@163.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>24</day>
<month>05</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>17</volume>
<elocation-id>1207149</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>04</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>05</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2023 Zhang, Zhang, Yang, Liu, Zheng, Bai and Ma.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Zhang, Zhang, Yang, Liu, Zheng, Bai and Ma</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Acoustic neuroma is one of the most common tumors in the cerebellopontine angle area. Patients with acoustic neuroma have clinical manifestations of the cerebellopontine angle occupying syndrome, such as tinnitus, hearing impairment and even hearing loss. Acoustic neuromas often grow in the internal auditory canal. Neurosurgeons need to observe the lesion contour with the help of MRI images, which not only takes a lot of time, but also is easily affected by subjective factors. Therefore, the automatic and accurate segmentation of acoustic neuroma in cerebellopontine angle on MRI is of great significance for surgical treatment and expected rehabilitation. In this paper, an automatic segmentation method based on Transformer is proposed, using TransUNet as the core model. As some acoustic neuromas are irregular in shape and grow into the internal auditory canal, larger receptive fields are thus needed to synthesize the features. Therefore, we added Atrous Spatial Pyramid Pooling to CNN, which can obtain a larger receptive field without losing too much resolution. Since acoustic neuromas often occur in the cerebellopontine angle area with relatively fixed position, we combined channel attention with pixel attention in the up-sampling stage so as to make our model automatically learn different weights by adding the attention mechanism. In addition, we collected 300 MRI sequence nuclear resonance images of patients with acoustic neuromas in Tianjin Huanhu hospital for training and verification. The ablation experimental results show that the proposed method is reasonable and effective. The comparative experimental results show that the Dice and Hausdorff 95 metrics of the proposed method reach 95.74% and 1.9476&#x2009;mm respectively, indicating that it is not only superior to the classical models such as UNet, PANet, PSPNet, UNet++, and DeepLabv3, but also show better performance than the newly-proposed SOTA (state-of-the-art) models such as CCNet, MANet, BiseNetv2, Swin-Unet, MedT, TransUNet, and UCTransNet.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>acoustic neuroma</kwd>
<kwd>image segmentation</kwd>
<kwd>transformer</kwd>
<kwd>UNet</kwd>
</kwd-group>
<contract-num rid="cn1">61201106</contract-num>
<contract-num rid="cn2">2022SKY126</contract-num>
<contract-sponsor id="cn1">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<contract-sponsor id="cn2">Tianjin Research Innovation Project for Postgraduate Students</contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="6"/>
<equation-count count="7"/>
<ref-count count="67"/>
<page-count count="13"/>
<word-count count="8708"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Translational Neuroscience</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="sec1" sec-type="intro">
<label>1.</label>
<title>Introduction</title>
<p>Acoustic neuroma is one of the most common tumors in the cerebellopontine angle area, accounting for about 85% of the tumors in this region. Although these tumors are typically non-life-threatening, postoperative morbidity can be associated with injury to the facial nerve, cochlear nerve, cerebrospinal fluid leaks, and other wound complications. Permanent facial paralysis can occur in 3 to 5% of cases, and up to 22% of patients may experience cerebrospinal fluid leaks (<xref ref-type="bibr" rid="ref33">North et al., 2022</xref>). Fortunately, the surgical mortality rate is low, with less than 1% of cases resulting in death (<xref ref-type="bibr" rid="ref31">McClelland et al., 2011</xref>). The main manifestation of acoustic neuroma is the thickening of the auditory nerve. Due to the limitation of bone canal, the tumor gradually grows to the cerebellopontine angle area with less resistance (<xref ref-type="bibr" rid="ref25">Ling et al., 2016</xref>). The tumor originates from the vestibular part of the VIII pair of cranial nerves. The early lesions are small and often grow in the internal auditory canal. Neurosurgeons need to use Magnetic Resonance Imaging (MRI), which not only takes a lot of time, but also is susceptible to subjective factors. Therefore, it is of great significance to realize the automatic and accurate segmentation of acoustic neuroma. MRI has the characteristics of no bony artifacts, multi-directional and multi angle imaging, clear anatomical structure and high-level resolution for tissues. It can clearly show the size, shape, edge contour, peritumoral edema and adjacent structural changes of tumor, providing information for the preoperative diagnosis of tumor. It has become a preferred method for the examination of space occupying lesions in cerebellopontine angle (<xref ref-type="bibr" rid="ref56">Xiaoxia et al., 2014</xref>).</p>
<p>At present, in the medical field, manual segmentation is mainly used in brain tumor segmentation. Manual segmentation is to manually outline the tumor area in all tumor MRI image slices. Although manual segmentation is accurate, it is time-consuming, laborious and subjective, which is not conducive to the timely diagnosis and treatment of patients. Therefore, scholars have been exploring automatic segmentation methods. In the early stage, people mainly focused on traditional segmentation methods, such as threshold segmentation (<xref ref-type="bibr" rid="ref55">Xiaobo et al., 2019</xref>), watershed segmentation (<xref ref-type="bibr" rid="ref58">Yongzhuo and Shuguang, 2018</xref>), region segmentation (<xref ref-type="bibr" rid="ref36">Qiulin and Xin, 2018</xref>). There are also more complex segmentation methods based on statistical shape model [6] and graph cut (<xref ref-type="bibr" rid="ref10">Corso et al., 2008</xref>). Despite the high speed of these segmentation methods, its result depends on the parameters specified by the user and the preprocessing of MRI images (<xref ref-type="bibr" rid="ref26">Lingmei et al., 2020</xref>), which greatly limits its generalization ability.</p>
<p>With the rapid development of artificial intelligence in recent years, deep learning methods have been successfully applied to the field of medical images. Deep learning models solve the problems of poor accuracy and strong dependence on data in traditional automatic segmentation methods, such as threshold segmentation, region segmentation, and clustering segmentation, and have made great progress in medical image segmentation. AlexNet (<xref ref-type="bibr" rid="ref23">Krizhevsky et al., 2017</xref>), VGG (<xref ref-type="bibr" rid="ref42">Simonyan and Zisserman, 2014</xref>), GoogLeNet (<xref ref-type="bibr" rid="ref45">Szegedy et al., 2014</xref>), ResNet (<xref ref-type="bibr" rid="ref17">He et al., 2016</xref>), DenseNet (<xref ref-type="bibr" rid="ref18">Huang et al., 2016</xref>), and other deep and wide network structures have been proposed one after another to learn deeper data features. UNet (<xref ref-type="bibr" rid="ref38">Ronneberger et al., 2015</xref>) is a network structure proposed by Ronneberger et al. in 2015, which was originally applied in the field of biomedical cell segmentation. In 2019, Mumtaz et al. used a new method based on 3D fully convolutional neural networks (FCNNs; <xref ref-type="bibr" rid="ref41">Shelhamer et al., 2016</xref>) and a 3D level set segmentation algorithm to classify and segment colon and rectal cancer. Their accuracy was 0.9378, which was 0.0755 lower than the previous accuracy of 0.8623 (<xref ref-type="bibr" rid="ref44">Soomro et al., 2018</xref>). <xref ref-type="bibr" rid="ref11">Cuixia et al. (2019)</xref> discussed and compared various classification models for breast tumors using deep learning in 2019 and proposed a novel method that combines deep learning features. Deep learning is also widely applied in brain tumor segmentation. <xref ref-type="bibr" rid="ref47">Thillaikkarasi and Saravanan (2019)</xref> proposed a brain tumor segmentation algorithm using a support vector machine to extract features and CNN segmentation in 2019, resulting in an accuracy of 84%. <xref ref-type="bibr" rid="ref13">Dong et al. (2017)</xref> used UNet to segment MRI images of brain tumors and achieved good results by splicing feature vectors of the expansion path and contraction path through skip connections. <xref ref-type="bibr" rid="ref26">Lingmei et al. (2020)</xref> improved the UNet structure in 2020 and applied it to the segmentation of glioma magnetic resonance images. Specifically, they used an attention module on the contraction path of UNet to distribute weight to convolution layers of different sizes, promoting the utilization of spatial and contextual information. Replacing the original convolution layer with the residual compact module can extract more features and promote network convergence. In 2021, <xref ref-type="bibr" rid="ref40">Russo et al. (2020)</xref> applied a spherical transformation preprocessing input training model, which was better than the Descartes input training model in predicting glioma tumor core segmentation and enhancing tumor category. The two models were combined to further improve prediction accuracy.</p>
<p>Undoubtedly, CNN represents a very promising method for image processing. However, its convolution operation has limitations, especially for samples with large texture differences, resulting in weak performance. In recent years, scholars have proposed several solutions to address this issue. For instance, <xref ref-type="bibr" rid="ref6">Chen et al. (2014)</xref> introduced the Atrous Spatial Pyramid Pooling (ASPP) module in DeepLabv3+ (<xref ref-type="bibr" rid="ref9">Chen et al., 2018a</xref>) after several generations of improvements (<xref ref-type="bibr" rid="ref8">Chen et al., 2017</xref>, <xref ref-type="bibr" rid="ref7">2018b</xref>). The addition of ASPP into CNN enables atrous convolution to expand the vision field of the filter without increasing computational demand. Therefore, ASPP can obtain feature information of different scales without using a pooling layer, overcoming the limitations of local information loss caused by grid effect and the lack of correlation between long-distance information when using a single atrous convolution. Moreover, some studies suggest building a self-attention mechanism based on CNN features (<xref ref-type="bibr" rid="ref52">Wang et al., 2017</xref>) as an effective means to solve the limitations of convolution operations. This method has also garnered much attention in the field of artificial intelligence. For instance, <xref ref-type="bibr" rid="ref48">Tian et al. (2020)</xref> used channel attention in ADNet to accurately extract useful information hidden in the complex background. <xref ref-type="bibr" rid="ref19">Huang et al. (2020)</xref> proposed the Criss-cross attention module in CCNet to capture contextual information of the complete image. <xref ref-type="bibr" rid="ref16">Fan et al. (2020)</xref> introduced the self-attention mechanism in 2020 and proposed Multi-scale Attention Net (MA-Net).</p>
<p>Furthermore, Transformer has emerged as an alternative architecture designed for sequence-to-sequence prediction, and its success has been widely demonstrated in various fields such as machine translation and natural language processing (NLP; <xref ref-type="bibr" rid="ref50">Vaswani et al., 2017</xref>; <xref ref-type="bibr" rid="ref12">Devlin et al., 2018</xref>). In various image recognition tasks, Transformer has proven to reach or even exceed the state-of-the-art (<xref ref-type="bibr" rid="ref65">Zheng et al., 2020</xref>; <xref ref-type="bibr" rid="ref14">Dosovitskiy et al., 2021</xref>). For example, <xref ref-type="bibr" rid="ref5">Chen et al. (2021)</xref> combined Transformer as a powerful encoder for medical image segmentation tasks with UNet in 2021, proposing TransUNet as a powerful alternative for medical image segmentation. Yang et al. added an attention mechanism to TransUNet (<xref ref-type="bibr" rid="ref57">Yang and Mehrkanoon, 2022</xref>), showing that the combination of attention mechanism and TransUNet can optimize the segmentation effect. Subsequently, Valanarasu et al. proposed the MedT (<xref ref-type="bibr" rid="ref49">Valanarasu et al., 2021</xref>) containing Local&#x2013;Global (Logo) training strategy based on Transformer, which further improved the model&#x2019;s performance. Cao H et al. fused high-resolution features from different scales of the encoder by skip connections, and Swin-Unet (<xref ref-type="bibr" rid="ref3">Cao et al., 2021</xref>) was proposed to mitigate the loss of spatial information due to the pooling operation.</p>
<p>It is worth noting that acoustic neuromas have different shapes and may grow into the inner auditory canal, which is challenging for accurate feature extraction. We believe that the combination of ASPP, attention mechanism and Transformer can solve this challenge well. Therefore, we propose a novel model called ACP-TransUNet for accurate segmentation of acoustic neuromas, with TransUNet as the core framework. Specifically, the ASPP module is added to increase the receptive field, enabling more accurate and noticeable extraction of tumor features during the segmentation process. We also incorporate the CPAT module, which combines channel attention (<xref ref-type="bibr" rid="ref21">Jie et al., 2019</xref>) and pixel attention (<xref ref-type="bibr" rid="ref62">Zhao et al., 2020</xref>) to better explore channel and pixel features of acoustic neuromas while recovering the original input image size. The use of feature multiplication between attentions enhances the ability of feature representation and improves the feature propagation strategy, resulting in higher performance under the same computational load (<xref ref-type="bibr" rid="ref62">Zhao et al., 2020</xref>; e.g., RCAN, <xref ref-type="bibr" rid="ref61">Zhang et al., 2018</xref>; CARN, <xref ref-type="bibr" rid="ref1">Ahn et al., 2018</xref>). By arranging the channel attention and pixel attention sequentially, we aim to improve the feature extraction capability of ACP-TransUNet.</p>
<p>Our main contributions are as follows:</p>
<list list-type="order">
<list-item>
<p>Our proposed ACP TransUNet combines Transformer and CNN to capture the global and local features of the segmentation target.</p>
</list-item>
<list-item>
<p>In the down-sampling process, the ASPP module is added after the convolutional neural network to gain contextual information at multiple scales and resolutions.</p>
</list-item>
<list-item>
<p>In the up-sampling process, channel attention and pixel attention are used to improve model performance and accuracy by weighting important features.</p>
</list-item>
</list>
</sec>
<sec id="sec2">
<label>2.</label>
<title>Related works</title>
<sec id="sec3">
<label>2.1.</label>
<title>TransUNet</title>
<p>UNet has become the most commonly used method to accurately segment lesions in medical segmentation tasks, and Transformer has also become a structural system that replaces the self-attention mechanism. TransUNet combines Transformer with UNet as a powerful alternative for medical image segmentation, possessing the advantages of both. To compensate for the loss of feature resolution due to Transformers, TransUNet adopted a hybrid CNN-Transformer architecture to exploit the detailed high-resolution spatial information of CNN features and the global context encoded by Transformers. Inspired by U-Shape, the attention features encoded by Transformers are combined with different high-resolution CNN features during upsampling to achieve precise localization. This design enables the model to preserve the advantages of Transformer and also facilitates the segmentation of medical images. On the one hand, Transformer encodes the tokenized image patches of the convolutional neural network (CNN) feature map as an input sequence for feature extraction; on the other hand, the decoder up-sampling the encoded features, and then combines them with the feature map in CNN to achieve accurate positioning (<xref ref-type="bibr" rid="ref5">Chen et al., 2021</xref>). Currently, TransUNet and its variants have achieved great success in image segmentation. Nur&#x00E7;in used TransUNet for the segmentation step of the red blood cells to improve the segmentation quality of overlapping cells (<xref ref-type="bibr" rid="ref34">Nur&#x00E7;in, 2022</xref>). MS-TransUNet++ (<xref ref-type="bibr" rid="ref53">Wang et al., 2022</xref>) employed a multi-scale and flexible feature fusion scheme between different levels of encoders and decoders to achieve competitive performance in prostate MR and liver CT image segmentation. Liu et al. proposed an efficient model called TransUNet+ (<xref ref-type="bibr" rid="ref29">Liu et al., 2022</xref>) through a redesigned skip connection, which has achieved promising results in medical image segmentation. Wang et al. proposed UCTransNet (<xref ref-type="bibr" rid="ref51">Wang et al., 2021</xref>), which used the CTrans block to replace the skip connection in U-Net and obtained a higher segmentation effect. DS-TransUNet (<xref ref-type="bibr" rid="ref24">Lin et al., 2022</xref>) applied swin transformer block (<xref ref-type="bibr" rid="ref30">Liu et al., 2021</xref>) to encoder and decoder. This may be the first attempt to combine the advantages of layered Swin Transformer into both encoder and decoder of standard U-shaped architecture with the aim of improving the segmentation quality of different medical images. In TransAttUnet (<xref ref-type="bibr" rid="ref4">Chen et al., 2021</xref>), multilevel guided attention and multiscale skip connection were co-developed to effectively improve the functionality and flexibility of the traditional U-shaped architecture. Zhao et al. proposed an automatic deep learning pipeline nn-TransUNet (<xref ref-type="bibr" rid="ref64">Zhao et al., 2022</xref>) for cardiac MRI segmentation by combining the experimental planning of nn-UNet and the network architecture of TransUNet. EG-TransUNet (<xref ref-type="bibr" rid="ref35">Pan et al., 2023</xref>) used progressive enhancement module, channel spatial attention, and semantic guidance attention to be able to capture object variability on different biomedical datasets. In summary, the architecture of TransUNet combines the advantages of Transformer and CNN, which is not only good for local information extraction, but also can explore long-range modeling.</p>
</sec>
<sec id="sec4">
<label>2.2.</label>
<title>Channel attention</title>
<p>Channel attention was first proposed in SE-Net and achieved excellent performance. In CBAM (<xref ref-type="bibr" rid="ref54">Woo et al., 2018</xref>), channel attention has been improved significantly. Specifically, channel attention compresses the feature of spatial dimension, i.e., each two-dimensional feature map becomes a real number, which is equivalent to the pooling operation with global receptive field. The number of feature channels remains unchanged, and the module structure is shown in <xref rid="fig1" ref-type="fig">Figure 1</xref>. Channel attention aggregates spatial information of feature maps based on global average pooling <inline-formula>
<mml:math id="M1">
<mml:mrow>
<mml:mi mathvariant="normal">AvgPool</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and maximum pooling <inline-formula>
<mml:math id="M2">
<mml:mrow>
<mml:mi mathvariant="normal">MaxPool</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> operations, generating two different spatial context descriptors: <inline-formula>
<mml:math id="M3">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">avg</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M4">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, representing average pool features and maximum pool features, respectively. After adding the two feature maps of the multilayer perceptron (MLP), the Sigmoid function is used to generate channel feature map, as follows in <xref ref-type="disp-formula" rid="EQ1">Eq. (1)</xref>:</p>
<disp-formula id="EQ1"><label>(1)</label><mml:math id="M5">
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">C</mml:mi>
</mml:msub>
<mml:mfenced>
<mml:mi mathvariant="normal">F</mml:mi>
</mml:mfenced>
<mml:mo>=</mml:mo>
<mml:mi>&#x03C3;</mml:mi>
<mml:mi mathvariant="normal">(MLP</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="normal">AvgPool</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="normal">MLP(MaxPool(F)))</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>&#x03C3;</mml:mi>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">avg</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math id="M6">
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
</mml:math>
</inline-formula> represents the Sigmoid function, <inline-formula>
<mml:math id="M7">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M8">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represent the two convolution operations, respectively, and <inline-formula>
<mml:math id="M9">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">avg</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M10">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mrow>
<mml:mi>max</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> represent the average pooling and max pooling, respectively. Sigmoid function can map the result to 0&#x2013;1 with the amplitude unchanged, so we can get the weight of each feature point of the input channel feature layer.</p>
<fig position="float" id="fig1">
<label>Figure 1</label>
<caption>
<p>Overview of the channel attention structure.</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g001.tif"/>
</fig>
<p>In recent years, channel attention has been widely used to solve medical challenges. Yuan et al. improved the accuracy of automatic vessel segmentation in fundus images by embedding an adaptive channel attention module to automatically rank the importance of each feature channel (<xref ref-type="bibr" rid="ref60">Yuan et al., 2021</xref>). Du et al. applied channel attention to the automatic segmentation of early gastric cancer (EGC) to extract subtle discriminative features of EGC lesions by capturing the interdependence between channel features (<xref ref-type="bibr" rid="ref15">Du et al., 2023</xref>). In addition, channel attention paired with other excellent attention mechanisms can also improve the quality of super-resolution reconstruction of medical images. Song et al. and Zhu et al. obtained high-quality reconstructed images for glioma MRI images and lung cancer CT images, respectively (<xref ref-type="bibr" rid="ref67">Zhu et al., 2022</xref>; <xref ref-type="bibr" rid="ref43">Song et al., 2023</xref>). Therefore, channel attention has great potential in the field of medical image processing.</p>
</sec>
<sec id="sec5">
<label>2.3.</label>
<title>Pixel attention</title>
<p>The channel attention aims to obtain a <inline-formula>
<mml:math id="M11">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>D</mml:mi>
<mml:mspace width="thickmathspace"/>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> vector of attentional features. In contrast, pixel attention (<xref ref-type="bibr" rid="ref62">Zhao et al., 2020</xref>) is able to generate <inline-formula>
<mml:math id="M12">
<mml:mrow>
<mml:mn>3</mml:mn>
<mml:mi>D</mml:mi>
<mml:mspace width="thickmathspace"/>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> matrices as attention features. Note that <inline-formula>
<mml:math id="M13">
<mml:mi>C</mml:mi>
</mml:math>
</inline-formula> is the number of channels, and <inline-formula>
<mml:math id="M14">
<mml:mi>H</mml:mi>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M15">
<mml:mi>W</mml:mi>
</mml:math>
</inline-formula> are the height and width of the features, respectively. Specifically, pixel attention generates attention coefficients for all pixels of the feature map. As shown in <xref rid="fig2" ref-type="fig">Figure 2</xref>, pixel attention uses only 1&#x2009;&#x00D7;&#x2009;1 convolutional layers and Sigmoid functions to obtain the attention map, and then multiplies the attention map with the input features, as follows in <xref ref-type="disp-formula" rid="EQ2">Eq. (2)</xref>:</p>
<fig position="float" id="fig2">
<label>Figure 2</label>
<caption>
<p>Overview of pixel attention structure.</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g002.tif"/>
</fig>
<disp-formula id="EQ2"><label>(2)</label><mml:math id="M16">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:msub>
<mml:mfenced>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mfenced>
<mml:mo>=</mml:mo>
<mml:mi>&#x03C3;</mml:mi>
<mml:mfenced>
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">PA</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mfenced>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math></disp-formula>
<p>where <inline-formula>
<mml:math id="M17">
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
</mml:math>
</inline-formula> represents the Sigmoid function and <inline-formula>
<mml:math id="M18">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">f</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">PA</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> represents a convolution operation with the filter size of <inline-formula>
<mml:math id="M19">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Pixel attention not only reduces the number of parameters, but also eliminates unnecessary pooling operations that can lead to image smoothing (<xref ref-type="bibr" rid="ref46">Tang et al., 2021</xref>). Relying on this advantage, pixel attention is widely used in the field of medical images for segmentation (<xref ref-type="bibr" rid="ref39">Roy et al., 2022</xref>) and super-resolution reconstruction tasks (<xref ref-type="bibr" rid="ref37">Rajeshwari and Shyamala, 2023</xref>).</p>
</sec>
</sec>
<sec id="sec6" sec-type="methods">
<label>3.</label>
<title>Methods</title>
<sec id="sec7">
<label>3.1.</label>
<title>Overview</title>
<p>In this section, we describe our ACP-TransUNet with more details. The ACP-TransUNet model proposed in this paper is based on the TransUNet (<xref ref-type="bibr" rid="ref5">Chen et al., 2021</xref>) model, and is improved and extended on the basis of the latter, as shown in <xref rid="fig3" ref-type="fig">Figure 3</xref>.</p>
<fig position="float" id="fig3">
<label>Figure 3</label>
<caption>
<p>Overview of ACP-TransUNet. The input is an acoustic neuroma MRI image, and the output is the corresponding prediction map generated by ACP-TransUNet.</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g003.tif"/>
</fig>
<p>Given an input image with resolution <inline-formula>
<mml:math id="M20">
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi mathvariant="normal">W</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and C number of channels, the segmentation map is obtained by down-sampling and up-sampling. The down-sampling process consists of five parts, which are CNN, ASPP, Image Sequentialization, Patch Embedding, and Transformer Layer. The input image is first extracted by CNN layer to get the feature map. After that, the ASPP module is used to increase the receptive field to obtain a feature map with different scales. Then, Hidden Feature and Linear Projection reshape the feature map into N flattened 2D patches for Image Sequentialization, with each patch of size <inline-formula>
<mml:math id="M21">
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math id="M22">
<mml:mrow>
<mml:mi mathvariant="normal">N</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#x2032;</mml:mo>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math id="M23">
<mml:msup>
<mml:mi mathvariant="normal">H</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math id="M24">
<mml:msup>
<mml:mi mathvariant="normal">W</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:math>
</inline-formula> being the length and width of each feature map. In order to encode the spatial information of the patches, we add positional embedding to the patch embedding to preserve the positional information, as follows in <xref ref-type="disp-formula" rid="EQ3">Eq. (3)</xref>:</p>
<disp-formula id="EQ3"><label>(3)</label> <mml:math id="M25">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">Z</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfenced close="]" open="[">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">x</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="normal">E; x</mml:mi>
</mml:mrow>
<mml:mi mathvariant="normal">p</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mi mathvariant="normal">E; </mml:mi>
<mml:mo>&#x2026;</mml:mo>
<mml:mspace width="0.25em"/>
<mml:mo>;</mml:mo>
<mml:mi mathvariant="normal">&#x2004;</mml:mi>
<mml:msubsup>
<mml:mi mathvariant="normal">x</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
<mml:mi mathvariant="normal">N</mml:mi>
</mml:msubsup>
<mml:mi mathvariant="normal">E</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">E</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math></disp-formula>
<p>where <inline-formula>
<mml:math id="M26">
<mml:mrow>
<mml:mi mathvariant="normal">E</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211D;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>.</mml:mo>
<mml:mi mathvariant="normal">C</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi mathvariant="normal">D</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the patch embedding projection, <inline-formula>
<mml:math id="M27">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">x</mml:mi>
<mml:mi mathvariant="normal">p</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the vectorized patch, and <inline-formula>
<mml:math id="M28">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">E</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211D;</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">N</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi mathvariant="normal">D</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the position embedding.</p>
<p>The Transformer (<xref ref-type="bibr" rid="ref50">Vaswani et al., 2017</xref>) layer is added at the end of the down-sampling to obtain the global features, which consists of Multi-head Attention (MSA) and Multi-layer Perceptron (MLP) as shown in <xref ref-type="disp-formula" rid="EQ4">Eqs. (4)</xref> and <xref ref-type="disp-formula" rid="EQ5">(5)</xref>:</p>
<disp-formula id="EQ4"><label>(4)</label><mml:math id="M29">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">Z</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="normal">MSA</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">LN</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mrow>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math></disp-formula>
<disp-formula id="EQ5"><label>(5)</label><mml:math id="M30">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">Z</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="normal">MLP</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">LN</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
</mml:mrow>
</mml:math> </disp-formula>
<p>where <inline-formula>
<mml:math id="M31">
<mml:mrow>
<mml:mi mathvariant="normal">LN</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mo>&#x00B7;</mml:mo>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes the layer normalization operator and <inline-formula>
<mml:math id="M32">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">z</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the encoded image representation.</p>
<p>In the up-sampling process, we added CPAT modules in each layer to weight the important features in recovering the image size to improve the performance and accuracy of the model.</p>
</sec>
<sec id="sec8">
<label>3.2.</label>
<title>ASPP module</title>
<p>Acoustic neuromas vary in shape. Some are irregular in shape and grow into the inner auditory canal, while some have clear boundary. Therefore, we need a larger receptive field to extract the feature of acoustic neuromas. The ordinary convolution structure cannot fully extract features, so in this paper we choose to use ASPP module to strengthen the ability of the model to segment objects at different scales. As shown in <xref rid="fig4" ref-type="fig">Figure 4</xref>, in this paper, ASPP module is equipped in the last layer of CNN, with dilation rate set to 2, 4, 8. The rate of atrous convolution is based on the ordinary convolution, and the interval between adjacent weights is <inline-formula>
<mml:math id="M33">
<mml:mrow>
<mml:mi mathvariant="normal">rate</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. The rate of ordinary convolution is defaulted to 1, so the actual size of atrous convolution is <inline-formula>
<mml:math id="M34">
<mml:mrow>
<mml:mi mathvariant="normal">k</mml:mi>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">rate</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, in which k is the size of the original convolution kernel. ASPP overcomes the shortcomings of local information loss and lack of correlation in remote information caused by grid effect when using single atrous convolution, making it possible to obtain different scale feature information without using pooling layer.</p>
<fig position="float" id="fig4">
<label>Figure 4</label>
<caption>
<p>Overview of ASPP Module.</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g004.tif"/>
</fig>
</sec>
<sec id="sec9">
<label>3.3.</label>
<title>CPAT module</title>
<p>Given an intermediate feature map <inline-formula>
<mml:math id="M35">
<mml:mrow>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211D;</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> as input, CPAT module sequentially infers a 1D channel attention map <inline-formula>
<mml:math id="M36">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211D;</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and a 3D pixel attention map <inline-formula>
<mml:math id="M37">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211D;</mml:mi>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>W</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> as illustrated in <xref rid="fig5" ref-type="fig">Figure 5</xref>. For the arrangement of attention modules, we found through experiments that the result is better when using two sequential attentions than using one attention, which will be discussed in the ablation experiments, as shown in <xref ref-type="disp-formula" rid="EQ6">Eqs. (6)</xref> and <xref ref-type="disp-formula" rid="EQ7">(7)</xref>:</p>
<fig position="float" id="fig5">
<label>Figure 5</label>
<caption>
<p>Overview of CPAT structure. This module has two submodules: channels and pixels, where <inline-formula>
<mml:math id="M43">
<mml:mo>&#x2297;</mml:mo>
</mml:math>
</inline-formula> denotes element-wise multiplication. The intermediate feature map is adaptively refined through our module (CPAT).</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g005.tif"/>
</fig>
<disp-formula id="EQ6"><label>(6)</label> <mml:math id="M38">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">c</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:mi mathvariant="normal">F</mml:mi>
</mml:mrow>
</mml:math></disp-formula>
<disp-formula id="EQ7"><label>(7)</label><mml:math id="M39">
<mml:mrow>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2033;</mml:mo>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">M</mml:mi>
<mml:mi mathvariant="normal">P</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2297;</mml:mo>
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math> </disp-formula>
<p>where <inline-formula>
<mml:math id="M40">
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:math>
</inline-formula> denotes the feature map obtained by channel attention, <inline-formula>
<mml:math id="M41">
<mml:msup>
<mml:mi mathvariant="normal">F</mml:mi>
<mml:mo>&#x2033;</mml:mo>
</mml:msup>
</mml:math>
</inline-formula> denotes the feature map obtained by pixel attention, and <inline-formula>
<mml:math id="M42">
<mml:mo>&#x2297;</mml:mo>
</mml:math>
</inline-formula> denotes element multiplication.</p>
</sec>
</sec>
<sec id="sec10">
<label>4.</label>
<title>Experimental results</title>
<p>In this section, we introduce the details of the experimental data and results. In order to verify whether ACP-TransUNet can effectively and accurately segment acoustic neuromas, we first performed comparative experiments and ablation experiments on all test sets (including coronal view, sagittal view, and transverse view). To test the accuracy of the model&#x2019;s segmentation effect in a single view, we also conducted multi-view evaluation, performing a comparison experiment and ablation experiment on the three views separately. The results are discussed in detail below. Among them, ACP-TransUNet achieves 95.74% Dice Similarity Coefficient on the test set, and Hausdorff 95 reaches 1.9476&#x2009;mm, which are superior than other models.</p>
<sec id="sec11">
<label>4.1.</label>
<title>Dataset</title>
<p>We selected MRI images of sagittal view, coronal view and transverse view of patients with cerebellopontine angle (CPA) acoustic neuroma diagnosed by experts in Tianjin Huanhu Hospital from January 2019 to January 2022, with all the patients signing informed consent. The scanning equipment we used was Siemens Skyra 3.0&#x2009;T MRI scanner, which could collect magnetic resonance images of multiple sequences. However, compared with other sequences, T1WI-SE could better distinguish the lesion and its surrounding adjacent tissues. Therefore, this paper adopts contrast - enhanced fast low-angle shot 2-dimensional sequence (T1_fl2d) with Gd-GDPA. Scanning parameters are as follows: slice thickness is 5&#x2009;mm; slice interval, 1.5&#x2009;mm; echo time (TE), 2.46&#x2009;ms; repetition time (TR), 220&#x2009;ms. After screening, a total of 300 magnetic resonance images of acoustic neuromas were selected in this paper, in which the ratio of training set, verification set and test set is 8: 1: 1 and each part has no cross.</p>
</sec>
<sec id="sec12">
<label>4.2.</label>
<title>Preprocessing</title>
<p>To avoid the deviation of the experimental results caused by the inconsistent data format, the training, verification and test MRI images in this paper are all set to the same format. Because the dataset is small, to improve the generalization ability of the model, the images are subjected to data augmentation processing such as inversion and flipping. In order to save training resources, the images are set to <inline-formula>
<mml:math id="M44">
<mml:mrow>
<mml:mn>512</mml:mn>
<mml:mo>&#x00D7;</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> pixels. The gray value visualization of the MRI image is shown in <xref rid="fig6" ref-type="fig">Figure 6</xref>.</p>
<fig position="float" id="fig6">
<label>Figure 6</label>
<caption>
<p>Gray visualization of MRI images in three directions. a-1, a-2, and a-3 represent coronal, sagittal and transverse MRI images, respectively. b-1, b-2, and b-3 are three-dimensional gray-scale visualization images of nuclear magnetic resonance, which represent the corresponding directions. The x-axis and y-axis represent the length and width of the image respectively, and the value range is [0, 512]. The z-axis represents the gray value distribution of the image, and the value range is [0, 255].</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g006.tif"/>
</fig>
</sec>
<sec id="sec13">
<label>4.3.</label>
<title>Experimental setup</title>
<p>In the experiment, the framework we used was Pytorch, and batchsize was set to 4. All networks trained 100 epochs on Nvidia Tesla V100 GPU. Specifically, we used a pre-training model (R50&#x2009;+&#x2009;ViT-B_16) that was trained on the ImageNet21k dataset. The pre-training model can be found at the following link: <ext-link xlink:href="https://console.cloud.google.com/storage/vit_models/" ext-link-type="uri">https://console.cloud.google.com/storage/vit_models/</ext-link>. In addition, we use the Adam optimizer (<xref ref-type="bibr" rid="ref22">Kingma and Ba, 2014</xref>) to optimize, the initial learning rate is <inline-formula>
<mml:math id="M45">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mn>10</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, and use the StepLR mechanism to set the learning rate attenuation according to epoch. The StepLR mechanism is a way to adjust the learning rate during training in machine learning. It reduces the learning rate by a certain factor after a fixed number of epochs or iterations. We set the &#x201C;step_size&#x201D; parameter to 7 and the &#x201C;gamma&#x201D; parameter to 0.1, which means that the learning rate was reduced by a factor of 0.1 every 7 epochs. By gradually reducing the learning rate, we aimed to improve the convergence of the model and prevent overfitting.</p>
</sec>
<sec id="sec14">
<label>4.4.</label>
<title>Evaluation metrics</title>
<p>In order to objectively evaluate the results of different models, this paper uses the Dice Similarity Coefficient (<xref ref-type="bibr" rid="ref32">Mehta, 2015</xref>; <xref ref-type="bibr" rid="ref27">Liu et al., 2020</xref>) and Hausdorff 95 (<xref ref-type="bibr" rid="ref20">Huttenlocher et al., 1993</xref>; <xref ref-type="bibr" rid="ref2">Beauchemin et al., 1998</xref>) as representative segmentation performance indicators, which measure the similarity and maximum mismatch between the segmentation result and the labeling result, respectively. These metrics are widely used in medical image segmentation studies and have been shown to be effective in evaluating segmentation performance.</p>
</sec>
<sec id="sec15">
<label>4.5.</label>
<title>Comparative experiment</title>
<p>To verify the validity of the proposed model, we compared several classical networks such as PANet (<xref ref-type="bibr" rid="ref28">Liu et al., 2018</xref>), PSPNet (<xref ref-type="bibr" rid="ref63">Zhao et al., 2016</xref>), UNet++ (<xref ref-type="bibr" rid="ref66">Zhou et al., 2018</xref>), and DeeplabV3 (<xref ref-type="bibr" rid="ref9">Chen et al., 2018a</xref>), as well as some emerging networks such as CCNet (<xref ref-type="bibr" rid="ref19">Huang et al., 2020</xref>), MANet (<xref ref-type="bibr" rid="ref16">Fan et al., 2020</xref>), BiseNetv2 (<xref ref-type="bibr" rid="ref59">Yu et al., 2021</xref>), Swin-Unet (<xref ref-type="bibr" rid="ref3">Cao et al., 2021</xref>), MedT (<xref ref-type="bibr" rid="ref49">Valanarasu et al., 2021</xref>), TransUNet (<xref ref-type="bibr" rid="ref5">Chen et al., 2021</xref>), and UCTransNet (<xref ref-type="bibr" rid="ref51">Wang et al., 2021</xref>), which have shown great performance on segmentation tasks in recent years. <xref rid="tab1" ref-type="table">Table 1</xref> summarizes the comparison results between our scheme and these representative networks. For each model, we visualized the segmentation effect in the coronal (cor), sagittal (sag), and transverse (tra) views, and the results are shown in <xref rid="fig7" ref-type="fig">Figure 7</xref>.</p>
<table-wrap position="float" id="tab1">
<label>Table 1</label>
<caption>
<p>Results of comparative experiment.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Model</th>
<th align="center" valign="top">Dice (%)</th>
<th align="center" valign="top">Hausdorff 95 (mm)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">UNet (2015)</td>
<td align="center" valign="top">94.65</td>
<td align="center" valign="top">4.4982</td>
</tr>
<tr>
<td align="left" valign="top">PSPNet (2016)</td>
<td align="center" valign="top">93.11</td>
<td align="center" valign="top">3.2145</td>
</tr>
<tr>
<td align="left" valign="top">DeepLabv3 (2017)</td>
<td align="center" valign="top">93.46</td>
<td align="center" valign="top">4.4438</td>
</tr>
<tr>
<td align="left" valign="top">UNet++ (2018)</td>
<td align="center" valign="top">94.66</td>
<td align="center" valign="top">3.7744</td>
</tr>
<tr>
<td align="left" valign="top">PANet (2018)</td>
<td align="center" valign="top">93.88</td>
<td align="center" valign="top">4.3399</td>
</tr>
<tr>
<td align="left" valign="top">CCNet (2020)</td>
<td align="center" valign="top">85.32</td>
<td align="center" valign="top">5.2548</td>
</tr>
<tr>
<td align="left" valign="top">MANet (2020)</td>
<td align="center" valign="top">94.95</td>
<td align="center" valign="top">3.7483</td>
</tr>
<tr>
<td align="left" valign="top">BiseNetv2 (2021)</td>
<td align="center" valign="top">89.86</td>
<td align="center" valign="top">5.6903</td>
</tr>
<tr>
<td align="left" valign="top">Swin-Unet (2021)</td>
<td align="center" valign="top">91.46</td>
<td align="center" valign="top">6.3458</td>
</tr>
<tr>
<td align="left" valign="top">MedT (2021)</td>
<td align="center" valign="top">93.26</td>
<td align="center" valign="top">4.7794</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet (2021)</td>
<td align="center" valign="top">95.02</td>
<td align="center" valign="top">4.0037</td>
</tr>
<tr>
<td align="left" valign="top">UCTransNet (2022)</td>
<td align="center" valign="top">95.06</td>
<td align="center" valign="top">4.0746</td>
</tr>
<tr>
<td align="left" valign="top">Ours</td>
<td align="center" valign="top"><bold>95.74</bold></td>
<td align="center" valign="top"><bold>1.9476</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column.</p>
</table-wrap-foot>
</table-wrap>
<fig position="float" id="fig7">
<label>Figure 7</label>
<caption>
<p>Examples of predictions for each network on acoustic neuromas in comparative experiments.</p>
</caption>
<graphic xlink:href="fnins-17-1207149-g007.tif"/>
</fig>
<p>The results show that ACP-TransUNet achieved the best performance on the test set, with a Dice value of 95.74% and a Hausdorff 95 value of 1.9476&#x2009;mm. Compared with the original UNet network proposed by <xref ref-type="bibr" rid="ref38">Ronneberger et al. (2015)</xref>, ACP-TransUNet achieved improvements of 1.09% and 2.5506&#x2009;mm in Dice and Hausdorff 95, respectively.</p>
<p>In the comparative experiments, our scheme achieved optimal Dice and Hausdorff 95 values, outperforming other network models. Specifically, our scheme improved Dice by 2.63% (PSPNet), 2.28% (DeepLabv3), 1.08% (UNet++), 1.86% (PANet), 10.42% (CCNet), 0.79% (MANet), 5.88% (BiseNetv2), 4.28% (Swin-Unet), 2.48% (MedT), 0.72% (TransUNet), and 0.68% (UCTransNet), respectively. Hausdorff 95 was increased by 1.2669&#x2009;mm (PSPNet), 2.4962&#x2009;mm (DeepLabv3), 1.8268&#x2009;mm (UNet++), 2.3923&#x2009;mm (PANet), 3.3072&#x2009;mm (CCNet), 1.8007&#x2009;mm (MANet), 3.7427&#x2009;mm (BiseNetv2), 4.3982&#x2009;mm (Swin-Unet), 2.8318&#x2009;mm (MedT), 2.0561&#x2009;mm (TransUNet), and 2.127&#x2009;mm (UCTransNet), respectively. The corresponding segmentation effect in <xref rid="fig7" ref-type="fig">Figure 7</xref> demonstrates the superior performance of ACP-TransUNet.</p>
<p>In comparison experiments, for some regular acoustic neuromas, such as the tumor shown in the sagittal view, it can be seen that the selected networks can achieve basic segmentation of the tumor except for BiseNetv2 and MedT. However, comparing the internal filling and boundary of the segmentation map, only ACP-TransUNet is closest to Ground Truth; for the part that shows irregular shape and grows into the internal auditory canal, as shown in the coronal view, PSPNet, PANet, CCNet, BiseNetv2 and MedT cannot well segment some tumors growing in the internal auditory canal. Although UNet++ and MANet could segment the tumors in the internal auditory tract, the segmentation results were inferior to the rest of the networks. DeepLabv3, UNet and TransUNet performed comparably to ACP-TransUNet for segmenting the tumors in the internal auditory tract, but UCTransNet and ACP-TransUNet outperformed the rest of the models in terms of edge detail. However, in the transverse (tra), only TransUNet and ACP-TransUNet can well segment the acoustic neuroma. We noticed that the models containing Transformer structures (such as MedT, Swin-Unet, TransNet, and UCTransNet) were deficient in processing edge details, which may be explained by the limited Transformer localization ability caused by insufficient low-level details. After adding the CPAT module and ASPP module, the segmentation map edge contours have been greatly improved.</p>
</sec>
<sec id="sec16">
<label>4.6.</label>
<title>Multi-view evaluation</title>
<p>To further verify the effectiveness of the model, we conduct comparative experiments and ablation experiments on the segmentation effects of the coronal, sagittal and transverse views in the test set, respectively. The results of the multi-view evaluation in the comparative experiments are shown in <xref rid="tab2" ref-type="table">Table 2</xref>.</p>
<table-wrap position="float" id="tab2">
<label>Table 2</label>
<caption>
<p>Comparative experiment from multiple perspectives.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top" rowspan="2">Model</th>
<th align="center" valign="top" colspan="3">Dice (%)</th>
<th align="center" valign="top" colspan="3">Hausdorff 95 (mm)</th>
</tr>
<tr>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">UNet (2015)</td>
<td align="center" valign="top">94.16</td>
<td align="center" valign="top">94.76</td>
<td align="center" valign="top">94.94</td>
<td align="center" valign="top">7.6671</td>
<td align="center" valign="top">1.6476</td>
<td align="center" valign="top">3.8948</td>
</tr>
<tr>
<td align="left" valign="top">PSPNet (2016)</td>
<td align="center" valign="top">92.3</td>
<td align="center" valign="top">92.11</td>
<td align="center" valign="top">94.13</td>
<td align="center" valign="top">3.454</td>
<td align="center" valign="top">2.6466</td>
<td align="center" valign="top">3.4861</td>
</tr>
<tr>
<td align="left" valign="top">DeepLabv3 (2017)</td>
<td align="center" valign="top">92.43</td>
<td align="center" valign="top">93.62</td>
<td align="center" valign="top">94.11</td>
<td align="center" valign="top">7.23</td>
<td align="center" valign="top">1.8961</td>
<td align="center" valign="top">3.9505</td>
</tr>
<tr>
<td align="left" valign="top">UNet++ (2018)</td>
<td align="center" valign="top">93.12</td>
<td align="center" valign="top">94.94</td>
<td align="center" valign="top">95.6</td>
<td align="center" valign="top">7.0516</td>
<td align="center" valign="top">1.4254</td>
<td align="center" valign="top">2.6112</td>
</tr>
<tr>
<td align="left" valign="top">PANet (2018)</td>
<td align="center" valign="top">93.14</td>
<td align="center" valign="top">94.03</td>
<td align="center" valign="top">94.33</td>
<td align="center" valign="top">7.4731</td>
<td align="center" valign="top">1.5437</td>
<td align="center" valign="top">3.7234</td>
</tr>
<tr>
<td align="left" valign="top">CCNet (2020)</td>
<td align="center" valign="top">90.9</td>
<td align="center" valign="top">90.15</td>
<td align="center" valign="top">93.76</td>
<td align="center" valign="top">8.0307</td>
<td align="center" valign="top">2.7355</td>
<td align="center" valign="top">4.7462</td>
</tr>
<tr>
<td align="left" valign="top">MANet (2020)</td>
<td align="center" valign="top">93.83</td>
<td align="center" valign="top">94.52</td>
<td align="center" valign="top">95.91</td>
<td align="center" valign="top">6.8507</td>
<td align="center" valign="top">2.4254</td>
<td align="center" valign="top"><bold>1.8365</bold></td>
</tr>
<tr>
<td align="left" valign="top">BiseNetv2 (2021)</td>
<td align="center" valign="top">89.84</td>
<td align="center" valign="top">84.93</td>
<td align="center" valign="top">92.07</td>
<td align="center" valign="top">8.3099</td>
<td align="center" valign="top">4.3865</td>
<td align="center" valign="top">4.2442</td>
</tr>
<tr>
<td align="left" valign="top">Swin-Unet (2021)</td>
<td align="center" valign="top">88.9</td>
<td align="center" valign="top">89.93</td>
<td align="center" valign="top">93.8</td>
<td align="center" valign="top">10.3311</td>
<td align="center" valign="top">4.7801</td>
<td align="center" valign="top">3.7695</td>
</tr>
<tr>
<td align="left" valign="top">MedT (2021)</td>
<td align="center" valign="top">91.52</td>
<td align="center" valign="top">92.32</td>
<td align="center" valign="top">94.89</td>
<td align="center" valign="top">8.2466</td>
<td align="center" valign="top">3.1183</td>
<td align="center" valign="top">2.8071</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet (2021)</td>
<td align="center" valign="top">94.31</td>
<td align="center" valign="top">94.34</td>
<td align="center" valign="top">95.81</td>
<td align="center" valign="top">6.7094</td>
<td align="center" valign="top">2.3035</td>
<td align="center" valign="top">2.8282</td>
</tr>
<tr>
<td align="left" valign="top">UCTransNet (2022)</td>
<td align="center" valign="top">94.4</td>
<td align="center" valign="top">94.01</td>
<td align="center" valign="top">95.99</td>
<td align="center" valign="top">7.5424</td>
<td align="center" valign="top">2.6062</td>
<td align="center" valign="top">1.9285</td>
</tr>
<tr>
<td align="left" valign="top">Ours</td>
<td align="center" valign="top"><bold>94.88</bold></td>
<td align="center" valign="top"><bold>95.45</bold></td>
<td align="center" valign="top"><bold>96.45</bold></td>
<td align="center" valign="top"><bold>2.541</bold></td>
<td align="center" valign="top"><bold>1.4056</bold></td>
<td align="center" valign="top">1.902</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column, and the coronal, sagittal, and transverse views are represented by cor, sag, and tra, respectively.</p>
</table-wrap-foot>
</table-wrap>
<p>It can be seen that although the Hausdorff 95 is not as good as MANet in the transverse view, our model is generally better than other models through the evaluation of dice and Hausdorff 95 values. Dice values of the coronal view, sagittal view and transverse view reached 94.88, 95.45 and 96.45% respectively; and the Hausdorff 95 values reached 2.541&#x2009;mm, 1.4056&#x2009;mm and 1.902&#x2009;mm, respectively.</p>
</sec>
<sec id="sec17">
<label>4.7.</label>
<title>Ablation experiment</title>
<p>To demonstrate the efficacy of the incorporation module, we performed two groups of ablation experiments based on the principle of &#x201C;fixing two items and changing one item.&#x201D;</p>
<sec id="sec18">
<label>4.7.1.</label>
<title>Ablation experiment of attention module</title>
<p>We examine four different experimental configurations to verify the efficacy of adding attention modules, i.e., TransUNet with ASPP (TransUNet+ASPP) as the baseline, and further with channel attention (TransUNet+ASPP+C), pixel attention (TransUNet+ASPP+P), and CPAT module (TransUNet+ASPP+CPAT). <xref rid="tab3" ref-type="table">Tables 3</xref>, <xref rid="tab4" ref-type="table">4</xref> show the segmentation results for the overall and multiple views, respectively.</p>
<table-wrap position="float" id="tab3">
<label>Table 3</label>
<caption>
<p>Results of ablation experiments with attentional module.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Model</th>
<th align="center" valign="top">Dice (%)</th>
<th align="center" valign="top">Hausdorff 95 (mm)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">TransUNet+ASPP</td>
<td align="center" valign="top">95.18</td>
<td align="center" valign="top">4.2574</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+C</td>
<td align="center" valign="top">95.23</td>
<td align="center" valign="top">3.7512</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+P</td>
<td align="center" valign="top">95.52</td>
<td align="center" valign="top">2.4821</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+CPAT</td>
<td align="center" valign="top"><bold>95.74</bold></td>
<td align="center" valign="top"><bold>1.9476</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="tab4">
<label>Table 4</label>
<caption>
<p>Results of ablation experiments with attentional module from multiple perspectives.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top" rowspan="2">Model</th>
<th align="center" valign="top" colspan="3">Dice (%)</th>
<th align="center" valign="top" colspan="3">Hausdorff 95 (mm)</th>
</tr>
<tr>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">TransUNet+ASPP</td>
<td align="center" valign="top">94.18</td>
<td align="center" valign="top">94.85</td>
<td align="center" valign="top">95.75</td>
<td align="center" valign="top">6.2425</td>
<td align="center" valign="top">1.9845</td>
<td align="center" valign="top">3.8742</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+C</td>
<td align="center" valign="top">94.24</td>
<td align="center" valign="top">94.98</td>
<td align="center" valign="top">95.91</td>
<td align="center" valign="top">5.9475</td>
<td align="center" valign="top">1.4863</td>
<td align="center" valign="top">2.8431</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+P</td>
<td align="center" valign="top">94.66</td>
<td align="center" valign="top">95.05</td>
<td align="center" valign="top">96.34</td>
<td align="center" valign="top">5.7424</td>
<td align="center" valign="top">1.8574</td>
<td align="center" valign="top">1.9527</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet+ASPP+CPAT</td>
<td align="center" valign="top"><bold>94.88</bold></td>
<td align="center" valign="top"><bold>95.45</bold></td>
<td align="center" valign="top"><bold>96.45</bold></td>
<td align="center" valign="top"><bold>2.541</bold></td>
<td align="center" valign="top"><bold>1.4056</bold></td>
<td align="center" valign="top"><bold>1.902</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column, and the coronal, sagittal, and transverse views are represented by cor, sag, and tra, respectively.</p>
</table-wrap-foot>
</table-wrap>
<p>From <xref rid="tab3" ref-type="table">Tables 3</xref>, <xref rid="tab4" ref-type="table">4</xref>, we have some observations as follows.</p>
<list list-type="order">
<list-item>
<p>When we added channel attention to &#x201C;TransUNet+ASPP,&#x201D; not only the Dice and Hausdorff 95 of &#x201C;TransUNet+ASPP+C&#x201D; in <xref rid="tab3" ref-type="table">Table 3</xref> improved by 0.05% and 0.5062&#x2009;mm, respectively, but also the experimental results of multiple views in <xref rid="tab4" ref-type="table">Table 4</xref> were better than those of &#x201C;TransUNet+ASPP&#x201D;&#xFF0C; which proves the effectiveness of adding channel attention.</p>
</list-item>
<list-item>
<p>When we added pixel attention to &#x201C;TransUNet+ASPP,&#x201D; the Dice and Hausdorff 95 of &#x201C;TransUNet+ASPP+P&#x201D; in <xref rid="tab3" ref-type="table">Table 3</xref> were 95.52% and 2.4821&#x2009;mm, respectively, and the experimental results in <xref rid="tab4" ref-type="table">Table 4</xref> were also improved significantly, thereby proving that the addition of pixel attention is effective.</p>
</list-item>
<list-item>
<p>The results of &#x201C;TransUNet+ASPP+CPAT&#x201D; in <xref rid="tab3" ref-type="table">Tables 3</xref>, <xref rid="tab4" ref-type="table">4</xref> are significantly better than those of &#x201C;TransUNet+ASPP+C&#x201D; and &#x201C;TransUNet+ASPP+P,&#x201D; demonstrating that the sequential connection of channel attention and pixel attention is better than using either attention module.</p>
</list-item>
</list>
</sec>
<sec id="sec19">
<label>4.7.2.</label>
<title>Ablation experiment of ASPP module</title>
<p>To demonstrate the efficacy of the ASPP module, two different experimental configurations were studied, i.e., TransUNet with CPAT (TransUNet +CPAT) as a baseline and further addition of the ASPP module (TransUNet +CPAT+ASPP). <xref rid="tab5" ref-type="table">Tables 5</xref>, <xref rid="tab6" ref-type="table">6</xref> show the segmentation results for the overall and multiple views, respectively. As can be seen from <xref rid="tab5" ref-type="table">Table 5</xref>, the addition of the ASPP module improves the &#x201C;TransUNet+CPAT+ASPP&#x201D; Dice and Hausdorff 95 by 0.2% and 1.5166&#x2009;mm, respectively. In addition, according to <xref rid="tab6" ref-type="table">Table 6</xref>, Hausdorff 95 with &#x201C;TransUNet +CPAT+ASPP&#x201D; is excellent in other views, although it is lower than &#x201C;TransUNet +CPAT&#x201D; in the transverse view. The above results prove the efficiency of ASPP module.</p>
<table-wrap position="float" id="tab5">
<label>Table 5</label>
<caption>
<p>Results of ablation experiments with ASPP module.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top">Model</th>
<th align="center" valign="top">Dice (%)</th>
<th align="center" valign="top">Hausdorff 95 (mm)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">TransUNet +CPAT</td>
<td align="center" valign="top">95.54</td>
<td align="center" valign="top">3.4642</td>
</tr>
<tr>
<td align="left" valign="top">TransUNet +CPAT+ASPP</td>
<td align="center" valign="top"><bold>95.74</bold></td>
<td align="center" valign="top"><bold>1.9476</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="tab6">
<label>Table 6</label>
<caption>
<p>Results of ablation experiments with ASPP module from multiple perspectives.</p>
</caption>
<table frame="hsides" rules="groups">
<thead>
<tr>
<th align="left" valign="top" rowspan="2">Model</th>
<th align="center" valign="top" colspan="3">Dice (%)</th>
<th align="center" valign="top" colspan="3">Hausdorff 95 (mm)</th>
</tr>
<tr>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
<th align="center" valign="top">cor</th>
<th align="center" valign="top">sag</th>
<th align="center" valign="top">tra</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left" valign="top">TransUNet +CPAT</td>
<td align="center" valign="top">94.68</td>
<td align="center" valign="top">95.34</td>
<td align="center" valign="top">96.23</td>
<td align="center" valign="top">7.0574</td>
<td align="center" valign="top">1.4682</td>
<td align="center" valign="top"><bold>1.8472</bold></td>
</tr>
<tr>
<td align="left" valign="top">TransUNet +CPAT+ASPP</td>
<td align="center" valign="top"><bold>94.88</bold></td>
<td align="center" valign="top"><bold>95.45</bold></td>
<td align="center" valign="top"><bold>96.45</bold></td>
<td align="center" valign="top"><bold>2.541</bold></td>
<td align="center" valign="top"><bold>1.4056</bold></td>
<td align="center" valign="top">1.902</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>Bold font is the best data for each column, and the coronal, sagittal, and transverse views are represented by cor, sag, and tra, respectively.</p>
</table-wrap-foot>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="sec20" sec-type="discussions">
<label>5.</label>
<title>Discussion</title>
<p>At present, the results of Dice and Hausdorff distance of our model in acoustic neuroma segmentation have reached our expectations. Given the fact that acoustic neuromas vary in shape--some with irregular shape and growing into the inner auditory canal, while some with clear boundary, we need a larger receptive field to extract the feature of acoustic neuromas. As ordinary convolution structure cannot fully extract features, we added the ASPP module. Furthermore, since acoustic neuromas often occur in the cerebellopontine angle area with relatively fixed position, we intended to make our model automatically learn the weights at different scales by adding the attention mechanism. Therefore, we added the channel attention and pixel attention in the up-sampling, so that the channel information and pixel information are combined to better explore the channel characteristics and pixel characteristics while restoring the original input image size. In the comparison experiments, we can see that most of the networks with the added Transformer structure achieve good results in segmentation of acoustic neuromas, for example, the Dice value of these networks is almost equal to that of ACP-TransUNet. However, Hausdorff 95 cannot be comparable to ACP-TransUNet. which is due to Transformer&#x2019;s inadequacy to capture low-level details and its limited positioning ability. Given that, we combined ASPP and attention mechanism to make up for this deficiency. In the ablation experiment, it is observed that the segmentation performance of the model becomes better and better with the addition of ASPP and CPAT modules, proving the effectiveness of our choice to add the modules.</p>
<p>However, there are still problems existing in the current work. For example, in the multi-view evaluation, we did not achieve desirable segmentation results in the transverse view. The Hausdorff 95 value of our model in the transverse view is 1.902&#x2009;mm. That figure is inferior to the MANet, which reached 1.8365&#x2009;mm in the comparison experimental. The reasons we believe are of two aspects. First, it could be explained by the relatively low importance of channel weight in the down-sampling of acoustic neuromas in the transverse view direction. But the addition of pixel attention could make all the pixels of the feature map generate attention coefficient, which makes up for the disadvantage of using channel attention alone. Second, although the addition of ASPP module would increase the receptive field, making each convolution output contain a large range of information, the information of smaller tumors in the transverse view could be lost. Given that, in our future work, we will gradually increase the dataset and study the performance changes when increasing or decreasing the single direction module. In addition, our current research task is to achieve accurate segmentation of acoustic neuromas. We hope that the application of ACP-TransUNet will not be limited to acoustic neuromas, so its effectiveness in segmenting other medical images will also be the focus of our future experimental research.</p>
<p>In our research work, the improvement of the accuracy of acoustic neuroma segmentation means that we need to abandon some indicators in some aspects. We have considered trade-offs in these issues. First, the addition of ASPP module, attention mechanism and deeper transformer layer means longer training time and larger model parameters. We believe that the medical segmentation task is different from other segmentation tasks that pursue timeliness (such as face segmentation). Between lightweight and precision, we prefer the latter. Second, since Transformer lacks the inductive bias of convolution, it requires more sample size than CNN. Transformer needs to learn this kind of information from a large amount of data. Considering the precious resources and insufficient data support of current medical images, instead of choosing to train from scratch, we resort to pre-trained models to achieve the same or even better performance than CNN. In the future, we will conduct research for Transformer on small-scale datasets.</p>
</sec>
<sec id="sec21" sec-type="conclusions">
<label>6.</label>
<title>Conclusion</title>
<p>In this paper, we proposed a novel model named ACP-TransUNet based on the improved TransUNet structure, with all the data on the basis of MRI images. Through deep learning, we realized the automatic and accurate segmentation of acoustic neuromas in the cerebellopontine angle region. Dice and Hausdorff 95 reached 95.74% and 1.9476&#x2009;mm respectively, and the dividing boundary was closer to the gold standard. The overall effect of segmentation was significantly improved, which was valuable for clinical application and auxiliary physician diagnosis. With decreased intervention of human factors, we greatly improved the diagnostic efficiency and reliability. In addition, the ASPP module was introduced into ACP-TransUNet, which not only increases the receptive field and obtains multi-scale and multi-resolution background information, but also makes the features contained in the sequence of the imported Transformer more accurate and significant. The CPAT module with sequential channel attention and pixel attention is added to the upsampling process so that channel information and pixel information are combined to improve model performance and accuracy by weighting important features. The experimental results show that our model can effectively segment acoustic neuroma. Compared with other methods, the proposed method has different degrees of performance improvement in the segmentation of acoustic neuroma.</p>
</sec>
<sec id="sec22" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="sec23">
<title>Ethics statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. The patients/participants provided their written informed consent to participate in this study.</p>
</sec>
<sec id="sec24">
<title>Author contributions</title>
<p>ZZ and HB: conceptualization. HB: methodology, formal analysis, supervision, and funding acquisition. ZZ: software, data, writing original draft preparation, and visualization. ZZ, HB, and QM: validation. JL and XZ: investigation. YY: resources. HB, QM, and CZ: writing review and editing. QM: project administration. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="sec25" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported in part by National Natural Science Foundation of China (61201106) and Tianjin Research Innovation Project for Postgraduate Students (2022SKY126).</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="sec100" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="ref1"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Ahn</surname> <given-names>N</given-names></name> <name><surname>Kang</surname> <given-names>B</given-names></name> <name><surname>Sohn</surname> <given-names>KA</given-names></name></person-group> (<year>2018</year>). Fast, accurate, and lightweight super-resolution with cascading residual network.</citation></ref>
<ref id="ref2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Beauchemin</surname> <given-names>M.</given-names></name> <name><surname>Thomson</surname> <given-names>K. P.</given-names></name> <name><surname>Edwards</surname> <given-names>G.</given-names></name></person-group> (<year>1998</year>). <article-title>On the Hausdorff distance used for the evaluation of segmentation results</article-title>. <source>Can. J. Remote. Sens.</source> <volume>24</volume>, <fpage>3</fpage>&#x2013;<lpage>8</lpage>. doi: <pub-id pub-id-type="doi">10.1080/07038992.1998.10874685</pub-id></citation></ref>
<ref id="ref3"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Cao</surname> <given-names>H</given-names></name> <name><surname>Wang</surname> <given-names>Y</given-names></name> <name><surname>Chen</surname> <given-names>J</given-names></name> <name><surname>Jiang</surname> <given-names>D</given-names></name> <name><surname>Zhang</surname> <given-names>X</given-names></name> <name><surname>Tian</surname> <given-names>Q</given-names></name> <name><surname>Wang</surname> <given-names>M</given-names></name></person-group> (<year>2021</year>). Swin-Unet: Unet-like pure transformer for medical image segmentation.</citation></ref>
<ref id="ref4"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>B</given-names></name> <name><surname>Liu</surname> <given-names>Y</given-names></name> <name><surname>Zhang</surname> <given-names>Z</given-names></name> <name><surname>Lu</surname> <given-names>G</given-names></name> <name><surname>Zhang</surname> <given-names>D</given-names></name></person-group> (<year>2021</year>). TransAttUnet: Multi-level attention-guided U-net with transformer for medical image segmentation. arXiv.</citation></ref>
<ref id="ref5"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>J</given-names></name> <name><surname>Lu</surname> <given-names>Y</given-names></name> <name><surname>Yu</surname> <given-names>Q</given-names></name> <name><surname>Luo</surname> <given-names>X</given-names></name> <name><surname>Zhou</surname> <given-names>Y</given-names></name></person-group> (<year>2021</year>). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv [Preprint].</citation></ref>
<ref id="ref6"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>LC</given-names></name> <name><surname>Papandreou</surname> <given-names>G</given-names></name> <name><surname>Kokkinos</surname> <given-names>I</given-names></name> <name><surname>Murphy</surname> <given-names>K</given-names></name> <name><surname>Yuille</surname> <given-names>AL</given-names></name></person-group> (<year>2014</year>). Semantic image segmentation with deep convolutional nets and fully connected CRFs. Computer science: 357&#x2013;361.</citation></ref>
<ref id="ref7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L. C.</given-names></name> <name><surname>Papandreou</surname> <given-names>G.</given-names></name> <name><surname>Kokkinos</surname> <given-names>I.</given-names></name> <name><surname>Murphy</surname> <given-names>K.</given-names></name> <name><surname>Yuille</surname> <given-names>A. L.</given-names></name></person-group> (<year>2018b</year>). <article-title>DeepLab: semantic image segmentation with deep convolutional nets, Atrous convolution, and fully connected CRFs</article-title>. <source>IEEE Trans. Pattern Anal. Mach. Intell.</source> <volume>40</volume>, <fpage>834</fpage>&#x2013;<lpage>848</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2699184</pub-id>, PMID: <pub-id pub-id-type="pmid">28463186</pub-id></citation></ref>
<ref id="ref8"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>LC</given-names></name> <name><surname>Papandreou</surname> <given-names>G</given-names></name> <name><surname>Schroff</surname> <given-names>F</given-names></name> <name><surname>Adam</surname> <given-names>H</given-names></name></person-group> (<year>2017</year>). Rethinking Atrous convolution for semantic image segmentation.</citation></ref>
<ref id="ref9"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>LC</given-names></name> <name><surname>Zhu</surname> <given-names>Y</given-names></name> <name><surname>Papandreou</surname> <given-names>G</given-names></name> <name><surname>Schroff</surname> <given-names>F</given-names></name> <name><surname>Adam</surname> <given-names>H</given-names></name></person-group>, (<year>2018a</year>). Encoder-decoder with Atrous separable convolution for semantic image segmentation, European Conference on Computer Vision.</citation></ref>
<ref id="ref10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Corso</surname> <given-names>J. J.</given-names></name> <name><surname>Sharon</surname> <given-names>E.</given-names></name> <name><surname>Dube</surname> <given-names>S.</given-names></name> <name><surname>El-Saden</surname> <given-names>S.</given-names></name> <name><surname>Sinha</surname> <given-names>U.</given-names></name> <name><surname>Yuille</surname> <given-names>A.</given-names></name></person-group> (<year>2008</year>). <article-title>Efficient multilevel brain tumor segmentation with integrated Bayesian model classification</article-title>. <source>IEEE Trans. Med. Imaging</source> <volume>27</volume>, <fpage>629</fpage>&#x2013;<lpage>640</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2007.912817</pub-id>, PMID: <pub-id pub-id-type="pmid">18450536</pub-id></citation></ref>
<ref id="ref11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cuixia</surname> <given-names>L.</given-names></name> <name><surname>Mingqiang</surname> <given-names>L.</given-names></name> <name><surname>Zhaoying</surname> <given-names>B.</given-names></name> <name><surname>Wenbing</surname> <given-names>L.</given-names></name> <name><surname>Dong</surname> <given-names>Z.</given-names></name> <name><surname>Jianhua</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Establishment of a deep feature-based classification model for distinguishing benign and malignant breast tumors on full-filed digital mammography</article-title>. <source>J. South Med. Univ</source> <volume>39</volume>, <fpage>88</fpage>&#x2013;<lpage>92</lpage>. doi: <pub-id pub-id-type="doi">10.12122/j.issn.1673-4254.2019.01.14</pub-id></citation></ref>
<ref id="ref12"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Devlin</surname> <given-names>J</given-names></name> <name><surname>Chang</surname> <given-names>MW</given-names></name> <name><surname>Lee</surname> <given-names>K</given-names></name> <name><surname>Toutanova</surname> <given-names>K</given-names></name></person-group> (<year>2018</year>). BERT: pre-training of deep bidirectional transformers for language understanding.</citation></ref>
<ref id="ref13"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>F.</given-names></name> <name><surname>Mo</surname> <given-names>Y.</given-names></name> <name><surname>Guo</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <source>Automatic brain tumor detection and segmentation using U-net based fully convolutional networks, annual conference on medical image understanding and analysis</source>. <publisher-loc>Edinburgh, UK</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>506</fpage>&#x2013;<lpage>517</lpage>.</citation></ref>
<ref id="ref14"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Dosovitskiy</surname> <given-names>A</given-names></name> <name><surname>Beyer</surname> <given-names>L</given-names></name> <name><surname>Kolesnikov</surname> <given-names>A</given-names></name> <name><surname>Weissenborn</surname> <given-names>D</given-names></name> <name><surname>Zhai</surname> <given-names>X</given-names></name> <name><surname>Unterthiner</surname> <given-names>T</given-names></name> <name><surname>Dehghani</surname> <given-names>M</given-names></name> <name><surname>Minderer</surname> <given-names>M</given-names></name> <etal/></person-group>., (<year>2021</year>). &#x201C;An image is worth 16x16 words: Transformers for image recognition at scale.&#x201D; in <italic>International Conference on Learning Representations</italic>.</citation></ref>
<ref id="ref15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>W.</given-names></name> <name><surname>Rao</surname> <given-names>N.</given-names></name> <name><surname>Yong</surname> <given-names>J.</given-names></name> <name><surname>Adjei</surname> <given-names>P. E.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2023</year>). <article-title>Early gastric cancer segmentation in gastroscopic images using a co-spatial attention and channel attention based triple-branch ResUnet</article-title>. <source>Comput. Methods Prog. Biomed.</source> <volume>231</volume>:<fpage>107397</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2023.107397</pub-id>, PMID: <pub-id pub-id-type="pmid">36753915</pub-id></citation></ref>
<ref id="ref16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fan</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>MA-net: a multi-scale attention network for liver and tumor segmentation</article-title>. <source>IEEE Access</source> <volume>8</volume>, <fpage>179656</fpage>&#x2013;<lpage>179665</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3025372</pub-id></citation></ref>
<ref id="ref17"><citation citation-type="book"><person-group person-group-type="author"><name><surname>He</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Ren</surname> <given-names>S.</given-names></name> <name><surname>Sun</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <source>Deep residual learning for image recognition</source>. <publisher-loc>Las Vegas, NV, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>.</citation></ref>
<ref id="ref18"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>G.</given-names></name> <name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Laurens</surname> <given-names>V.</given-names></name> <name><surname>Weinberger</surname> <given-names>K. Q.</given-names></name></person-group> (<year>2016</year>). <source>Densely connected convolutional networks</source>. <publisher-loc>Honolulu, HI, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>.</citation></ref>
<ref id="ref19"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Z</given-names></name> <name><surname>Wang</surname> <given-names>X</given-names></name> <name><surname>Wei</surname> <given-names>Y</given-names></name> <name><surname>Huang</surname> <given-names>L</given-names></name> <name><surname>Huang</surname> <given-names>TS</given-names></name></person-group> (<year>2020</year>). CCNet: Criss-cross attention for semantic segmentation. IEEE transactions on pattern analysis and machine intelligence PP:1.</citation></ref>
<ref id="ref20"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Huttenlocher</surname> <given-names>D. P</given-names></name> <name><surname>Klanderman</surname> <given-names>G. A</given-names></name> <name><surname>Rucklidge</surname> <given-names>W. J</given-names></name></person-group> (<year>1993</year>), <article-title>Comparing images using the Hausdorff distance</article-title>. <source>Pattern analysis and machine intelligence, IEEE transactions</source> on, <volume>15</volume>, 850, 863, doi: <pub-id pub-id-type="doi">10.1109/34.232073</pub-id>.</citation></ref>
<ref id="ref21"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Jie</surname> <given-names>Shen</given-names></name> <name><surname>Samuel</surname> <given-names>Albanie</given-names></name> <name><surname>Gang</surname> <given-names>Sun</given-names></name> <name><surname>Enhua</surname></name></person-group> (<year>2019</year>). Squeeze-and-excitation networks. IEEE transactions on pattern analysis and machine intelligence.</citation></ref>
<ref id="ref22"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D</given-names></name> <name><surname>Ba</surname> <given-names>J</given-names></name></person-group> (<year>2014</year>). Adam: A Method for Stochastic Optimization. Computer Science.</citation></ref>
<ref id="ref23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krizhevsky</surname> <given-names>A.</given-names></name> <name><surname>Sutskever</surname> <given-names>I.</given-names></name> <name><surname>Hinton</surname> <given-names>G. E.</given-names></name></person-group> (<year>2017</year>). <article-title>ImageNet classification with deep convolutional neural networks</article-title>. <source>Commun. ACM</source> <volume>60</volume>, <fpage>84</fpage>&#x2013;<lpage>90</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3065386</pub-id></citation></ref>
<ref id="ref24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lin</surname> <given-names>A.</given-names></name> <name><surname>Chen</surname> <given-names>B.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Lu</surname> <given-names>G.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name></person-group> (<year>2022</year>). <article-title>Ds-TransUNet: dual swin transformer u-net for medical image segmentation</article-title>. <source>IEEE Trans. Instrum. Meas.</source> <volume>71</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIM.2022.3178991</pub-id></citation></ref>
<ref id="ref25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ling</surname> <given-names>C.</given-names></name> <name><surname>Chao</surname> <given-names>Z.</given-names></name> <name><surname>Mengling</surname> <given-names>T.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Qingxiang</surname> <given-names>L.</given-names></name></person-group> (<year>2016</year>). <article-title>Effect analysis of MRI in differential diagnosis of cerebellopontine angle meningioma and acoustic neuroma</article-title>. <source>Contemp. Med. Symp.</source> <volume>14</volume>, <fpage>134</fpage>&#x2013;<lpage>136</lpage>. doi: CNKI:SUN:QYWA.0.2016-22-093</citation></ref>
<ref id="ref26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lingmei</surname> <given-names>A.</given-names></name> <name><surname>Tiandong</surname> <given-names>L.</given-names></name> <name><surname>Fuyuan</surname> <given-names>L.</given-names></name> <name><surname>Kangzhen</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>Magnetic resonance brain tumor image segmentation based on attention U-net</article-title>. <source>Laser Optoelectronics Progress</source> <volume>57</volume>, <fpage>141030</fpage>&#x2013;<lpage>141286</lpage>. doi: <pub-id pub-id-type="doi">10.3788/LOP57.141030</pub-id></citation></ref>
<ref id="ref27"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z</given-names></name> <name><surname>Chen</surname> <given-names>L</given-names></name> <name><surname>Tong</surname> <given-names>L</given-names></name> <name><surname>Zhou</surname> <given-names>F</given-names></name> <name><surname>Jiang</surname> <given-names>Z</given-names></name> <name><surname>Zhang</surname> <given-names>Q</given-names></name> <name><surname>Shan</surname> <given-names>C</given-names></name> <name><surname>Wang</surname> <given-names>Y</given-names></name> <etal/></person-group>. (<year>2020</year>). Deep learning based brain tumor segmentation: A survey.</citation></ref>
<ref id="ref28"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>S</given-names></name> <name><surname>Qi</surname> <given-names>L</given-names></name> <name><surname>Qin</surname> <given-names>H</given-names></name> <name><surname>Shi</surname> <given-names>J</given-names></name> <name><surname>Jia</surname> <given-names>J</given-names></name></person-group> (<year>2018</year>). Path aggregation network for instance segmentation. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).</citation></ref>
<ref id="ref29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Huangliang</surname> <given-names>K.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name></person-group> (<year>2022</year>). <article-title>TransUNet+: redesigning the skip connection to enhance features in medical image segmentation</article-title>. <source>Knowl.-Based Syst.</source> <volume>256</volume>:<fpage>109859</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2022.109859</pub-id></citation></ref>
<ref id="ref30"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>Z.</given-names></name> <name><surname>Lin</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>H.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <etal/></person-group>. (<year>2021</year>). &#x201C;Swin transformer: hierarchical vision transformer using shifted windows.&#x201D; in <italic>IEEE/CVF International Conference on Computer Vision (ICCV)</italic>. pp. 9992&#x2013;10002.</citation></ref>
<ref id="ref31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McClelland</surname> <given-names>S.</given-names></name> <name><surname>Guo</surname> <given-names>H.</given-names></name> <name><surname>Okuyemi</surname> <given-names>K. S.</given-names></name></person-group> (<year>2011</year>). <article-title>Morbidity and mortality following acoustic neuroma excision in the United States: analysis of racial disparities during a decade in the radiosurgery era</article-title>. <source>Neuro-Oncology</source> <volume>13</volume>, <fpage>1252</fpage>&#x2013;<lpage>1259</lpage>. doi: <pub-id pub-id-type="doi">10.1093/neuonc/nor118</pub-id>, PMID: <pub-id pub-id-type="pmid">21856684</pub-id></citation></ref>
<ref id="ref32"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Mehta</surname> <given-names>R</given-names></name></person-group> (<year>2015</year>). Introducing dice, Jaccard, and other label overlap measures to ITK.</citation></ref>
<ref id="ref33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>North</surname> <given-names>M.</given-names></name> <name><surname>Weishaar</surname> <given-names>J.</given-names></name> <name><surname>Nuru</surname> <given-names>M.</given-names></name> <name><surname>Anderson</surname> <given-names>D.</given-names></name> <name><surname>Leonetti</surname> <given-names>J. P.</given-names></name></person-group> (<year>2022</year>). <article-title>Assessing surgical approaches for acoustic neuroma resection: do patients perceive a difference in quality-of-life outcomes?</article-title> <source>Otol. Neurotol.</source> <volume>43</volume>, <fpage>1245</fpage>&#x2013;<lpage>1251</lpage>. doi: <pub-id pub-id-type="doi">10.1097/MAO.0000000000003720</pub-id>, PMID: <pub-id pub-id-type="pmid">36351229</pub-id></citation></ref>
<ref id="ref34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nur&#x00E7;in</surname> <given-names>F. V.</given-names></name></person-group> (<year>2022</year>). <article-title>Improved segmentation of overlapping red blood cells on malaria blood smear images with TransUNet architecture</article-title>. <source>Int. J. Imaging Syst. Technol.</source> <volume>32</volume>, <fpage>1673</fpage>&#x2013;<lpage>1680</lpage>. doi: <pub-id pub-id-type="doi">10.1002/ima.22739</pub-id></citation></ref>
<ref id="ref35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pan</surname> <given-names>S. M.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Xie</surname> <given-names>N. D.</given-names></name> <name><surname>Chong</surname> <given-names>Y. W.</given-names></name></person-group> (<year>2023</year>). <article-title>EG-TransUNet: a transformer-based U-net with enhanced and guided models for biomedical image segmentation</article-title>. <source>BMC Bioinform.</source> <volume>24</volume>:<fpage>85</fpage>. doi: <pub-id pub-id-type="doi">10.1186/s12859-023-05196-1</pub-id>, PMID: <pub-id pub-id-type="pmid">36882688</pub-id></citation></ref>
<ref id="ref36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Qiulin</surname> <given-names>J.</given-names></name> <name><surname>Xin</surname> <given-names>W.</given-names></name></person-group> (<year>2018</year>). <article-title>Brain tumor image segmentation based on region growing algorithm</article-title>. <source>J. Changchun Univ. Technol.</source> <volume>39</volume>, <fpage>490</fpage>&#x2013;<lpage>493</lpage>. doi: CNKI:SUN:JLGX.0.2018-05-013</citation></ref>
<ref id="ref37"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Rajeshwari</surname> <given-names>P</given-names></name> <name><surname>Shyamala</surname> <given-names>K</given-names></name></person-group>. (<year>2023</year>). &#x201C;Pixel attention based deep neural network for chest CT image super resolution.&#x201D; in <italic>Advanced Network Technologies and Intelligent Computing: Second International Conference (ANTIC)</italic>. pp. 393&#x2013;407.</citation></ref>
<ref id="ref38"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Ronneberger</surname> <given-names>O</given-names></name> <name><surname>Fischer</surname> <given-names>P</given-names></name> <name><surname>Brox</surname> <given-names>T</given-names></name></person-group> (<year>2015</year>). U-net: Convolutional networks for biomedical image segmentation. ArXiv abs/1505.04597.</citation></ref>
<ref id="ref39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roy</surname> <given-names>K.</given-names></name> <name><surname>Banik</surname> <given-names>D.</given-names></name> <name><surname>Bhattacharjee</surname> <given-names>D.</given-names></name> <name><surname>Krejcar</surname> <given-names>O.</given-names></name> <name><surname>Kollmann</surname> <given-names>C.</given-names></name></person-group> (<year>2022</year>). <article-title>LwMLA-NET: a lightweight multi-level attention-based network for segmentation of COVID-19 lungs abnormalities from CT images</article-title>. <source>IEEE Trans. Instrum. Meas.</source> <volume>71</volume>, <fpage>1</fpage>&#x2013;<lpage>13</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIM.2022.3161690</pub-id></citation></ref>
<ref id="ref40"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Russo</surname> <given-names>C</given-names></name> <name><surname>Liu</surname> <given-names>S</given-names></name> <name><surname>Di Ieva</surname> <given-names>A</given-names></name></person-group> (<year>2020</year>). Spherical coordinates transformation pre-processing in deep convolution neural networks for brain tumor segmentation in MRI. arXiv preprint arXiv:200807090.</citation></ref>
<ref id="ref41"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Shelhamer</surname> <given-names>E</given-names></name> <name><surname>Long</surname> <given-names>J</given-names></name> <name><surname>Darrell</surname> <given-names>T</given-names></name></person-group> (<year>2016</year>). Fully convolutional networks for semantic segmentation.</citation></ref>
<ref id="ref42"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Simonyan</surname> <given-names>K</given-names></name> <name><surname>Zisserman</surname> <given-names>A</given-names></name></person-group> (<year>2014</year>). Very deep convolutional networks for large-scale image recognition. Computer Science.</citation></ref>
<ref id="ref43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>Z.</given-names></name> <name><surname>Qiu</surname> <given-names>D.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name> <name><surname>Hui</surname> <given-names>Y.</given-names></name></person-group> (<year>2023</year>). <article-title>Channel attention generative adversarial network for super-resolution of glioma magnetic resonance image</article-title>. <source>Comput. Methods Prog. Biomed.</source> <volume>229</volume>:<fpage>107255</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2022.107255</pub-id>, PMID: <pub-id pub-id-type="pmid">36462426</pub-id></citation></ref>
<ref id="ref44"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Soomro</surname> <given-names>MH</given-names></name> <name><surname>De Cola</surname> <given-names>G</given-names></name> <name><surname>Conforto</surname> <given-names>S</given-names></name> <name><surname>Schmid</surname> <given-names>M</given-names></name> <name><surname>Giunta</surname> <given-names>G</given-names></name> <name><surname>Guidi</surname> <given-names>E</given-names></name> <name><surname>Neri</surname> <given-names>E</given-names></name> <name><surname>Caruso</surname> <given-names>D</given-names></name> <etal/></person-group>., (<year>2018</year>). &#x201C;Automatic segmentation of colorectal cancer in 3D MRI by combining deep learning and 3D level-set algorithm-a preliminary study.&#x201D; in <italic>2018 IEEE 4th Middle East Conference on Biomedical Engineering (MECBME), IEEE.</italic> pp. 198&#x2013;203.</citation></ref>
<ref id="ref45"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Szegedy</surname> <given-names>C.</given-names></name> <name><surname>Liu</surname> <given-names>W.</given-names></name> <name><surname>Jia</surname> <given-names>Y.</given-names></name> <name><surname>Sermanet</surname> <given-names>P.</given-names></name> <name><surname>Rabinovich</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <source>Going deeper with convolutions</source>. <publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>.</citation></ref>
<ref id="ref46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname> <given-names>J.</given-names></name> <name><surname>Zou</surname> <given-names>B.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name> <name><surname>Feng</surname> <given-names>S.</given-names></name> <name><surname>Peng</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Plane-wave image reconstruction via generative adversarial network and attention mechanism</article-title>. <source>IEEE Trans. Instrum. Meas.</source> <volume>70</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TIM.2021.3087819</pub-id></citation></ref>
<ref id="ref47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Thillaikkarasi</surname> <given-names>R.</given-names></name> <name><surname>Saravanan</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>An enhancement of deep learning algorithm for brain tumor segmentation using kernel based CNN with M-SVM</article-title>. <source>J. Med. Syst.</source> <volume>43</volume>, <fpage>1</fpage>&#x2013;<lpage>7</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s10916-019-1223-7</pub-id></citation></ref>
<ref id="ref48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>C.</given-names></name> <name><surname>Xu</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Zuo</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>Attention-guided CNN for image denoising</article-title>. <source>Neural Netw.</source> <volume>124</volume>, <fpage>117</fpage>&#x2013;<lpage>129</lpage>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2019.12.024</pub-id>, PMID: <pub-id pub-id-type="pmid">31991307</pub-id></citation></ref>
<ref id="ref49"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Valanarasu</surname> <given-names>JMJ</given-names></name> <name><surname>Oza</surname> <given-names>P</given-names></name> <name><surname>Hacihaliloglu</surname> <given-names>I</given-names></name> <name><surname>Patel</surname> <given-names>VM</given-names></name></person-group>, (<year>2021</year>). &#x201C;Medical transformer: gated axial-attention for medical image segmentation.&#x201D; in <italic>International Conference on Medical Image Computing and Computer-Assisted Intervention</italic>.</citation></ref>
<ref id="ref50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vaswani</surname> <given-names>A.</given-names></name> <name><surname>Shazeer</surname> <given-names>N.</given-names></name> <name><surname>Parmar</surname> <given-names>N.</given-names></name> <name><surname>Uszkoreit</surname> <given-names>J.</given-names></name> <name><surname>Jones</surname> <given-names>L.</given-names></name> <name><surname>Gomez</surname> <given-names>A. N.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>Attention is all you need</article-title>. <source>Adv. Neural Inf. Proces. Syst.</source> <fpage>6000</fpage>&#x2013;<lpage>6010</lpage>. doi: <pub-id pub-id-type="doi">10.48550/arXiv.1706.03762</pub-id></citation></ref>
<ref id="ref51"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>H</given-names></name> <name><surname>Cao</surname> <given-names>P</given-names></name> <name><surname>Wang</surname> <given-names>J</given-names></name> <name><surname>Zaiane</surname> <given-names>OR</given-names></name></person-group> (<year>2021</year>). UCTransNet: Rethinking the skip connections in U-net from a channel-wise perspective with transformer.</citation></ref>
<ref id="ref52"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X</given-names></name> <name><surname>Girshick</surname> <given-names>R</given-names></name> <name><surname>Gupta</surname> <given-names>A</given-names></name> <name><surname>He</surname> <given-names>K</given-names></name></person-group> (<year>2017</year>). Non-local neural networks.</citation></ref>
<ref id="ref53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Wang</surname> <given-names>F.</given-names></name> <name><surname>Dong</surname> <given-names>P.</given-names></name> <name><surname>Li</surname> <given-names>C.</given-names></name></person-group> (<year>2022</year>). <article-title>Multiscale transunet++: dense hybrid U-net with transformer for medical image segmentation</article-title>. <source>SIViP</source> <volume>16</volume>, <fpage>1607</fpage>&#x2013;<lpage>1614</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11760-021-02115-w</pub-id></citation></ref>
<ref id="ref54"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Woo</surname> <given-names>S</given-names></name> <name><surname>Park</surname> <given-names>J</given-names></name> <name><surname>Lee</surname> <given-names>JY</given-names></name> <name><surname>Kweon</surname> <given-names>IS</given-names></name></person-group> (<year>2018</year>). <source>CBAM: Convolutional block attention module</source>. <publisher-name>Springer</publisher-name>, <publisher-loc>Cham</publisher-loc>.</citation></ref>
<ref id="ref55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiaobo</surname> <given-names>L.</given-names></name> <name><surname>Maosheng</surname> <given-names>X.</given-names></name> <name><surname>Xiaomei</surname> <given-names>X.</given-names></name></person-group> (<year>2019</year>). <article-title>Automatic segmentation for Glioblastoma Multiforme using multimodal MR images and multiple features</article-title>. <source>J. Comp.Aided Design Comp. Graphics</source> <volume>31</volume>, <fpage>421</fpage>&#x2013;<lpage>430</lpage>. doi: CNKI:SUN:JSJF.0.2019-03-008</citation></ref>
<ref id="ref56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xiaoxia</surname> <given-names>P.</given-names></name> <name><surname>Qian</surname> <given-names>M.</given-names></name> <name><surname>Xia</surname> <given-names>T.</given-names></name> <name><surname>Ziwei</surname> <given-names>L.</given-names></name> <name><surname>Daohai</surname> <given-names>X.</given-names></name></person-group> (<year>2014</year>). <article-title>MRI findings of lesions in the cerebellopontine angle</article-title>. <source>J. Med. Imaging</source> <volume>24</volume>, <fpage>12</fpage>&#x2013;<lpage>15</lpage>.</citation></ref>
<ref id="ref57"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>Y</given-names></name> <name><surname>Mehrkanoon</surname> <given-names>S</given-names></name></person-group> (<year>2022</year>). AA-TransUNet: Attention augmented TransUNet for Nowcasting tasks.</citation></ref>
<ref id="ref58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yongzhuo</surname> <given-names>L.</given-names></name> <name><surname>Shuguang</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>Application of improved watershed algorithm in segmentation of brain tumor CT images</article-title>. <source>Software Guide</source> <volume>17</volume>, <fpage>157</fpage>&#x2013;<lpage>159</lpage>. doi: <pub-id pub-id-type="doi">10.11907/rjdk.172913</pub-id></citation></ref>
<ref id="ref59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>C.</given-names></name> <name><surname>Gao</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>G.</given-names></name> <name><surname>Shen</surname> <given-names>C.</given-names></name> <name><surname>Sang</surname> <given-names>N.</given-names></name></person-group> (<year>2021</year>). <article-title>BiSeNet V2: bilateral network with guided aggregation for real-time semantic segmentation</article-title>. <source>Int. J. Comput. Vis.</source> <volume>129</volume>, <fpage>3051</fpage>&#x2013;<lpage>3068</lpage>. doi: <pub-id pub-id-type="doi">10.1007/s11263-021-01515-2</pub-id></citation></ref>
<ref id="ref60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yuan</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Multi-level attention network for retinal vessel segmentation</article-title>. <source>IEEE J. Biomed. Health Inform.</source> <volume>26</volume>, <fpage>312</fpage>&#x2013;<lpage>323</lpage>. doi: <pub-id pub-id-type="doi">10.1109/JBHI.2021.3089201</pub-id></citation></ref>
<ref id="ref61"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y</given-names></name> <name><surname>Li</surname> <given-names>K</given-names></name> <name><surname>Li</surname> <given-names>K</given-names></name> <name><surname>Wang</surname> <given-names>L</given-names></name> <name><surname>Zhong</surname> <given-names>B</given-names></name> <name><surname>Fu</surname> <given-names>Y</given-names></name></person-group> (<year>2018</year>). Image super-resolution using very deep Residual Channel attention networks.</citation></ref>
<ref id="ref62"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>H</given-names></name> <name><surname>Kong</surname> <given-names>X</given-names></name> <name><surname>He</surname> <given-names>J</given-names></name> <name><surname>Qiao</surname> <given-names>Y</given-names></name> <name><surname>Dong</surname> <given-names>C</given-names></name></person-group> (<year>2020</year>). Efficient image super-resolution using pixel attention.</citation></ref>
<ref id="ref63"><citation citation-type="book"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Qi</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Jia</surname> <given-names>J.</given-names></name></person-group> (<year>2016</year>). <source>Pyramid scene parsing network</source>. <publisher-loc>Honolulu, HI, USA</publisher-loc>: <publisher-name>IEEE Computer Society</publisher-name>.</citation></ref>
<ref id="ref64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>D. M.</given-names></name> <name><surname>Jin</surname> <given-names>X.</given-names></name> <name><surname>Zhu</surname> <given-names>W. N.</given-names></name></person-group> (<year>2022</year>). <article-title>Nn-TransUNet: an automatic deep learning pipeline for heart MRI segmentation</article-title>. <source>Life</source> <volume>12</volume>:<fpage>1570</fpage>. doi: <pub-id pub-id-type="doi">10.3390/life12101570</pub-id>, PMID: <pub-id pub-id-type="pmid">36295005</pub-id></citation></ref>
<ref id="ref65"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>S</given-names></name> <name><surname>Lu</surname> <given-names>J</given-names></name> <name><surname>Zhao</surname> <given-names>H</given-names></name> <name><surname>Zhu</surname> <given-names>X</given-names></name> <name><surname>Zhang</surname> <given-names>L</given-names></name></person-group> (<year>2020</year>). Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers.</citation></ref>
<ref id="ref66"><citation citation-type="other"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>Z</given-names></name> <name><surname>Siddiquee</surname> <given-names>M</given-names></name> <name><surname>Tajbakhsh</surname> <given-names>N</given-names></name> <name><surname>Liang</surname> <given-names>J</given-names></name></person-group> (<year>2018</year>). UNet++: A nested U-net architecture for medical image segmentation.</citation></ref>
<ref id="ref67"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>D.</given-names></name> <name><surname>Sun</surname> <given-names>D.</given-names></name> <name><surname>Wang</surname> <given-names>D.</given-names></name></person-group> (<year>2022</year>). <article-title>Dual attention mechanism network for lung cancer images super-resolution</article-title>. <source>Comput. Methods Prog. Biomed.</source> <volume>226</volume>:<fpage>107101</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.cmpb.2022.107101</pub-id>, PMID: <pub-id pub-id-type="pmid">36367483</pub-id></citation></ref>
</ref-list>
</back>
</article>