<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Immunol.</journal-id>
<journal-title>Frontiers in Immunology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Immunol.</abbrev-journal-title>
<issn pub-type="epub">1664-3224</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fimmu.2023.1225557</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Immunology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A study on the recognition of monkeypox infection based on deep convolutional neural networks</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Junkang</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/2318086"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Han</surname>
<given-names>Junying</given-names>
</name>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<institution>College of Information Science and Technology, Gansu Agricultural University</institution>, <addr-line>Lanzhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Muhammad Umar, Hazara University, Pakistan</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Krishna Kumar Mohbey, Central University of Rajasthan, India</p>
<p>Talha Burak Alaku&#x15f;, K&#x131;rklareli University, T&#xfc;rkiye</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Junying Han, <email xlink:href="mailto:hanjy@gsau.edu.cn">hanjy@gsau.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>12</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1225557</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>10</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Chen and Han</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Chen and Han</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>The World Health Organization (WHO) has assessed the global public risk of monkeypox as moderate, and 71 WHO member countries have reported more than 14,000 cases of monkeypox infection. At present, the identification of clinical symptoms of monkeypox mainly depends on traditional medical means, which has the problems of low detection efficiency and high detection cost. The deep learning algorithm is excellent in image recognition and can extract and recognize image features quickly and reliably.</p>
</sec>
<sec>
<title>Methods</title>
<p>Therefore, this paper proposes a residual convolutional neural network based on the &#x3bb; function and contextual transformer (LaCTResNet) for the image recognition of monkeypox cases.</p>
</sec>
<sec>
<title>Results</title>
<p>The average recognition accuracy of the neural network model is 91.85%, which is 15.82% higher than that of the baseline model ResNet50 and better than the classical convolutional neural networks models such as AlexNet, VGG16, Inception-V3, and EfficientNet-B5.</p>
</sec>
<sec>
<title>Discussion</title>
<p>This method realizes high-precision identification of skin symptoms of the monkeypox virus to provide a fast and reliable auxiliary diagnosis method for monkeypox cases for front-line medical staff.</p>
</sec>
</abstract>
<kwd-group>
<kwd>monkeypox images</kwd>
<kwd>residual convolutional networks</kwd>
<kwd>deep learning</kwd>
<kwd>aided diagnosis</kwd>
<kwd>contextual transformer</kwd>
</kwd-group>
<contract-num rid="cn001">Grant No.32360437</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="11"/>
<table-count count="8"/>
<equation-count count="17"/>
<ref-count count="27"/>
<page-count count="14"/>
<word-count count="7182"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Viral Immunology</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>According to the World Health Organization (WHO) epidemic records, in 1958, the pathogen &#x201c;monkeypox virus&#x201d; of Macaca cynomolgus monkeys was first identified in the laboratory of Copenhagen, Denmark (<xref ref-type="bibr" rid="B1">1</xref>). In 1970, the first human infection with the monkeypox virus was found in the Democratic Republic of Congo. Since then, monkeypox has been prevalent in Central and West African countries. In 2003, the United States reported the first outbreak of monkeypox outside Africa. At the beginning of May 2022, the spread of the monkeypox virus was first discovered in Britain, and then there were more than 100 cases of monkeypox, and there was a phenomenon of community transmission. Then, monkeypox cases were reported in the United States, Portugal, Spain, Canada, Belgium, Sweden, Italy, and other countries (<xref ref-type="bibr" rid="B2">2</xref>). From January 2022 to January 2023, WHO reported 84,733 laboratory-confirmed cases in 110 countries and regions, including 80 deaths.</p>
<p>Monkeypox virus (MPXV) is an enveloped double-stranded DNA virus that belongs to the genus Orthopoxvirus of Poxviridae, together with Variola virus (VARA) and Cowpox virus (CPXV) (<xref ref-type="bibr" rid="B3">3</xref>). Monkeypox is a zoonotic disease, and African rodents are the primary hosts of the monkeypox virus. Its infection route is similar to smallpox, and it can be spread through respiratory droplets, body fluids, infected animals, or articles contaminated by infected people (<xref ref-type="bibr" rid="B4">4</xref>, <xref ref-type="bibr" rid="B5">5</xref>). After humans are infected with the monkeypox virus, the incubation period is usually 7~14 days, and the longest is 21 days. Then entering the prodromal stage, there will be prodromal symptoms such as lymph node enlargement, fever, headache, and muscle pain, which generally last for 1~2 days (<xref ref-type="bibr" rid="B6">6</xref>); Finally, it enters the eruption period, when the patient is highly contagious, that is, the period of high probability of infection, during which it is most dangerous for uninfected people to contact the patient. History has proved that unidentified or misdiagnosed infectious diseases are the decisive factors for super-transmission events (<xref ref-type="bibr" rid="B7">7</xref>). Therefore, it is urgent to strengthen the global emergency response-ability to public health events and efficiently provide reliable diagnostic information for patients suspected of monkeypox infection.</p>
<p>Currently, the diagnosis of monkeypox cases mainly depends on traditional medical equipment and the artificial experience judgment of doctors. For example, medical and health institutions detect the sequence-specific DNA sequence of the monkeypox virus by Polymerase Chain Reaction (PCR) and analyze its structure and function or isolate the monkeypox virus from clinical and animal specimens by cell culture (<xref ref-type="bibr" rid="B8">8</xref>). With the unremitting efforts of researchers, laboratory detection technology has made a breakthrough. Lv et&#xa0;al. (<xref ref-type="bibr" rid="B9">9</xref>) proposed a loop-mediated isothermal amplification (LAMP) for detecting the monkeypox virus, and its sensitivity was about ten times higher than that of standard PCR. The above method is a direct biochemical analysis of the monkeypox virus. Although the experimental results are highly reliable, the whole operation process has strict requirements on laboratory grade, high-risk factors, and long detection time, which is not conducive to investigating suspected monkeypox virus infection on a large scale. Moreover, most patients are in the eruption stage at the time of initial diagnosis, and their transmission risk is the highest, which will seriously affect the personal safety of doctors.</p>
<p>In recent years, with the updated iteration of computer vision technology and hardware equipment, image recognition using deep learning algorithms has become the mainstream method, and it has also achieved good results in disease investigation (<xref ref-type="bibr" rid="B10">10</xref>&#x2013;<xref ref-type="bibr" rid="B13">13</xref>). The experimental sample is usually a two-dimensional picture of the diseased part taken by medical equipment, which avoids the close contact between medical staff and patients as much as possible and dramatically guarantees the safety of front-line workers. The related research of deep learning technology in auxiliary medical diagnosis includes El Asnaoui et&#xa0;al. (<xref ref-type="bibr" rid="B14">14</xref>) using various CNN models to rapidly diagnose novel coronavirus, among which the recognition accuracy of the models exceeds 96%. Ardakani et&#xa0;al. (<xref ref-type="bibr" rid="B15">15</xref>) evaluated ten deep-learning models using a small data set, including 108 COVID-19 patients and 86 non-COVID-19 patients, and achieved 99% accuracy. Prellberg et&#xa0;al. (<xref ref-type="bibr" rid="B16">16</xref>) used the ResNeXt network to efficiently classify the microscopic images of white blood cells, and the F1 score reached 88.89%. Feng Wang et&#xa0;al. (<xref ref-type="bibr" rid="B17">17</xref>) designed a method based on CNN to differentiate and diagnose benign and malignant nodules in lung CT images and predict the degree of malignancy, and all indicators showed promising results. Roy et&#xa0;al. (<xref ref-type="bibr" rid="B18">18</xref>) used different segmentation techniques to detect skin diseases like acne, candidiasis, cellulitis, chickenpox, etc. Ahsan et&#xa0;al. (<xref ref-type="bibr" rid="B19">19</xref>) used the improved VGG16 model to detect monkeypox. Ali et&#xa0;al. (<xref ref-type="bibr" rid="B20">20</xref>) used VGG16, ResNet50, and Inception-V3 to classify monkeypox and other diseases, among which ResNet50 achieved the best overall accuracy of 82.96%. Mohbey et&#xa0;al. (<xref ref-type="bibr" rid="B21">21</xref>) presenting a hybrid technique based on Convolutional Neural Networks (CNN) and Long Short-Term Memory Networks (LSTM). In this study a knowledge graph of related events based on Twitter data, which provides a real-time and eventful source of new information. The recommended model&#x2019;s accuracy was 94% on the monkeypox tweet dataset. The findings of this research contribute to an increased awareness of monkeypox infection in the general population. Diponkor et&#xa0;al. (<xref ref-type="bibr" rid="B22">22</xref>) proposed an improved model MonkeyNet based on DenseNet-201, which classifies monkeypox from various skin images, and implements the model in a reliable mobile application, and really supports the diagnosis of medical staff. The above research aims to show that deep learning is effective in the medical field and the auxiliary diagnosis of skin infection symptoms of viruses, which can improve disease diagnosis efficiency.</p>
<p>The remarkable achievements of the above-mentioned deep learning technology in disease detection provide a solid basis for using deep learning technology to recognize the image of monkeypox cases. Therefore, this paper proposes a residual convolutional neural network based on the &#x3bb; function and contextual transformer (LaCTResNet) for the image recognition of monkeypox cases. The series of network models are tested on an independent test set of monkeypox images, and the recognition accuracy of the optimal model reaches 91.85%, which is 15.82%, 7.79%, and 29.89% higher than that of the benchmark models ResNet50, CoTResNet50, and LambdaResNet50, respectively. Compared with the similar models AlexNet, VGG16, Inception-V3, and EfficientNet-B5, the recognition accuracy is improved by 25.03%, 20.66%, 34.00%, and 16.05%, respectively. The experimental results fully prove the feasibility and effectiveness of this method in clinical image recognition of monkeypox skin infection to provide a low-cost, high-efficiency, safe, and reliable auxiliary diagnosis method for medical personnel.</p>
<p>The following is the main work of this study:</p>
<list list-type="bullet">
<list-item>
<p>First, to address the problems of the latest monkeypox public dataset (which consists of images from six different categories, including Monkeypox, Varicella, Cowpox, Measles, Variola, and Health images.) on the Kaggle platform, which suffers from the blurring of some of the images as well as the reoccurrence of the images with high similarity, we manually filtered the dataset and utilized the data enhancement strategy to construct a reliable monkeypox dataset.</p>
</list-item>
<list-item>
<p>Secondly, for the weak ability of traditional convolutional layers for sequence modeling and the lack of ability to deal with long-distance dependencies, we introduce the &#x3bb; function layer and CoT function layer and define the residual convolution module based on &#x3bb; function as well as the residual convolution module based on contextual transformer. The optimal collocation form of the two new modules is derived after many experiments, and the optimal collocation ratio of the backbone modules is derived on this basis. Our proposed model has higher recognition accuracy and can assist medical workers in diagnosing and treating more safely and efficiently.</p>
</list-item>
<list-item>
<p>Finally, the optimal model proposed in this study is compared with the benchmark model and the classical model based on the same conditions. The model&#x2019;s performance is evaluated from all aspects and multiple perspectives through metrics such as precision, recall, F1 score, and AUC score, and finally, it is concluded that the model proposed in this study has a more excellent recognition performance.</p>
</list-item>
</list>
<p>The following section is organized: details of the experimental dataset are given in Section 2. After that, Section 3 presents the design ideas of the proposed model, the overall architecture, the algorithmic ideas of the residual convolution module based on the &#x3bb; function, and the residual convolution module based on the contextual transformer. Section 4 describes the experimental equipment setup hyperparameter settings and presents the data from the ablation and comparison experiments. Section 5 presents the evaluation metrics of the model and provides a comprehensive and objective discussion and analysis of the experimental results of the model on the test set. Section 6 briefly summarizes the work process and results of this study.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Dataset</title>
<p>The dataset of the monkeypox in this study mainly comes from the Kaggle platform, which has attracted the attention of 800,000 data scientists. In order to ensure the authenticity and validity of the data, we invited immunologists to identify the image data set one by one, and it further confirmed the authenticity of the image data and the accuracy of the disease categories. The symptoms of various viruses in the monkeypox data set are shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Sample images of the dataset. <bold>(A)</bold> Varicella, <bold>(B)</bold> Cowpox, <bold>(C)</bold> Healthy, <bold>(D)</bold> Measles, <bold>(E)</bold> Monkeypox, <bold>(F)</bold> Variola.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g001.tif"/>
</fig>
<p>The above figure shows that these viruses have seriously eroded the skin, and the differences in symptoms are not noticeable. It is easy for doctors to make misdiagnoses only by naked eye observation in the initial diagnosis. Similarly, in the training of the model, it will also lead to misjudgment due to the high similarity of image features. Therefore, we enhanced the original data set and adjusted the image&#x2019;s brightness, contrast, and position information. The final image data set includes 4,235 clinical images of five kinds of virus infections (Varicella, Cowpox, Measles, Monkeypox, and Variola) and healthy skin images. In order to avoid the uneven distribution of image data caused by subjective interference when manually dividing the data set, we use a random screening algorithm to separate the images in the data set. Each case category is divided into training sets, verification sets, and test sets according to the ratio of 6:2:2 to ensure mutual independence between images. The case categories and the number of images included in the data set are shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. This study focuses on distinguishing monkeypox from other types of cases. However, the five different types of cases are not classified into one category to cope with the situation that the monkeypox virus may mutate from two types to multiple types in the future (refer to the multiple mutations in Covid-19, which will lead to more violent global epidemic transmission events). Therefore, we individually divide the remaining five categories to train a model with solid robustness and generalization.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>The details of the dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Classification</th>
<th valign="middle" align="center">Before</th>
<th valign="middle" align="center">After</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Varicella</td>
<td valign="middle" align="center">178</td>
<td valign="middle" align="center">890</td>
</tr>
<tr>
<td valign="middle" align="center">Cowpox</td>
<td valign="middle" align="center">54</td>
<td valign="middle" align="center">270</td>
</tr>
<tr>
<td valign="middle" align="center">Healthy</td>
<td valign="middle" align="center">50</td>
<td valign="middle" align="center">250</td>
</tr>
<tr>
<td valign="middle" align="center">Measles</td>
<td valign="middle" align="center">47</td>
<td valign="middle" align="center">235</td>
</tr>
<tr>
<td valign="middle" align="center">Monkeypox</td>
<td valign="middle" align="center">160</td>
<td valign="middle" align="center">800</td>
</tr>
<tr>
<td valign="middle" align="center">Variola</td>
<td valign="middle" align="center">358</td>
<td valign="middle" align="center">1790</td>
</tr>
<tr>
<td valign="middle" align="center">
<bold>Total</bold>
</td>
<td valign="middle" align="center">847</td>
<td valign="middle" align="center">4235</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<label>3</label>
<title>The proposed mode</title>
<sec id="s3_1">
<label>3.1</label>
<title>The design idea of the backbone</title>
<p>To meet the requirements of high efficiency and high precision in monkeypox case identification, we define a residual convolution module based on the &#x3bb; function and a residual convolution module based on the contextual transformer, as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. The design idea is to replace the 3&#xd7;3 convolution layer in the traditional residual module with &#x3bb; calculation layer and CoT calculation layer. These two computing layers will be detailed in sections 3.3 and 3.4.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>The design idea of the backbone.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g002.tif"/>
</fig>
<p>In this paper, LaCTResNet neural network model is defined by improving the feature extraction module of the ResNet backbone network, and the overall framework of the network model is shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>. For the infected image of the monkeypox virus to be identified, this network finally outputs the category of infected virus&#xa0;predicted by the network model after passing through a convolution layer, two residual convolution modules based on the &#x3bb; function, two residual convolution modules of the contextual transformer, an average pooling layer, a fully connected layer, and a Softmax layer. In the figure, Conv1 represents the ordinary convolution layer, LamResConv2 and LamResConv3 represent the &#x3bb; function residual convolution module, and CoTResConv4 and CoTResConv5 represent the contextual transformer residual convolution module.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>The overall architecture of the LaCTResNet model. The cube represents the output tensor; H, W, and Channel represent the height, width, and number of channels of the tensor, respectively. The circle represents the output category.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g003.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>The principle of residual learning</title>
<p>The working mechanism of residual learning (<xref ref-type="bibr" rid="B23">23</xref>) is realized by residual connection. There are no redundant branches in the traditional convolutional neural network. From top to bottom, the input signal will be transformed nonlinearly at each layer and directly transmitted to the next layer. The structure of residual learning includes the main branch and the residual difference branch, as shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4</bold>
</xref>. The input signal <italic>x</italic> enters the main branch, passes through the <italic>Weight Layer</italic> and the activation function <italic>ReLU</italic>, and outputs the signal <italic>F(x)</italic>; The residual differential branch directly copies the input signal <italic>x</italic> and transmits it to the output end of the main branch through the cross-layer residual connection, and adds it with the output signal <italic>F(x)</italic> of the main branch to get <italic>F(x)+x</italic>, which is transmitted to the next layer after the activation function <italic>ReLU</italic>. Through residual connection, the main branch&#x2019;s output signal contains the characteristics of traditional nonlinear transformation and residual characteristics. The residual feature represents the difference between the input signal and the expected output. By adding it to the output signal of the main branch, the network can learn these differences more efficiently, improve the accuracy of network feature extraction and recognition, and effectively avoid network degradation.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>Residual structure.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g004.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Residual convolution module based on &#x3bb; function</title>
<p>Modeling long-distance information interaction is an essential topic in deep learning. At present, the mainstream paradigm is the attention mechanism. It is easy to find that the high occupation of secondary memory is not conducive to dealing with long sequences or multi-dimensional input by analyzing the calculation diagram of the attention mechanism. Considering the limitation of self-attention, we introduce the &#x3bb; function layer (<xref ref-type="bibr" rid="B24">24</xref>) into the model, which provides a novel general framework for capturing long-range information interaction between input signals and structured contexts. The &#x3bb; function layer captures the information interaction by transforming the available context into a linear function &#x3bb; and applying these linear functions to each input value respectively. The attention mechanism defines a convolution kernel of similarity between input and context. At the same time, the &#x3bb; function layer aggregates the context information into a linear function with a fixed size, thus skillfully solving the situation that attention tries to occupy much memory. The detailed calculation diagram of the &#x3bb; function layer is shown in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5</bold>
</xref>. The left part is &#x3bb; function generated by context, and the right amount is &#x3bb; function applied to the query.</p>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>Calculation diagram of the &#x3bb; function layer.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g005.tif"/>
</fig>
<p>From the above information, we can draw that, firstly, the Context gets <italic>V</italic>(Value) and <italic>K</italic>(Key) through linear projection, and the mathematical expressions are shown in formulas (1) and (2):</p>
<disp-formula>
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>C</mml:mi>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>K</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>C</mml:mi>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>V</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Next, <italic>V</italic> is multiplied with the position code <italic>E<sub>n</sub>
</italic> to obtain the position-based <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula>, and the mathematical expression is shown in equation (3):</p>
<disp-formula>
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:mi>V</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mi>,</mml:mi>
<mml:msub>
<mml:mi>E</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<italic>K</italic> is then obtained by the normalization operation <inline-formula>
<mml:math display="inline" id="im2">
<mml:mover accent="true">
<mml:mi>K</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> with the mathematical expression shown in equation (4):</p>
<disp-formula>
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>K</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>After that, <inline-formula>
<mml:math display="inline" id="im3">
<mml:mover accent="true">
<mml:mi>K</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
</mml:math>
</inline-formula> is multiplied with <italic>V</italic> to obtain the content-based <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> with the mathematical expression shown in equation (5):</p>
<disp-formula>
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mover accent="true">
<mml:mi>K</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mi>V</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Finally, <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> are multiplied to obtain <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> containing location and context information, and the mathematical expression is shown in equation (6):</p>
<disp-formula>
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mover accent="true">
<mml:mi>K</mml:mi>
<mml:mo>&#xaf;</mml:mo>
</mml:mover>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mi>V</mml:mi>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>E</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:mi>V</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The right part is to map the input value <italic>X</italic> (<italic>X</italic> <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>), by features to <italic>Q</italic>(Query), <italic>Q</italic> <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Then, <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> (<inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>) on the left side is applied to <italic>Q</italic> one by one. The final output <italic>Y</italic> (<italic>Y</italic> <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>) is mathematically expressed as shown in equations (7) and (8):</p>
<disp-formula>
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>X</mml:mi>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>Q</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msubsup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>T</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>c</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mi>&#x3bb;</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>p</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>q</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi>v</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Residual convolution module based on contextual transformer</title>
<p>
<xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6A</bold>
</xref> shows that the traditional self-attention can trigger feature interaction in different spatial positions. However, all the attention matrices of <italic>Query</italic> and <italic>Key</italic> are realized by calculating independent query key pairs, ignoring the rich context information between keys, thus limiting the visual learning ability of the self-attention mechanism in two-dimensional images. Therefore, we introduce the Contextual Transformer Layer (CoT) (<xref ref-type="bibr" rid="B25">25</xref>) into the traditional residual module and construct a residual convolution module based on contextual transformation, the structure of which is shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6B</bold>
</xref>. Through the fusion modeling of local static context and dynamic global context, the CoT computing layer can enlarge the feature distance between samples, improve the problems of significant parameters and weak feature extraction ability in the original ResNet network, and realize the effective extraction of feature information.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>
<bold>(A)</bold> Calculation diagram of traditional attention mechanism, <bold>(B)</bold> Calculation diagram of contextual transformer layer.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g006.tif"/>
</fig>
<p>From the information in Figure (b), it can be seen that the input feature X of size <italic>H &#xd7; W &#xd7; C</italic> is given, and three variables <italic>Q = X</italic>, <italic>K = X</italic>, and <italic>V = XW<sub>V</sub>
</italic> are defined (here only <italic>V</italic> is mapped to the feature, and the original <italic>X</italic> values are still used for <italic>Q</italic> and <italic>K</italic>). A grouped convolution of <italic>k &#xd7; k</italic> (the value of <italic>k</italic> is taken as 3 in this experiment) is performed on <italic>K</italic> to obtain <italic>K</italic> with local context information representation (denoted as <italic>K<sup>1</sup>
</italic>, <italic>K<sup>1</sup>
</italic>&#x2208;<italic>R<sup>H&#xd7;W&#xd7;C</sup>
</italic>), and this <italic>K<sup>1</sup>
</italic> can be seen as static modeling on the local information. Then <italic>K<sup>1</sup>
</italic> and <italic>Q</italic> were Concat, and the result of Concat was subjected to two successive <italic>1&#xd7;1</italic> convolution operations to obtain the attention matrix <italic>A</italic>. The mathematical expression is shown in equation (9):</p>
<disp-formula>
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>K</mml:mi>
<mml:mn>1</mml:mn>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mi>Q</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>&#x3b8;</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>W</mml:mi>
<mml:mi>&#x3b4;</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Unlike the traditional self-attentive mechanism, the A-matrix here is obtained from the interaction of Query information and local context information <italic>K<sup>1</sup>
</italic>, rather than just a simple correspondence between Query and Key. That is, the guidance of regional context modeling enhances the self-attention mechanism. This attentional computational graph <italic>A</italic> and <italic>V</italic> are then subjected to a matrix product operation, which yields <italic>K<sup>2</sup>
</italic> for dynamic context modeling, with the mathematical expression shown in equation (10):</p>
<disp-formula>
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:msup>
<mml:mi>K</mml:mi>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mi>V</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Then the final feature result <italic>Y</italic> is obtained by fusing <italic>K<sup>1</sup>
</italic> from local static context modeling and <italic>K<sup>2</sup>
</italic> from global dynamic context modeling. The design of the CoT layer unifies the context mining between adjacent keys and the self-attentive learning of 2D feature maps, using the context information between input keys to guide the self-attentive learning, thus avoiding the introduction of extra branches for context mining and improving the representational power of the network.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment</title>
<sec id="s4_1">
<label>4.1</label>
<title>Configuration of experimental environment</title>
<p>We conducted experiments on a computer with a GPU model NVIDIA GeForce GTX 3090, and the details are shown in <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Experimental parameters and configuration.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Name</th>
<th valign="middle" align="center">Configuration Information</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Operating System</td>
<td valign="middle" align="center">Ubuntu 18.04</td>
</tr>
<tr>
<td valign="middle" align="center">CPU</td>
<td valign="middle" align="center">Intel(R) Xeon(R) CPU E5-2680 v4 @2.40GHz (7 Cores)</td>
</tr>
<tr>
<td valign="middle" align="center">RAM</td>
<td valign="middle" align="center">30 GB</td>
</tr>
<tr>
<td valign="middle" align="center">GPU</td>
<td valign="middle" align="center">NVIDIA GeForce GTX 3090</td>
</tr>
<tr>
<td valign="middle" align="center">Video Memory</td>
<td valign="middle" align="center">24 GB</td>
</tr>
<tr>
<td valign="middle" align="center">Code management software</td>
<td valign="middle" align="center">PyCharm Community 2021.1.1</td>
</tr>
<tr>
<td valign="middle" align="center">Computer Language</td>
<td valign="middle" align="center">Python 3.8</td>
</tr>
<tr>
<td valign="middle" align="center">Deep learning framework</td>
<td valign="middle" align="center">PyTorch 1.9.0</td>
</tr>
<tr>
<td valign="middle" align="center">Learning Rate</td>
<td valign="middle" align="center">0.001</td>
</tr>
<tr>
<td valign="middle" align="center">Batch Size</td>
<td valign="middle" align="center">16</td>
</tr>
<tr>
<td valign="middle" align="center">Epoch</td>
<td valign="middle" align="center">100</td>
</tr>
<tr>
<td valign="middle" align="center">Loss Function</td>
<td valign="middle" align="center">Cross Entropy Loss</td>
</tr>
<tr>
<td valign="middle" align="center">Optimization Algorithm</td>
<td valign="middle" align="center">SGD</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Ablation experiments</title>
<p>Firstly, the ratio of four groups of residual convolution modules is fixed at 3:4:6:3. That is, the module configuration ratio of ResNet50 is followed. Then, the 3&#xd7;3 convolution in four groups of traditional residual convolution modules is replaced by the &#x3bb; function layer or CoT layer for discussion. A total of five groups of experiments were carried out, and there were sixteen situations: in the first group of experiments, all four groups were replaced by the CoT layer; in The second group of experiments, replacing one group with &#x3bb; function layer and the other three groups with CoT layer, has four situations; The third group of experiments, replacing two of them with &#x3bb; function layer and the other two with CoT layer, has six situations; The fourth group of experiments, replacing three of them with &#x3bb; function layer and the other with CoT layer, has four situations; In the fifth group of experiments, all four groups were replaced by &#x3bb; function layer. The experimental results are shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>The recognition accuracy and parameters of models with different types of residual modules on monkeypox independent test set.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Group</th>
<th valign="middle" align="center">Model Name</th>
<th valign="middle" align="center">&#x3bb; Function Layer</th>
<th valign="middle" align="center">CoT Layer</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">CoTResNet50</td>
<td valign="middle" align="center">[0, 0, 0, 0]</td>
<td valign="middle" align="center">[1, 1, 1, 1]</td>
<td valign="middle" align="center">84.06%</td>
<td valign="middle" align="center">33.79</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">2</td>
<td valign="middle" align="center">LaCTResNet5001</td>
<td valign="middle" align="center">[1, 0, 0, 0]</td>
<td valign="middle" align="center">[0, 1, 1, 1]</td>
<td valign="middle" align="center">87.37%</td>
<td valign="middle" align="center">33.61</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5002</td>
<td valign="middle" align="center">[0, 1, 0, 0]</td>
<td valign="middle" align="center">[1, 0, 1, 1]</td>
<td valign="middle" align="center">87.25%</td>
<td valign="middle" align="center">32.82</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5003</td>
<td valign="middle" align="center">[0, 0, 1, 0]</td>
<td valign="middle" align="center">[1, 1, 0, 1]</td>
<td valign="middle" align="center">83.94%</td>
<td valign="middle" align="center">27.89</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5004</td>
<td valign="middle" align="center">[0, 0, 0, 1]</td>
<td valign="middle" align="center">[1, 1, 1, 0]</td>
<td valign="middle" align="center">82.29%</td>
<td valign="middle" align="center">21.90</td>
</tr>
<tr>
<td valign="middle" rowspan="6" align="center">3</td>
<td valign="middle" align="center">LaCTResNet5012</td>
<td valign="middle" align="center">[1, 1, 0, 0]</td>
<td valign="middle" align="center">[0, 0, 1, 1]</td>
<td valign="middle" align="center">89.02%</td>
<td valign="middle" align="center">32.65</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5013</td>
<td valign="middle" align="center">[1, 0, 1, 0]</td>
<td valign="middle" align="center">[1, 0, 0, 1]</td>
<td valign="middle" align="center">84.89%</td>
<td valign="middle" align="center">27.72</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5014</td>
<td valign="middle" align="center">[1, 0, 0, 1]</td>
<td valign="middle" align="center">[0, 1, 1, 0]</td>
<td valign="middle" align="center">86.19%</td>
<td valign="middle" align="center">21.72</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5023</td>
<td valign="middle" align="center">[0, 1, 1, 0]</td>
<td valign="middle" align="center">[1, 0, 0, 1]</td>
<td valign="middle" align="center">79.69%</td>
<td valign="middle" align="center">26.93</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5024</td>
<td valign="middle" align="center">[0, 1, 0, 1]</td>
<td valign="middle" align="center">[1, 0, 1, 0]</td>
<td valign="middle" align="center">84.53%</td>
<td valign="middle" align="center">20.93</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet5034</td>
<td valign="middle" align="center">[0, 0, 1, 1]</td>
<td valign="middle" align="center">[1, 1, 0, 0]</td>
<td valign="middle" align="center">68.71%</td>
<td valign="middle" align="center">16.00</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">4</td>
<td valign="middle" align="center">LaCTResNet50123</td>
<td valign="middle" align="center">[1, 1, 1, 0]</td>
<td valign="middle" align="center">[0, 0, 0, 1]</td>
<td valign="middle" align="center">83.00%</td>
<td valign="middle" align="center">26.75</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet50124</td>
<td valign="middle" align="center">[1, 1, 0, 1]</td>
<td valign="middle" align="center">[0, 0, 1, 0]</td>
<td valign="middle" align="center">88.43%</td>
<td valign="middle" align="center">20.76</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet50134</td>
<td valign="middle" align="center">[1, 0, 1, 1]</td>
<td valign="middle" align="center">[0, 1, 0, 0]</td>
<td valign="middle" align="center">76.86%</td>
<td valign="middle" align="center">15.83</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet50234</td>
<td valign="middle" align="center">[0, 1, 1, 1]</td>
<td valign="middle" align="center">[1, 0, 0, 0]</td>
<td valign="middle" align="center">62.34%</td>
<td valign="middle" align="center">15.04</td>
</tr>
<tr>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">LambdaResNet50</td>
<td valign="middle" align="center">[1, 1, 1, 1]</td>
<td valign="middle" align="center">[0, 0, 0, 0]</td>
<td valign="middle" align="center">61.98%</td>
<td valign="middle" align="center">14.86</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Looking at the above table, we can see that the effect of replacing all four groups of modules with the &#x3bb; function layer is not ideal, and the recognition accuracy is only 61.98%; The CoT layer replaces all of them, and the recognition accuracy is 84.06%. In the second group of experiments, the best effect is that the &#x3bb; function layer replaces the first group, and the other three groups are the CoT layer, and the recognition accuracy reaches 87.37%. In the third group of experiments, the best effect is that the &#x3bb; function layer replaces the first and second groups, and the CoT layer replaces the other two groups, and the recognition accuracy reaches 89.02%. In the fourth group of experiments, the best effect is that the &#x3bb; function layer replaces the first, second, and fourth groups, and the CoT layer replaces the third group, and the recognition accuracy reaches 88.43%. From this, we get the optimal convolution module collocation form: the first two groups are replaced by the &#x3bb; function layer, and the CoT layer replaces the last two groups. We renamed the optimal model LaCTResNet5012 as LaCTResNet50.</p>
<p>Next, we use the 1:1:1:1 module distribution ratio of ResNet18 and Swin Transformers&#x2019; 1:1:3:1 module distribution ratio (<xref ref-type="bibr" rid="B26">26</xref>) to discuss the LaCTResNet50 model further. These two ratios are proved effective in improving the model&#x2019;s accuracy (<xref ref-type="bibr" rid="B23">23</xref>, <xref ref-type="bibr" rid="B27">27</xref>), and the experimental results are shown in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>The recognition accuracy and parameters of the models with different module ratios on the monkeypox independent test set.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Model Name</th>
<th valign="middle" align="center">Number of modules</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">LaCTResNet50</td>
<td valign="middle" align="center">[3, 4, 6, 3]</td>
<td valign="top" align="center">89.02%</td>
<td valign="top" align="center">32.65</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet2222</td>
<td valign="middle" align="center">[2, 2, 2, 2]</td>
<td valign="middle" align="center">89.26%</td>
<td valign="middle" align="center">19.95</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet3333</td>
<td valign="middle" align="center">[3, 3, 3, 3]</td>
<td valign="middle" align="center">87.84%</td>
<td valign="middle" align="center">27.86</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet2262</td>
<td valign="middle" align="center">[2, 2, 6, 2]</td>
<td valign="middle" align="center">89.02%</td>
<td valign="middle" align="center">26.14</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet3393</td>
<td valign="middle" align="center">[3, 3, 9, 3]</td>
<td valign="middle" align="center">91.85%</td>
<td valign="middle" align="center">37.14</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to the above table, the recognition accuracy of LaCTResNet2222 and LaCTResNet3393 is higher than that of LaCTResNet50, with an increase of 0.24% and 2.83%, respectively. Although the recognition accuracy of LaCTResNet2262 is the same as that of LaCTResNet50, it has certain advantages in parameters. In the end, the best model we got was LaCTResNet3393, and the recognition accuracy was 91.85% on the independent test set of the monkeypox data set. According to the overall structural characteristics of the network, we renamed LaCTResNet2222 to LaCTResNet26, LaCTResNet2262 to LaCTResNet38, and LaCTResNet3393 to LaCTResNet56.</p>
<p>The experimental data mentioned above show that the LaCTResNet model we designed has the advantages of both lightweight and high accuracy. In order to be able to disclose the details of the model&#x2019;s framework more intuitively, we output the framework structure of the LaCTResNet series of models through the program and list the parameters in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>Details of LaCTResNet series network parameters.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Layer Name</th>
<th valign="middle" align="center">Output Size</th>
<th valign="middle" align="center">LaCTResNet26</th>
<th valign="middle" align="center">LaCTResNet38</th>
<th valign="middle" align="center">LaCTResNet56</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">Conv1</td>
<td valign="middle" align="center">112&#xd7;112</td>
<td valign="middle" colspan="3" align="center">7&#xd7;7, 64, stride2</td>
</tr>
<tr>
<td valign="middle" rowspan="2" align="center">LamResConv2</td>
<td valign="middle" rowspan="2" align="center">56&#xd7;56</td>
<td valign="middle" colspan="3" align="center">3&#xd7;3, max pool, stride2</td>
</tr>
<tr>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>64</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="middle" align="center">LamResConv3</td>
<td valign="middle" align="center">28&#xd7;28</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>128</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="middle" align="center">CoTResConv4</td>
<td valign="middle" align="center">14&#xd7;14</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1024</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1024</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>6</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>256</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1024</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>9</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="middle" align="center">CoTResConv5</td>
<td valign="middle" align="center">7&#xd7;7</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2048</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im23">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2048</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
<td valign="middle" align="left">
<inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:mrow>
<mml:mo  stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>CoT</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>512</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2048</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo  stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>
</td>
</tr>
<tr>
<td valign="middle" align="center">Average Pooling</td>
<td valign="middle" align="center">1&#xd7;1</td>
<td valign="middle" align="center">/</td>
<td valign="middle" align="center">/</td>
<td valign="middle" align="center">/</td>
</tr>
<tr>
<td valign="middle" align="center">Fully Connected</td>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
</tr>
<tr>
<td valign="middle" align="center">Softmax</td>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
<td valign="middle" align="center"/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>5-fold cross-validation</title>
<p>To ensure the reliability and stability of the model, we evaluated the model using cross-validation experiments. For a limited sample dataset, five-fold cross-validation is commonly used to evaluate or compare the performance of models. In 5-fold cross-validation, the dataset is divided into five mutually exclusive subsets (i.e., <italic>D = D<sub>1</sub>
</italic> &#x222a; <italic>D<sub>2</sub>
</italic> &#x222a;&#x2026; &#x222a; <italic>D<sub>5</sub>, D<sub>i</sub>
</italic> &#x2229; <italic>D<sub>j</sub>
</italic> = &#x2205; (<italic>i</italic> &#x2260; <italic>j</italic>)), where <italic>D</italic> - <italic>D<sub>i</sub>
</italic> is used as the training set and <italic>D<sub>i</sub>
</italic> (<italic>i</italic> = <italic>1</italic>, <italic>2</italic>,<italic>&#x2026;</italic>, <italic>5</italic>) as the validation set. The cross-validation process is repeated five times, and the results of the five times are averaged to evaluate the model&#x2019;s performance. The 5-fold cross-validation of our proposed model achieves an average recognition accuracy of 91.25% on the training set and 90.54% on the test set. The standard deviations of the above experimental data are minor, only 0.0091 and 0.0119. The experimental data are shown in <xref ref-type="table" rid="T6">
<bold>Table&#xa0;6</bold>
</xref>.</p>
<table-wrap id="T6" position="float">
<label>Table&#xa0;6</label>
<caption>
<p>Recognition accuracy of 5-fold cross-validation.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Fold</th>
<th valign="middle" align="center">Train Accuracy</th>
<th valign="middle" align="center">Test Accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">1</td>
<td valign="middle" align="center">0.9205</td>
<td valign="middle" align="center">0.9185</td>
</tr>
<tr>
<td valign="middle" align="center">2</td>
<td valign="middle" align="center">0.8977</td>
<td valign="middle" align="center">0.8863</td>
</tr>
<tr>
<td valign="middle" align="center">3</td>
<td valign="middle" align="center">0.9233</td>
<td valign="middle" align="center">0.9168</td>
</tr>
<tr>
<td valign="middle" align="center">4</td>
<td valign="middle" align="center">0.9091</td>
<td valign="middle" align="center">0.8992</td>
</tr>
<tr>
<td valign="middle" align="center">5</td>
<td valign="middle" align="center">0.9119</td>
<td valign="middle" align="center">0.9064</td>
</tr>
<tr>
<td valign="middle" align="center">Mean</td>
<td valign="middle" align="center">0.9125</td>
<td valign="middle" align="center">0.9054</td>
</tr>
<tr>
<td valign="middle" align="center">Standard deviation</td>
<td valign="middle" align="center">0.0091</td>
<td valign="middle" align="center">0.0119</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Contrast experiments</title>
<p>We selected the classic deep convolution neural networks AlexNet, VGG16, Inception-V3, ResNet50, and EfficientNet-B5 as the comparative experimental models, and the experimental results are shown in <xref ref-type="table" rid="T7">
<bold>Table&#xa0;7</bold>
</xref>.</p>
<table-wrap id="T7" position="float">
<label>Table&#xa0;7</label>
<caption>
<p>The recognition accuracy and parameters of the classical network models and LaCTResNet series models in the monkeypox independent test set.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Model Name</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" align="center">AlexNet</td>
<td valign="middle" align="center">66.82%</td>
<td valign="middle" align="center">61.10</td>
</tr>
<tr>
<td valign="middle" align="center">VGG16</td>
<td valign="middle" align="center">71.19%</td>
<td valign="middle" align="center">138.36</td>
</tr>
<tr>
<td valign="middle" align="center">Inception-V3</td>
<td valign="middle" align="center">57.85%</td>
<td valign="middle" align="center">22.32</td>
</tr>
<tr>
<td valign="middle" align="center">ResNet50</td>
<td valign="middle" align="center">76.03%</td>
<td valign="middle" align="center">25.56</td>
</tr>
<tr>
<td valign="middle" align="center">EfficientNet-B5</td>
<td valign="middle" align="center">75.80%</td>
<td valign="middle" align="center">2.22</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet26</td>
<td valign="middle" align="center">89.26%</td>
<td valign="middle" align="center">19.95</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet38</td>
<td valign="middle" align="center">89.02%</td>
<td valign="middle" align="center">26.14</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet50</td>
<td valign="middle" align="center">89.02%</td>
<td valign="middle" align="center">36.65</td>
</tr>
<tr>
<td valign="middle" align="center">LaCTResNet56</td>
<td valign="middle" align="center">91.85%</td>
<td valign="middle" align="center">37.14</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The above table shows that the recognition accuracy of classical network models on the independent test set of monkeypox is all below 80%. In comparison, the recognition accuracy of the LaCTResNet series models proposed by us is about 90%, which is 13.23%, 13.03%, and 15.82% higher than that of the ResNet50 model. In order to more intuitively reflect the relationship between the accuracy and parameters of each network model, we draw the relationship diagram between accuracy and parameters, as shown in <xref ref-type="fig" rid="f7">
<bold>Figure&#xa0;7</bold>
</xref>, through visualization software. Considering these two parameters comprehensively, we find that LaCTResNet series models have advantages.</p>
<fig id="f7" position="float">
<label>Figure&#xa0;7</label>
<caption>
<p>The recognition accuracy and parameter quantity of the models. The vertical coordinate indicates the average recognition accuracy of the model, and the horizontal coordinate indicates the number of parameters of the model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g007.tif"/>
</fig>
</sec>
</sec>
<sec id="s5" sec-type="results">
<label>5</label>
<title>Results analysis and performance evaluation</title>
<sec id="s5_1">
<label>5.1</label>
<title>Performance of the model on the training set</title>
<p>In this section, we analyze the accuracy and convergence of the loss value of the model on the monkeypox training set in detail. In the experiment, we adopt a dynamic learning rate strategy. That is, the learning rate decays by 90% every 30 training cycles. This strategy can make the model jump out of the &#x201c;trap&#x201d; of optimal local value and avoid local oscillation, thus effectively improving the convergence speed of the model. The two-dimensional line chart of each model&#x2019;s accuracy and loss value during training is shown in <xref ref-type="fig" rid="f8">
<bold>Figure&#xa0;8</bold>
</xref>. Observing all the curves in the diagram, we can find that the recognition accuracy of each model has significantly jumped after the 30th cycle, and the loss values have decreased significantly. After the 60th cycle, the accuracy and loss values gradually converged and approached a stable value. Let us compare the classic model and LaCTResNet series models as two groups. We can find that the recognition accuracy of the classic model fluctuates between 0.5 and 0.8 after the 60th cycle, while that of LaCTResNet series models fluctuates between 0.7 and 0.9. LaCTResNet series models have apparent advantages in convergence speed and recognition accuracy. The overall trend of the accuracy and loss value curve is consistent with the target expectation and fluctuates in the normal range, showing that the model can deal with local optimal traps, performs well in over-fitting, and can fully grasp the potential &#x201c;universal law&#x201d; in the sample.</p>
<fig id="f8" position="float">
<label>Figure&#xa0;8</label>
<caption>
<p>The recognition accuracy and loss curve of models on the training set. The vertical coordinates in <bold>(A)</bold> indicate the accuracy rate; the vertical coordinates in <bold>(B)</bold> indicate the loss value; the horizontal coordinates all indicate the training period.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g008.tif"/>
</fig>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Performance of the model on the independent test set</title>
<p>In order to verify the generalization ability and robustness of the model, we conducted prediction experiments on the LaCTResNet series model on an independent test set of monkeypox images, and the confusion matrix heat map was drawn by a visualization tool, as shown in <xref ref-type="fig" rid="f9">
<bold>Figure&#xa0;9</bold>
</xref>, which can visually demonstrate the recognition ability of the model.</p>
<fig id="f9" position="float">
<label>Figure&#xa0;9</label>
<caption>
<p>Confusion matrix heat map. The number represents the total number of sample images predicted as that class by the model; the larger the value, the darker the color.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g009.tif"/>
</fig>
<p>In the heat map, the vertical coordinates represent the predicted values of the samples, and the horizontal coordinates represent the actual values of the samples. The main diagonal numbers represent the number of images when the predicted values agree with the actual values, i.e., the number of images correctly recognized by the model, and the remaining position numbers represent the number of images when the predicted values differ from the actual values, i.e., the number of images incorrectly recognized by the model. Observing the confusion matrix heat map of all models, we can find that the classical model has a significant difference in the color depth of the main diagonal blocks of the heat map, which indicates that the model incorrectly recognizes more images; the main diagonal blocks of the heat map of the LaCTResNet series of models all consistently show a darker color, which indicates that the series of models proposed in this paper show better recognition of all six types of images on the monkeypox data set. It has certain technical reference value and application significance.</p>
<p>The model&#x2019;s performance can be evaluated more quantitatively by further analyzing and processing the values in the confusion matrix. In this regard, we must also define the following metrics: <italic>TP</italic>, <italic>FP</italic>, <italic>TN</italic>, and <italic>FN</italic>, representing true positive, false positive, true negative, and false negative, respectively. Among them, <italic>Precision</italic> is the proportion of positive cases correctly predicted by the model to the actual positive cases, reflecting the checking accuracy of the model, and the mathematical expression is shown in equation (11):</p>
<disp-formula>
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>
<italic>Recall</italic>, also known as True Positive Rate (<italic>TPR</italic>), refers to the proportion of positive cases correctly predicted by the model to all positive cases predicted by the model, reflecting the model&#x2019;s check-all rate and the mathematical expression is shown in equation (12):</p>
<disp-formula>
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The False Positive Rate (<italic>FPR</italic>) refers to the proportion of positive cases that the model incorrectly predicts to all positive cases, also called the false identification rate and false alarm rate, and the mathematical expression is shown in equation (13):</p>
<disp-formula>
<label>(13)</label>
<mml:math display="block" id="M13">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>R</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The F1 Score (<italic>F1_score</italic>) is the summed average of precision and recall, which is used to measure the comprehensive performance of the model and the mathematical expression is shown in equation (14):</p>
<disp-formula>
<label>(14)</label>
<mml:math display="block" id="M14">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>_</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Since this experiment is a multiclassification problem, the macro average of the above evaluation metrics is used to measure the &#x201c;global&#x201d; performance of the model, which are the macro accuracy rate (<italic>macro-P</italic>), macro completeness rate (<italic>macro-R</italic>) and macro F1 score (<italic>macro-F1_score</italic>), and the mathematical expressions are shown in equations (15), (16) and (17), respectively:</p>
<disp-formula>
<label>(15)</label>
<mml:math display="block" id="M15">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(16)</label>
<mml:math display="block" id="M16">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:mstyle displaystyle="true">
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:msubsup>
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula>
<label>(17)</label>
<mml:math display="block" id="M17">
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mo>-</mml:mo><mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi><mml:mo>-</mml:mo><mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi><mml:mo>-</mml:mo><mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi><mml:mo>-</mml:mo><mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi><mml:mo>-</mml:mo><mml:mi>R</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Where, <italic>P<sub>i</sub>
</italic> refers to the precision of the <italic>i</italic>th class; <italic>R<sub>i</sub>
</italic> refers to the recall of the <italic>i</italic>th class. Detailed data on the recognition performance of the LaCTResNet family of models are shown in <xref ref-type="table" rid="T8">
<bold>Table&#xa0;8</bold>
</xref>.</p>
<table-wrap id="T8" position="float">
<label>Table&#xa0;8</label>
<caption>
<p>Accuracy, Recall, and F1 Score of each category.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="top" align="left"/>
<th valign="top" align="left">Model Name</th>
<th valign="top" align="center">Varicella</th>
<th valign="top" align="center">Cowpox</th>
<th valign="top" align="center">Healthy</th>
<th valign="top" align="center">Measles</th>
<th valign="top" align="center">Monkeypox</th>
<th valign="top" align="center">Variola</th>
<th valign="top" align="center">Average</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Precision</td>
<td valign="top" align="left">LaCTResNet26</td>
<td valign="top" align="center">0.8404</td>
<td valign="top" align="center">0.9038</td>
<td valign="top" align="center">0.9423</td>
<td valign="top" align="center">0.8511</td>
<td valign="top" align="center">0.8627</td>
<td valign="top" align="center">0.9296</td>
<td valign="top" align="center">0.8883</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet38</td>
<td valign="top" align="center">0.8177</td>
<td valign="top" align="center">0.8364</td>
<td valign="top" align="center">0.9615</td>
<td valign="top" align="center">0.8261</td>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.9345</td>
<td valign="top" align="center">0.8794</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet50</td>
<td valign="top" align="center">0.8407</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.8611</td>
<td valign="top" align="center">0.9214</td>
<td valign="top" align="center">0.8801</td>
<td valign="top" align="center">0.9139</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet56</td>
<td valign="top" align="center">0.8681</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.9423</td>
<td valign="top" align="center">0.7679</td>
<td valign="top" align="center">0.9474</td>
<td valign="top" align="center">0.9521</td>
<td valign="top" align="center">0.8996</td>
</tr>
<tr>
<td valign="top" align="left">Recall/TPR</td>
<td valign="top" align="left">LaCTResNet26</td>
<td valign="top" align="center">0.8876</td>
<td valign="top" align="center">0.8704</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.8511</td>
<td valign="top" align="center">0.825</td>
<td valign="top" align="center">0.9218</td>
<td valign="top" align="center">0.8893</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet38</td>
<td valign="top" align="center">0.9326</td>
<td valign="top" align="center">0.8519</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.8085</td>
<td valign="top" align="center">0.7875</td>
<td valign="top" align="center">0.9162</td>
<td valign="top" align="center">0.8828</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet50</td>
<td valign="top" align="center">0.8596</td>
<td valign="top" align="center">0.8704</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.6596</td>
<td valign="top" align="center">0.8062</td>
<td valign="top" align="center">0.9637</td>
<td valign="top" align="center">0.8566</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet56</td>
<td valign="top" align="center">0.8876</td>
<td valign="top" align="center">0.8519</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.9149</td>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.9441</td>
<td valign="top" align="center">0.9131</td>
</tr>
<tr>
<td valign="top" align="left">F1_score</td>
<td valign="top" align="left">LaCTResNet26</td>
<td valign="top" align="center">0.8634</td>
<td valign="top" align="center">0.8868</td>
<td valign="top" align="center">0.9608</td>
<td valign="top" align="center">0.8511</td>
<td valign="top" align="center">0.8434</td>
<td valign="top" align="center">0.9257</td>
<td valign="top" align="center">0.8885</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet38</td>
<td valign="top" align="center">0.8714</td>
<td valign="top" align="center">0.8441</td>
<td valign="top" align="center">0.9804</td>
<td valign="top" align="center">0.8172</td>
<td valign="top" align="center">0.84</td>
<td valign="top" align="center">0.9253</td>
<td valign="top" align="center">0.8797</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet50</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.9307</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.747</td>
<td valign="top" align="center">0.86</td>
<td valign="top" align="center">0.92</td>
<td valign="top" align="center">0.8813</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet56</td>
<td valign="top" align="center">0.8777</td>
<td valign="top" align="center">0.8846</td>
<td valign="top" align="center">0.9608</td>
<td valign="top" align="center">0.835</td>
<td valign="top" align="center">0.9231</td>
<td valign="top" align="center">0.9481</td>
<td valign="top" align="center">0.9049</td>
</tr>
<tr>
<td valign="top" align="left">FPR</td>
<td valign="top" align="left">LaCTResNet26</td>
<td valign="top" align="center">0.0448</td>
<td valign="top" align="center">0.0063</td>
<td valign="top" align="center">0.0038</td>
<td valign="top" align="center">0.0088</td>
<td valign="top" align="center">0.0306</td>
<td valign="top" align="center">0.0511</td>
<td valign="top" align="center">/</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet38</td>
<td valign="top" align="center">0.0553</td>
<td valign="top" align="center">0.0113</td>
<td valign="top" align="center">0.0025</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.0204</td>
<td valign="top" align="center">0.047</td>
<td valign="top" align="center">
</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet50</td>
<td valign="top" align="center">0.0433</td>
<td valign="top" align="center">0</td>
<td valign="top" align="center">0.0013</td>
<td valign="top" align="center">0.0063</td>
<td valign="top" align="center">0.016</td>
<td valign="top" align="center">0.0961</td>
<td valign="top" align="center">
</td>
</tr>
<tr>
<td valign="top" align="left">
</td>
<td valign="top" align="left">LaCTResNet56</td>
<td valign="top" align="center">0.0359</td>
<td valign="top" align="center">0.005</td>
<td valign="top" align="center">0.0038</td>
<td valign="top" align="center">0.0163</td>
<td valign="top" align="center">0.0116</td>
<td valign="top" align="center">0.0348</td>
<td valign="top" align="center">
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Observing the information in the above table, we can find that the LaCTResNet56 model has the most significant <italic>macro-F1_score</italic> value as well as the highest <italic>F1_score</italic> value for monkeypox case recognition, the model has the best overall recognition performance, and the overall recognition of monkeypox cases is better than other models.</p>
<p>As shown in <xref ref-type="fig" rid="f10">
<bold>Figure&#xa0;10</bold>
</xref>, the combined recognition ability of the LaCTResNet family of models for each category is demonstrated. The models are less effective in recognizing measles cases. It is because the clinical symptoms of measles, chickenpox, and smallpox all show a large rash on the skin with redness and swelling. The characteristics are highly similar, which can easily lead to model misidentification.</p>
<fig id="f10" position="float">
<label>Figure&#xa0;10</label>
<caption>
<p>F1 Score of the LaCTResNet family of models.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g010.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="f11">
<bold>Figure&#xa0;11</bold>
</xref> shows the Receiver Operating Characteristic (ROC) curve of the optimal model, LaCTResNet56, for each category of cases on the monkeypox dataset. The vertical coordinate represents the <italic>TPR</italic>, the horizontal coordinate represents the <italic>FPR</italic>, and the Area Under Curve (AUC) represents the recognition effect of the model for that category, and the more significant the area, the better the effect. If the AUC is less than 0.5, it means that the model does not have realistic reference significance for identifying this category, while the closer the AUC is to 1, the better the model is for identifying this category. It can be seen from the figure that the AUC of measles cases is 0.9235. It is worth mentioning that the AUC of monkeypox cases reaches 0.9442, and the average AUC of all categories reaches 0.9476, proving that the model can identify monkeypox cases with high accuracy.</p>
<fig id="f11" position="float">
<label>Figure&#xa0;11</label>
<caption>
<p>ROC curve of LaCTResNet56 model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fimmu-14-1225557-g011.tif"/>
</fig>
</sec>
</sec>
<sec id="s6" sec-type="conclusions">
<label>6</label>
<title>Conclusions</title>
<p>In this study, we introduced the &#x3bb; function layer and CoT function layer to replace the 3&#xd7;3 convolutional layer in the original model based on the original ResNet model and defined a new residual convolution module. We tried to replace all the original convolution modules with new modules containing CoT function layers. The accuracy of this model on the dataset reached 84.06%. By replacing the first module with new modules containing &#x3bb; function layers and the rest with new modules containing CoT function layers, the accuracy of this model on the dataset reached 87.37%. Based on the good results of the above experiments, we continued to optimize the model structure by discussing the position of the two new modules and the ratio of the number of new modules, respectively and finally found that the model constructed by replacing the first two modules with modules containing the &#x3bb; function layer, and replacing the last two modules with modules containing the CoT function layer, and using either a 1:1:1:1 ratio of the modules or a 1:1:3:1 ratio, was able to exhibit excellent performance. Among them, our constructed optimal model LaCTResNet56 achieves an average recognition accuracy of 91.85% on the test set, which is 15.82%, 7.79%, and 29.89% better than the baseline models ResNet50, CoTResNet50, and LambdaResNet50, and better than the similar models AlexNet, VGG16, Inception-V3, and EfficientNet-B5 by 25.03%, 20.66%, 34.00%, and 16.05%, respectively. We evaluated the recognition accuracy of the optimal model using the ROC curve and concluded that the model has a strong recognition ability for monkeypox cases, with an AUC of 0.9442. The above experimental data aim to show that our model has excellent comprehensive performance, can effectively extract the feature information of monkeypox cases and identify similar cases efficiently and reliably, and has certain practical significance in the auxiliary diagnosis of monkeypox cases.</p>
<p>Our future work is mainly based on the following considerations: (1) To collect as many images of monkeypox clinical cases as possible and continuously optimize the model. (2) Since there are already mutated strains of the monkeypox virus, we will collect clinical images to classify cases of different monkeypox strains individually and conduct experiments to support the precision of monkeypox epidemic tracking. (3) The model will be deployed to intelligent devices such as cell phones.</p>
</sec>
<sec id="s7" sec-type="data-availability">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/maxmelichov/monkeypox-2022-remastered">https://www.kaggle.com/datasets/maxmelichov/monkeypox-2022-remastered</ext-link>.</p>
</sec>
<sec id="s8" sec-type="ethics-statement">
<title>Ethics statement</title>
<p>The studies involving humans were approved by ethics committee of the Experimental Animal Center of Gansu Agricultural University(GSAU-Eth-ASF2022-008). The studies were conducted in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required from the participants or the participants&#x2019; legal guardians/next of kin because Our experiment used a public dataset (see: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/maxmelichov/monkeypox-2022-remastered">https://www.kaggle.com/datasets/maxmelichov/monkeypox-2022-remastered</ext-link>). We did not have direct contact with patients to collect clinical data. Anyone who may be identified in the dataset has been blindfolded.</p>
</sec>
<sec id="s9" sec-type="author-contributions">
<title>Author contributions</title>
<p>Conceptualization, JC and JH; Data curation, JC and JH; Formal analysis, JC; Funding acquisition, JH; Investigation, JC; Methodology, JC; Project administration, JC; Resources, JC and JH; Software, JC; Supervision, JC; Validation, JC and JH; Visualization, JC; Writing&#x2014;original draft, JC; Writing&#x2014;review and editing, JC and JH. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s10" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare financial support was received for the research, authorship, and/or publication of this article. This work was supported by the National Natural Science Foundation of China (Grant No.32360437); by the Innovation Fund Project of Colleges and Universities in Gansu of China (Grant No.2021A-056); and by the Industrial Support and Guidance Project of Universities in Gansu Province, China (Grant No.2021CYZC-57).</p>
</sec>
<sec id="s11" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s12" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<label>1</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname> <given-names>N</given-names>
</name>
<name>
<surname>Acharya</surname> <given-names>A</given-names>
</name>
<name>
<surname>Gendelman</surname> <given-names>HE</given-names>
</name>
<name>
<surname>Byrareddy</surname> <given-names>SN</given-names>
</name>
</person-group>. <article-title>The 2022 outbreak and the pathobiology of the monkeypox virus</article-title>. <source>J Autoimmun</source> (<year>2022</year>) <volume>131</volume>:<fpage>102855</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.jaut.2022.102855</pub-id>
</citation>
</ref>
<ref id="B2">
<label>2</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>M</given-names>
</name>
<name>
<surname>Cui</surname> <given-names>F</given-names>
</name>
</person-group>. <article-title>Status of epidemiology and prevention and control of monkeypox</article-title>. <source>Jiangsu J Prev Med</source> (<year>2023</year>) <volume>34</volume>(<issue>01</issue>):<fpage>8</fpage>&#x2013;<lpage>11</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.13668/j.issn.1006-9070.2023.01.002</pub-id>
</citation>
</ref>
<ref id="B3">
<label>3</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shchelkunova</surname> <given-names>GA</given-names>
</name>
<name>
<surname>Shchelkunov</surname> <given-names>SN</given-names>
</name>
</person-group>. <article-title>Smallpox, monkeypox and other human orthopoxvirus infections</article-title>. <source>Viruses</source> (<year>2022</year>) <volume>15</volume>(<issue>1</issue>):<fpage>103</fpage>. doi: <pub-id pub-id-type="doi">10.3390/v15010103</pub-id>
</citation>
</ref>
<ref id="B4">
<label>4</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumar</surname> <given-names>P</given-names>
</name>
<name>
<surname>Chaudhary</surname> <given-names>B</given-names>
</name>
<name>
<surname>Yadav</surname> <given-names>N</given-names>
</name>
<name>
<surname>Devi</surname> <given-names>S</given-names>
</name>
<name>
<surname>Pareek</surname> <given-names>A</given-names>
</name>
<name>
<surname>Alla</surname> <given-names>S</given-names>
</name>
<etal/>
</person-group>. <article-title>Recent advances in research and management of human monkeypox virus: an emerging global health threat</article-title>. (<year>2023</year>) <volume>15</volume>(<issue>4</issue>):<fpage>937</fpage>. doi: <pub-id pub-id-type="doi">10.3390/v15040937</pub-id>
</citation>
</ref>
<ref id="B5">
<label>5</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Algarate</surname> <given-names>S</given-names>
</name>
<name>
<surname>Bueno</surname> <given-names>J</given-names>
</name>
<name>
<surname>Crusells</surname> <given-names>MJ</given-names>
</name>
<name>
<surname>Ara</surname> <given-names>M</given-names>
</name>
<name>
<surname>Alonso</surname> <given-names>H</given-names>
</name>
<name>
<surname>Alvarado</surname> <given-names>E</given-names>
</name>
<etal/>
</person-group>. <article-title>Usefulness of non-skin samples in the PCR diagnosis of Mpox (Monkeypox)[J]</article-title>. <source>Viruses</source> (<year>2023</year>) <volume>15</volume>(<issue>5</issue>):<fpage>1107</fpage>. doi: <pub-id pub-id-type="doi">10.3390/v15051107</pub-id>
</citation>
</ref>
<ref id="B6">
<label>6</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Candela</surname> <given-names>C</given-names>
</name>
<name>
<surname>Raccagni</surname> <given-names>AR</given-names>
</name>
<name>
<surname>Bruzzesi</surname> <given-names>E</given-names>
</name>
<name>
<surname>Bertoni</surname> <given-names>C</given-names>
</name>
<name>
<surname>Rizzo</surname> <given-names>A</given-names>
</name>
<name>
<surname>Gagliardi</surname> <given-names>G</given-names>
</name>
<etal/>
</person-group>. <article-title>Human monkeypox experience in a tertiary level hospital in Milan, Italy, between May and October 2022: epidemiological features and clinical characteristics</article-title>. <source>Viruses</source> (<year>2023</year>) <volume>15</volume>(<issue>3</issue>):<fpage>667</fpage>. doi: <pub-id pub-id-type="doi">10.3390/v15030667</pub-id>
</citation>
</ref>
<ref id="B7">
<label>7</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lloyd-Smith</surname> <given-names>JO</given-names>
</name>
<name>
<surname>Schreiber</surname> <given-names>SJ</given-names>
</name>
<name>
<surname>Kopp</surname> <given-names>PE</given-names>
</name>
<name>
<surname>Getz</surname> <given-names>WM</given-names>
</name>
</person-group>. <article-title>Superspreading and the effect of individual variation on disease emergence</article-title>. <source>Nature</source> (<year>2005</year>) <volume>438</volume>(<issue>7066</issue>):<page-range>355&#x2013;9</page-range>. doi: <pub-id pub-id-type="doi">10.1038/nature04153</pub-id>
</citation>
</ref>
<ref id="B8">
<label>8</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>Y</given-names>
</name>
</person-group>. <article-title>Monkeypox</article-title>. In: <source>Progress in Microbiology and Immunology</source>, vol. <volume>04</volume>. <publisher-name>Lanzhou Institute of biological Products CO., LTD</publisher-name>, <publisher-loc>China</publisher-loc> (<year>2022</year>). p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>. Available at: <uri xlink:href="http://kns.cnki.net/kcms/detail/62.1120.R.20220608.1908.002.html">http://kns.cnki.net/kcms/detail/62.1120.R.20220608.1908.002.html</uri>. 2022-09-05.</citation>
</ref>
<ref id="B9">
<label>9</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname> <given-names>Q</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>W</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>P</given-names>
</name>
<name>
<surname>He</surname> <given-names>L</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>Q</given-names>
</name>
<etal/>
</person-group>. <article-title>Establishment of loop-mediated isothermal amplification method for Monkeypox virus</article-title>. <source>Chin J Health Lab Technol</source> (<year>2013</year>) <volume>23</volume>(<issue>05</issue>):<page-range>1170&#x2013;3</page-range>. Available at: <uri xlink:href="https://kns.cnki.net/kcms2/article/abstract?v=LD-wYsOa3DhxRG7dsHpRGFZoAIkiNZLjyU8iuNToTVQzLQFwlsn62epHE1FCT4XFWmGPPv2Wzf8HFHrAQvoFUFpPR8cPIZWet7p4LYYr-fIqUkwsT41hZTnH3hwrB3Gw&amp;uniplatform=NZKPT&amp;language=CHS">https://kns.cnki.net/kcms2/article/abstract?v=LD-wYsOa3DhxRG7dsHpRGFZoAIkiNZLjyU8iuNToTVQzLQFwlsn62epHE1FCT4XFWmGPPv2Wzf8HFHrAQvoFUFpPR8cPIZWet7p4LYYr-fIqUkwsT41hZTnH3hwrB3Gw&amp;uniplatform=NZKPT&amp;language=CHS</uri>.</citation>
</ref>
<ref id="B10">
<label>10</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsuneki</surname> <given-names>M</given-names>
</name>
<name>
<surname>Abe</surname> <given-names>M</given-names>
</name>
<name>
<surname>Kanavati</surname> <given-names>F</given-names>
</name>
</person-group>. <article-title>Deep learning-based screening of urothelial carcinoma in whole slide images of liquid-based cytology urine specimens</article-title>. <source>Cancers</source> (<year>2023</year>) <volume>15</volume>(<issue>1</issue>):<fpage>226</fpage>. doi: <pub-id pub-id-type="doi">10.3390/cancers15010226</pub-id>
</citation>
</ref>
<ref id="B11">
<label>11</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mosquera-Zamudio</surname> <given-names>A</given-names>
</name>
<name>
<surname>Launet</surname> <given-names>L</given-names>
</name>
<name>
<surname>Tabatabaei</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Parra-Medina</surname> <given-names>R</given-names>
</name>
<name>
<surname>Colomer</surname> <given-names>A</given-names>
</name>
<name>
<surname>Moll</surname> <given-names>JO</given-names>
</name>
<etal/>
</person-group>. <article-title>Deep learning for skin melanocytic tumors in whole-slide images: A systematic review</article-title>. <source>Cancers</source> (<year>2022</year>) <volume>15</volume>(<issue>1</issue>):<fpage>42</fpage>. doi: <pub-id pub-id-type="doi">10.3390/cancers15010042</pub-id>
</citation>
</ref>
<ref id="B12">
<label>12</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ibraheim</surname> <given-names>MK</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>R</given-names>
</name>
<name>
<surname>Gardner</surname> <given-names>JM</given-names>
</name>
<name>
<surname>Elsensohn</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Artificial intelligence in dermatopathology: an analysis of its practical application</article-title>. <source>Dermatopathology</source> (<year>2023</year>) <volume>10</volume>(<issue>1</issue>):<page-range>93&#x2013;4</page-range>. doi: <pub-id pub-id-type="doi">10.3390/dermatopathology10010014</pub-id>
</citation>
</ref>
<ref id="B13">
<label>13</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>J</given-names>
</name>
<name>
<surname>Ko</surname> <given-names>S</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>M</given-names>
</name>
<name>
<surname>Park</surname> <given-names>NJY</given-names>
</name>
<name>
<surname>Han</surname> <given-names>H</given-names>
</name>
<name>
<surname>Cho</surname> <given-names>J</given-names>
</name>
<etal/>
</person-group>. <article-title>Deep learning prediction of TERT promoter mutation status in thyroid cancer using histologic images</article-title>. <source>Medicina</source> (<year>2023</year>) <volume>59</volume>(<issue>3</issue>):<fpage>536</fpage>. doi: <pub-id pub-id-type="doi">10.3390/medicina59030536</pub-id>
</citation>
</ref>
<ref id="B14">
<label>14</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>El Asnaoui</surname> <given-names>K</given-names>
</name>
<name>
<surname>Chawki</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Idri</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Automated methods for detection and classification pneumonia based on x-ray images using deep learning</article-title>. In: <source>Artificial intelligence and blockchain for future cybersecurity applications</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2021</year>). p. <page-range>257&#x2013;84</page-range>.</citation>
</ref>
<ref id="B15">
<label>15</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ardakani</surname> <given-names>AA</given-names>
</name>
<name>
<surname>Kanafi</surname> <given-names>AR</given-names>
</name>
<name>
<surname>Acharya</surname> <given-names>UR</given-names>
</name>
<name>
<surname>Khadem</surname> <given-names>N</given-names>
</name>
<name>
<surname>Mohammadi</surname> <given-names>A</given-names>
</name>
</person-group>. <article-title>Application of deep learning technique to manage COVID-19 in routine clinical practice using CT images: Results of 10 convolutional neural networks</article-title>. <source>Comput Biol Med</source> (<year>2020</year>) <volume>121</volume>:<fpage>103795</fpage>. doi: <pub-id pub-id-type="doi">10.1016/j.compbiomed.2020.103795</pub-id>
</citation>
</ref>
<ref id="B16">
<label>16</label>
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Prellberg</surname> <given-names>J</given-names>
</name>
<name>
<surname>Kramer</surname> <given-names>O</given-names>
</name>
</person-group>. <article-title>Acute lymphoblastic leukemia classification from microscopic images using convolutional neural networks</article-title>. In: <source>ISBI 2019 C-NMC Challenge: Classification in Cancer Cell Imaging</source>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name> (<year>2019</year>). p. <fpage>53</fpage>&#x2013;<lpage>61</lpage>.</citation>
</ref>
<ref id="B17">
<label>17</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>F</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>L</given-names>
</name>
<name>
<surname>Li</surname> <given-names>N</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Z</given-names>
</name>
</person-group>. <article-title>Deep learning based on three-dimensional convolutional neural network for differential diagnosis of benign and Malignant pulmonary nodules</article-title>. <source>Chin J Med Imaging</source> (<year>2019</year>) <volume>27</volume>(<issue>10</issue>):<fpage>779</fpage>&#x2013;<lpage>782+787</lpage>. Available at: <uri xlink:href="https://kns.cnki.net/kcms2/article/abstract?v=LD-wYsOa3DhjYjnB-4D5RXfFntwuPA5AjovreeLW6jyetF_H78HepkWPwiufX0W8srwz7a6sPs2DLIHw2jiSRLVz2Mdl9LN8qR-UFiE3c55FNTQphyL5FP_1cv-Lcb7pOijaA0eB8j0=&amp;uniplatform=NZKPT&amp;language=CHS">https://kns.cnki.net/kcms2/article/abstract?v=LD-wYsOa3DhjYjnB-4D5RXfFntwuPA5AjovreeLW6jyetF_H78HepkWPwiufX0W8srwz7a6sPs2DLIHw2jiSRLVz2Mdl9LN8qR-UFiE3c55FNTQphyL5FP_1cv-Lcb7pOijaA0eB8j0=&amp;uniplatform=NZKPT&amp;language=CHS</uri>.</citation>
</ref>
<ref id="B18">
<label>18</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Roy</surname> <given-names>K</given-names>
</name>
<name>
<surname>Chaudhuri</surname> <given-names>SS</given-names>
</name>
<name>
<surname>Ghosh</surname> <given-names>S</given-names>
</name>
<name>
<surname>Dutta</surname> <given-names>SK</given-names>
</name>
<name>
<surname>Chakraborty</surname> <given-names>P</given-names>
</name>
<name>
<surname>Sarkar</surname> <given-names>R</given-names>
</name>
</person-group>. (<year>2019</year>). <article-title>Skin Disease detection based on different Segmentation Techniques</article-title>, in: <conf-name>2019 International Conference on Opto-Electronics and Applied Optics (Optronix). IEEE</conf-name>, <publisher-loc>America</publisher-loc>. pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi: <pub-id pub-id-type="doi">10.1109/OPTRONIX.2019.8862403</pub-id>
</citation>
</ref>
<ref id="B19">
<label>19</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahsan</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Uddin</surname> <given-names>MR</given-names>
</name>
<name>
<surname>Farjana</surname> <given-names>M</given-names>
</name>
<name>
<surname>Sakib</surname> <given-names>AN</given-names>
</name>
<name>
<surname>Momin</surname> <given-names>KA</given-names>
</name>
<name>
<surname>Luna</surname> <given-names>SA</given-names>
</name>
<etal/>
</person-group>. <article-title>Image Data collection and implementation of deep learning-based model in detecting Monkeypox disease using modified VGG16</article-title>. <source>arXiv preprint arXiv</source> (<year>2022</year>). doi: <pub-id pub-id-type="doi">10.48550/arXiv.2206.01862</pub-id>
</citation>
</ref>
<ref id="B20">
<label>20</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ali</surname> <given-names>SN</given-names>
</name>
<name>
<surname>Ahmed</surname> <given-names>M</given-names>
</name>
<name>
<surname>Paul</surname> <given-names>J</given-names>
</name>
<name>
<surname>Jahan</surname> <given-names>T</given-names>
</name>
<name>
<surname>Sani</surname> <given-names>SMS</given-names>
</name>
<name>
<surname>Noor</surname> <given-names>N</given-names>
</name>
<etal/>
</person-group>. <article-title>Monkeypox skin lesion detection using deep learning models: A feasibility study</article-title>. <source>arXiv preprint arXiv</source> (<year>2022</year>). doi: <pub-id pub-id-type="doi">10.48550/arXiv.2207.03342</pub-id>
</citation>
</ref>
<ref id="B21">
<label>21</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mohbey</surname> <given-names>KK</given-names>
</name>
<name>
<surname>Meena</surname> <given-names>G</given-names>
</name>
<name>
<surname>Kumar</surname> <given-names>S</given-names>
</name>
<name>
<surname>Lokesh</surname> <given-names>K</given-names>
</name>
</person-group>. <article-title>A CNN-LSTM-based hybrid deep learning approach to detect sentiment polarities on Monkeypox tweets</article-title>. <source>arXiv preprint arXiv</source> (<year>2022</year>). doi: <pub-id pub-id-type="doi">10.1007/s00354-023-00227-0</pub-id>
</citation>
</ref>
<ref id="B22">
<label>22</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bala</surname> <given-names>D</given-names>
</name>
<name>
<surname>Hossain</surname> <given-names>MS</given-names>
</name>
<name>
<surname>Hossain</surname> <given-names>MA</given-names>
</name>
<name>
<surname>Abdullah</surname> <given-names>MI</given-names>
</name>
<name>
<surname>Rahman</surname> <given-names>MM</given-names>
</name>
<name>
<surname>Manavalan</surname> <given-names>B</given-names>
</name>
<etal/>
</person-group>. <article-title>MonkeyNet: A robust deep convolutional neural network for monkeypox disease detection and classification</article-title>. <source>Neural Networks</source> (<year>2023</year>) <volume>161</volume>:<page-range>757&#x2013;75</page-range>. doi: <pub-id pub-id-type="doi">10.1016/j.neunet.2023.02.022</pub-id>
</citation>
</ref>
<ref id="B23">
<label>23</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>He</surname> <given-names>K</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>S</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>J</given-names>
</name>
</person-group>. (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>, in: <conf-name>Proceedings of the IEEE conference on computer vision and pattern recognition</conf-name>, <publisher-loc>America</publisher-loc>. pp. <page-range>770&#x2013;8</page-range>.</citation>
</ref>
<ref id="B24">
<label>24</label>
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bello</surname> <given-names>I</given-names>
</name>
</person-group>. <article-title>Lambdanetworks: Modeling long-range interactions without attention</article-title>. <source>arXiv preprint arXiv</source> (<year>2021</year>). doi: <pub-id pub-id-type="doi">10.48550/arXiv.2102.08602</pub-id>
</citation>
</ref>
<ref id="B25">
<label>25</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>T</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Mei</surname> <given-names>T</given-names>
</name>
</person-group>. (<year>2022</year>). <article-title>Contextual transformer networks for visual recognition</article-title>, in: <conf-name>IEEE Transactions on Pattern Analysis and Machine Intelligence</conf-name>, <publisher-loc>America</publisher-loc>. <volume>45</volume>(<issue>2</issue>):<page-range>1489&#x2013;1500</page-range>. doi: 10.1109/TPAMI.2022.3164083</citation>
</ref>
<ref id="B26">
<label>26</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>Y</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>Y</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Swin transformer: Hierarchical vision transformer using shifted windows</article-title>, in: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>, <publisher-loc>America</publisher-loc>. pp. <page-range>10012&#x2013;22</page-range>.</citation>
</ref>
<ref id="B27">
<label>27</label>
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Z</given-names>
</name>
<name>
<surname>Mao</surname> <given-names>H</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>CY</given-names>
</name>
<name>
<surname>Feichtenhofer</surname> <given-names>C</given-names>
</name>
<name>
<surname>Darrell</surname> <given-names>T</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>S</given-names>
</name>
</person-group>. (<year>2022</year>). <article-title>A convnet for the 2020s</article-title>, in: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>, <publisher-loc>America</publisher-loc>. pp. <page-range>11976&#x2013;86</page-range>.</citation>
</ref>
</ref-list>
</back>
</article>