<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">875112</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.875112</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>PredMHC: An Effective Predictor of Major Histocompatibility Complex Using Mixed Features</article-title>
<alt-title alt-title-type="left-running-head">Chen and Li</alt-title>
<alt-title alt-title-type="right-running-head">PredMHC</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Dong</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/838100/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Yanjuan</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1465511/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>College of Electrical and Information Engineering</institution>, <institution>Quzhou University</institution>, <addr-line>Quzhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/531759/overview">Quan Zou</ext-link>, University of Electronic Science and Technology of China, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/578368/overview">Chunyu Wang</ext-link>, Harbin Institute of Technology, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1473670/overview">Haiying Zhang</ext-link>, Xiamen University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Yanjuan Li, <email>lyjuan5@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>25</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>875112</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Chen and Li.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Chen and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>The major histocompatibility complex (MHC) is a large locus on vertebrate DNA that contains a tightly linked set of polymorphic genes encoding cell surface proteins essential for the adaptive immune system. The groups of proteins encoded in the MHC play an important role in the adaptive immune system. Therefore, the accurate identification of the MHC is necessary to understand its role in the adaptive immune system. An effective predictor called PredMHC is established in this study to identify the MHC from protein sequences. Firstly, PredMHC encoded a protein sequence with mixed features including 188D, APAAC, KSCTriad, CKSAAGP, and PAAC. Secondly, three classifiers including SGD, SMO, and random forest were trained on the mixed features of the protein sequence. Finally, the prediction result was obtained by the voting of the three classifiers. The experimental results of the 10-fold cross-validation test in the training dataset showed that PredMHC can obtain 91.69% accuracy. Experimental results on comparison with other features, classifiers, and existing methods showed the effectiveness of PredMHC in predicting the MHC.</p>
</abstract>
<kwd-group>
<kwd>protein classification</kwd>
<kwd>major histocompatibility complex</kwd>
<kwd>machine learning</kwd>
<kwd>feature extraction</kwd>
<kwd>identification</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>As a large locus on vertebrate DNA, the major histocompatibility complex (MHC) contains a tightly linked set of polymorphic genes encoding cell surface proteins that are essential for immune surveillance. These cell surface proteins are called MHC molecules (<xref ref-type="bibr" rid="B23">Kubiniok et al., 2022</xref>). MHC molecules are classified into MHC class I, MHC class II, and MHC class III according to variation in molecular structure, function, and distribution (<xref ref-type="bibr" rid="B33">Marcoux et al., 2021</xref>). MHC class I molecules are expressed in all nucleated cells and platelets&#x2014;essentially all cells except red blood cells, which display antigens to signal cytotoxic T lymphocytes, including clusters of differentiation (CD8<sup>&#x2b;</sup>) (<xref ref-type="bibr" rid="B34">McShan et al., 2021</xref>). MHC class II molecules are expressed in antigen-presenting cells, such as B cells, dendritic cells, and macrophages, where they normally bind to CD4<sup>&#x2b;</sup> receptors on helper T cells to clear foreign antigens. MHC class III genes are interleaved with class I and class II genes on the short arm of chromosome 6, but their proteins play different physiological roles.</p>
<p>MHC molecules are cell surface glycoproteins with a three-dimensional structure and are of vital importance to infection, autoimmunity, transplantation, and tumor immunotherapy. MHC-binding prediction plays an important role in identifying potential novel therapeutic strategies. <xref ref-type="bibr" rid="B32">Mahoney et al. (2021)</xref> pointed out that MHC phosphopeptides can be considered potential immunotherapeutic targets for cancer and other chronic diseases. Therefore, many scholars carried out a lot of research work on MHC-binding prediction. The first computational method (<xref ref-type="bibr" rid="B7">Altuvia et al., 1995</xref>) to uncover the MHC-binding peptide was developed by Altuvia et al., which is based on protein structure and is further improved to distinguish candidate peptides that bind to hydrophobic binding pockets of the MHC molecules (<xref ref-type="bibr" rid="B8">Altuvia et al., 1997</xref>). The SVRMHC (<xref ref-type="bibr" rid="B25">Liu et al., 2006</xref>) is an MHC-binding peptide model which encoded peptides with physicochemical properties and trained support vector machines to construct a prediction model on mice. NetMHC-3.0 (<xref ref-type="bibr" rid="B26">Lundegaard et al., 2008</xref>) is a web server with high performance for predicting peptide binders based on artificial neural networks. Boehm et al. proposed a method named ForestMHC (<xref ref-type="bibr" rid="B10">Boehm et al., 2019</xref>) to identify immunogenic peptides. ForestMHC encoded a peptide sequence with physicochemical properties and trained a random forest classifier to construct an identification model. <xref ref-type="bibr" rid="B41">Saxena et al. (2020)</xref> predicted the binding potential of peptides to the MHC, which is critical for designing peptide-based therapeutics, using a deep learning model named OnionMHC. In consideration of the importance of structural information, the OnionMHC represents peptides with its sequence and structure-based features for peptide-HLA-A&#x2a;02:01 binding predictions. (<xref ref-type="bibr" rid="B30">Lv et al., 2020</xref>) <xref ref-type="bibr" rid="B21">Jiang et al. (2021)</xref> gave a comprehensive review of the state-of-the-art literature on MHC-binding peptide prediction and an in-depth evaluation of feature representation methods, prediction models, and model training strategies on benchmark datasets. Based on the limitation of only handling peptide sequences with fixed length, Jiang et al. proposed a novel variable-length MHC-binding prediction model named BVLSTM-MHC. Experimental results on an independent validation dataset showed that BVLSTM-MHC has better performance than the ten mainstream prediction tools.</p>
<p>Scientists are devoted to discover MHC molecules in various vertebrate genomes. <xref ref-type="bibr" rid="B20">Hopkins et al. (1986)</xref> described a rat monoclonal antibody which can recognize MHC class II antigens in sheep and seems to recognize determinants which are nonpolymorphic. Moreover, based on the antibody, the distribution of sheep class II molecules is investigated, and the class II- expression variations by cells in efferent lymph and peripheral is also investigated. <xref ref-type="bibr" rid="B53">Westbrook et al. (2015)</xref> combined the SMRT sequencing technology and CCS and introduced and validated the technology of SMRT-CCS on identifying class I transcripts in Mauritian-origin cynomolgus macaques. Furthermore, SMRT-CCS was applied to characterize 60 new full-length class I transcriptional sequences expressed in the Chinese cynomolgus monkey population. By using pyrosequencing with high-resolution and Sanger sequencing technology, <xref ref-type="bibr" rid="B42">Shiina et al. (2015)</xref> genotyped 127 unrelated animals and identified 112 different alleles. Moreover, the International Society for Animal Genetics (ISAG) standardized the nomenclature and established the IPD-MHC database which is used to scientifically manage the MHC allele sequences and genes from nonhuman organisms (<xref ref-type="bibr" rid="B19">Giuseppe et al., 2017</xref>; <xref ref-type="bibr" rid="B31">Maccari et al., 2018</xref>; <xref ref-type="bibr" rid="B5">Ali et al., 2021</xref>; <xref ref-type="bibr" rid="B13">Burton et al., 2021</xref>; <xref ref-type="bibr" rid="B22">Karcioglu and Bulut, 2021</xref>; <xref ref-type="bibr" rid="B38">Roy et al., 2021</xref>; <xref ref-type="bibr" rid="B39">Safaei et al., 2021</xref>; <xref ref-type="bibr" rid="B51">Wang et al., 2021</xref>).</p>
<p>At early stages, the research studies related to the MHC are developed based on mice experiments. With the availability of a large amount of data and development of machine learning, developing a machine learning&#x2013;based model to research the MHC was feasible. <xref ref-type="bibr" rid="B24">Li et al. (2019)</xref> proposed an identification method of the MHC based on an extreme learning machine algorithm. Although high accuracy has been achieved, there are still many aspects worthy of further investigation (<xref ref-type="bibr" rid="B27">Lv et al., 2019</xref>; <xref ref-type="bibr" rid="B28">Lv et al., 2021a</xref>; <xref ref-type="bibr" rid="B29">Lv et al., 2021b</xref>). In this study, we aim to propose a new MHC predictor, PredMHC, to further improve prediction performance.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>Materials and Methods</title>
<sec id="s2-1">
<title>Framework of PredMHC</title>
<p>In this study, we introduced a novel MHC predictor named PredMHC, the framework of which is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. First, PredMHC encoded a protein sequence with mixed features including 188D, APAAC, KSCTriad, CKSAAGP, and PAAC. Second, three classifiers including SGD, SMO, and random forest were trained on the mixed features of protein sequence. Finally, the prediction result was obtained by the voting of the three classifiers. We will introduce the datasets, feature extraction, and classifiers in detail in the following section.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Framework of PredMHC.</p>
</caption>
<graphic xlink:href="fgene-13-875112-g001.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>Dataset</title>
<p>The dataset constructed by <xref ref-type="bibr" rid="B24">Li et al. (2019)</xref> is used in this study. A web server called ELM-MHC was developed by Li et al., from which the dataset can be downloaded. The reason that we used the same dataset as ELM-MHC is as follows. First, the dataset is constructed by searching for MHC sequences on the Uniprot database, and it is reliable. Second, the dataset is used cd-hit to de-duplication processing. The protein sequences are clustered based on the parameter setting, and the sequence with the maximum length in every cluster is used as a representative sequence. The redundant and homology-biased sequences are removed in this dataset. Finally, the most important inference was that we can fairly compare with the existing method by using the same dataset. The final dataset contained 13,488 protein sequences, which consists of 6,712 MHC protein sequences (positive examples) and 6,776 nonMHC protein sequences (negative examples). All protein sequences were divided into two groups: 10,790 sequences as a set of 10-fold cross-validation and 2,698 sequences as a set of independent validation. The training dataset (Train-10790) comprised 5,370 MHC protein sequences and 5,420 nonMHC protein sequences, all randomly selected from the set of positive and negative examples, respectively. They were then further randomly divided into five sets for the input of 10-fold cross-validation. The independent testing dataset (Test-2698) contained 1,342 positive and 1,356 negative examples.</p>
</sec>
<sec id="s2-3">
<title>Feature Extraction</title>
<p>To classify a protein sequence into different categories using the machine learning method, the first step is to encode the protein sequence with features. A feature that can effectively discriminate positive examples from negative examples can greatly improve the prediction performance of the model. In this study, we try to encode protein sequences with mixed features including 188D, APAAC, KSCTriad, CKSAAGP, and PAAC. The mixed features can represent a protein sequence from different prospectives; thus, it can better distinguish different protein sequences.</p>
<sec id="s2-3-1">
<title>SVMProt-188D</title>
<p>SVMProt-188D is a feature extraction method based on the amino acid composition and physicochemical properties (<xref ref-type="bibr" rid="B18">Dubchak et al., 1995</xref>; <xref ref-type="bibr" rid="B40">Saxena et al., 2021</xref>). It encodes each protein sequence as a 188-dimensional feature vector. The first 20 features are the frequencies of the 20 amino acids (A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, and Y in alphabetical order) occurring in the sequence. The formula is defined as<disp-formula id="equ1">
<mml:math id="m1">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mtext>&#x200a;</mml:mtext>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mi>L</mml:mi>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>N</italic>
<sub>i</sub> denotes the number of the <italic>i</italic>th amino acid in the protein sequence and L denotes the length of a sequence. Obviously, <inline-formula id="inf1">
<mml:math id="m2">
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The latter dimensions are correlated with eight physicochemical properties, namely, hydrophobicity, normalized Van der Waals volume, polarity, polarizability, charge, surface tension, secondary structure, and solvent accessibility. Each physicochemical property consists of 21 numbers. In detail, each property consists of three descriptors, composition (C), transition (T), and distribution (D). C indicates the proportion of amino acids with specific physicochemical properties to all amino acids, and the dimension of C is 3; T represents the percentage frequency of amino acids with a specific property behind amino acids with another property, and its dimension is 3; and D represents the proportions of the chain length of 0, 25, 50, 75, and 100% amino acids with a specific property, and its dimension is 8. Therefore, after analyzing the composition and eight physicochemical properties of amino acids, we can obtain a total of 20&#x2b;(3 &#x2b; 5&#x2b;8)&#xd7;8 &#x3d; 188 features.</p>
</sec>
<sec id="s2-3-2">
<title>Amphiphilic Pseudo Amino Acid Composition</title>
<p>The concept of amphiphilic pseudo amino acid composition (APAAC), originally proposed by Chou (<xref ref-type="bibr" rid="B17">Chou, 2005</xref>; <xref ref-type="bibr" rid="B28">Lv et al., 2021a</xref>; <xref ref-type="bibr" rid="B9">Awais et al., 2021</xref>; <xref ref-type="bibr" rid="B35">Naseer et al., 2021</xref>; <xref ref-type="bibr" rid="B54">Yan et al., 2021</xref>), is an effective protein descriptor and has been applied for diverse protein sequence analysis. APAAC is different from traditional AAC. It can incorporate a partial sequence-order effect by using the hydrophobicity and hydrophilicity of the constituent amino acids in a protein. For the convenience of the readers, we will briefly introduce the concept of APAAC. Let R<sub>1</sub>R<sub>2</sub>R<sub>3</sub>...R<sub>L</sub> be a protein sequence with length L, where R<sub>1</sub> denotes the residue at position 1, R<sub>2</sub> denotes the residue at positon 2, and so forth. According to the definition of APAAC, a protein can be denoted as a vector P with dimension (20&#x2b;2&#x3bb;). Vector P is defined as follows.<disp-formula id="e1">
<mml:math id="m3">
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="normal">&#x3bb;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi mathvariant="normal">&#x3bb;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where P<sub>1,</sub> P<sub>2</sub>, &#x2026; , P<sub>20</sub> in <xref ref-type="disp-formula" rid="e1">Eq. 1</xref> represent the classic AAC and the next 2&#x3bb; discrete numbers describe the sequence correlation factor.</p>
</sec>
<sec id="s2-3-3">
<title>K-Spaced Conjoint Triad</title>
<p>The k-spaced conjoint triad (KSCTriad) (<xref ref-type="bibr" rid="B14">Chao et al., 2018</xref>; <xref ref-type="bibr" rid="B64">Zhen et al., 2020</xref>) is an effective protein descriptor and has been comprehensively applied for diverse biological sequence analyses. Different from the conjoint triad descriptor, KSCTriad not only calculates the number of three continuous amino acid units but also incorporates the continuous amino acid units that are separated by any k-residues.</p>
</sec>
<sec id="s2-3-4">
<title>Composition of K-Spaced Amino Acid Group Pairs</title>
<p>The composition of k-spaced amino acid pairs (CKSAAP) (<xref ref-type="bibr" rid="B15">Chen et al., 2010</xref>; <xref ref-type="bibr" rid="B1">Ahmad et al., 2021</xref>; <xref ref-type="bibr" rid="B2">Akbar et al., 2021</xref>; <xref ref-type="bibr" rid="B3">Al-Qazzaz et al., 2021</xref>; <xref ref-type="bibr" rid="B4">Alar and Fernandez, 2021</xref>; <xref ref-type="bibr" rid="B6">Alim et al., 2021</xref>; <xref ref-type="bibr" rid="B12">Buriro et al., 2021</xref>) method describes the order-related information of the protein sequence, which takes the occurrence frequency of two amino acids separated by k-residues in the sequence as a feature element. The protein contains 20 amino acids; thus, a 400-dimensional feature vector can be obtained for each interval. The composition of k-spaced amino acid group pairs (CKSAAGP) is a variation of the CKSAAP method. The 20 amino acids can be classified into five groups based on the chemical properties of their side chains: the aliphatic group, aromatic group, positive charged group, negative charged group, and uncharged group. The CKSAAGP method is based on the frequency of the two groups separated by a k-spaced amino acid.</p>
</sec>
<sec id="s2-3-5">
<title>Pseudo-Amino Acid Composition</title>
<p>The conventional amino acid composition is defined in a 20-D space, and each dimension represents the frequency of the occurrence of one of the 20 native amino acids. Different from the conventional amino acid protein composition, the pseudo-amino acid composition (<xref ref-type="bibr" rid="B16">Chou, 2001</xref>; <xref ref-type="bibr" rid="B9">Awais et al., 2021</xref>), which is a vector with 20&#x2b;&#x3bb; discrete components, will contain much more sequence-order and sequence-length information. According to the concept of pseudo-amino acid composition, the feature is given by<disp-formula id="equ2">
<mml:math id="m4">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x22ee;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mo>&#x22ee;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mn>20</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>where the first 20 components are the occurrence frequencies of the 20 amino acids in the protein which is the same as in the conventional amino acid composition, while the additional components p<sub>20&#x2b;1</sub> &#x2026; p<sub>20&#x2b;&#x3bb;</sub> are the sequence-order correlation factors of the different ranks.</p>
</sec>
</sec>
<sec id="s2-4">
<title>Classifier</title>
<p>To obtain better classification results, we adopted the voting of three base classifiers as the final classification result. The three classifiers were, respectively, random forest, SMO, and SGD. The three classifiers are popular and have been successfully used in bioinformatics many times.</p>
<p>Random forest is an ensemble classifier based on the decision tree algorithm proposed by Breiman in 2001 (<xref ref-type="bibr" rid="B11">Breiman, 2001</xref>). To solve regression or classification tasks, random forests construct many decision trees by extracting subsets from all the samples through the bootstrap technique and obtain the prediction result by voting on these decision trees. Random forests are widely used in bioinformatics because of their low computational overhead and ability of handling unbalanced data.</p>
<p>The support vector machine (SVM) (<xref ref-type="bibr" rid="B36">Hearst et al., 1998</xref>) is a well-known machine learning algorithm that completes various classification tasks by constructing a separating hyperplane in the high-dimensional space. However, the training speed of support vector machines is heavily influenced by data size. To solve this problem, the sequential minimum optimization (SMO) (<xref ref-type="bibr" rid="B37">Platt, 1999</xref>) algorithm was proposed, which decomposes large quadratic programming problems (OPs) of an original SVM into a series of the smallest possible QP problems. Moreover, the solution process of SMO needs no additional matrix storage, thus saving both time and space costs.</p>
<p>The goal of the stochastic gradient descent (SGD) algorithm is to find a path that leads to optimal result. When using this algorithm, the parameter values are first initialized, and then these values are continuously changed until the target function converges. The SGD algorithm is widely used to process large-scale sparse data, such as text classification tasks.</p>
</sec>
<sec id="s2-5">
<title>Measurement</title>
<p>To evaluate the performance of the proposed method, we introduced four indicators commonly used in bioinformatics: sensitivity (SE), specificity (SP), accuracy (ACC), and Matthew&#x2019;s correlation coefficient (MCC). The formulae of these indicators are as follows (<xref ref-type="bibr" rid="B61">Zhang et al., 2021a</xref>; <xref ref-type="bibr" rid="B29">Lv et al., 2021b</xref>; <xref ref-type="bibr" rid="B60">Zhang et al., 2021b</xref>; <xref ref-type="bibr" rid="B59">Zhang et al., 2021c</xref>; <xref ref-type="bibr" rid="B58">Zhang et al., 2021d</xref>; <xref ref-type="bibr" rid="B57">Zhang et al., 2021e</xref>; <xref ref-type="bibr" rid="B62">Zhao et al., 2021</xref>; <xref ref-type="bibr" rid="B65">Zhu et al., 2021</xref>; <xref ref-type="bibr" rid="B66">Zou et al., 2021</xref>; <xref ref-type="bibr" rid="B63">Zhao et al., 2022</xref>).<disp-formula id="equ3">
<mml:math id="m5">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>E</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ4">
<mml:math id="m6">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ5">
<mml:math id="m7">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ6">
<mml:math id="m8">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>where TP is an abbreviation for true positives, representing the number of MHC proteins predicted in positive examples; FP is an abbreviation for false positives, representing the number of MHC proteins predicted in negative examples; TN is an abbreviation for true negatives, representing nonMHC proteins predicted in negative examples; and FN is an abbreviation for false negatives and indicates the number of predicted nonMHC proteins in positive examples. SE and SP represent the predictive accuracy of the model in positive and negative samples, respectively. Both ACC and MCC represent the overall performance of the model. For all the aforementioned metrics , the higher the score they get the better the performance of the model.</p>
</sec>
</sec>
<sec id="s3">
<title>Result and Discussion</title>
<sec id="s3-1">
<title>Cross-Validation Results of Train-10790</title>
<p>In many experiments, we tried a variety of methods to extract highly recognizable features from protein sequences in the training set and used several algorithms to train the model to achieve optimal accuracy. The experimental comparison results of different features are explained in <italic>Performance of Different Features on Cross-Validation</italic>, and the experimental comparison results of different classifiers are explained in <italic>Performance of Different Classifiers on Cross-Validation</italic>.</p>
<sec id="s3-1-1">
<title>Performance of Different Features on Cross-Validation</title>
<p>Using the voting of random forest, SMO, and SGD as the classification model, we first tried 188D, APAAC, KSCTriad, CKSAAGP, PAAC, and their combinations. <xref ref-type="table" rid="T1">Table 1</xref> shows the performance of the five single features and several combinations of features with good performance in the 10-fold cross-validation. As shown in <xref ref-type="table" rid="T1">Table 1</xref>, according to the indexes MCC and ACC, the mixed features proposed in this study have the highest score; thus, our method has better overall performance. According to the indicator of SE, the feature of APAAC has the highest score, whereas its value of ACC, MCC, and SP is lower; it verifies that the feature of APAAC was bias to classify a protein into the MHC protein. Similar to APAAC, PAAC also has higher value on the indicator SE and lower value on other indicators. Therefore, from the overall perspective, our method obviously performs better than all other methods.</p>
<table-wrap id="T1" position="float">
<label>TABLE1</label>
<caption>
<p>Result of different features on Train-10790.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feaures</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">SE</th>
<th align="center">SP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">(1)-188D</td>
<td align="char" char=".">0.8953</td>
<td align="char" char=".">0.7927</td>
<td align="char" char=".">0.8596</td>
<td align="char" char=".">0.9310</td>
</tr>
<tr>
<td align="left">(2)-APAAC</td>
<td align="char" char=".">0.8329</td>
<td align="char" char=".">0.6824</td>
<td align="char" char=".">0.9494</td>
<td align="char" char=".">0.7108</td>
</tr>
<tr>
<td align="left">(3)-KSCTriad</td>
<td align="char" char=".">0.8764</td>
<td align="char" char=".">0.7580</td>
<td align="char" char=".">0.8177</td>
<td align="char" char=".">0.9350</td>
</tr>
<tr>
<td align="left">(4)-CKSAAGP</td>
<td align="char" char=".">0.8682</td>
<td align="char" char=".">0.7469</td>
<td align="char" char=".">0.7826</td>
<td align="char" char=".">0.9529</td>
</tr>
<tr>
<td align="left">(5)-PAAC</td>
<td align="char" char=".">0.8283</td>
<td align="char" char=".">0.6739</td>
<td align="char" char=".">0.9485</td>
<td align="char" char=".">0.7018</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC</td>
<td align="char" char=".">0.9003</td>
<td align="char" char=".">0.8019</td>
<td align="char" char=".">0.8735</td>
<td align="char" char=".">0.9276</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad</td>
<td align="char" char=".">0.8872</td>
<td align="char" char=".">0.7782</td>
<td align="char" char=".">0.8386</td>
<td align="char" char=".">0.9360</td>
</tr>
<tr>
<td align="left">KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.8993</td>
<td align="char" char=".">0.8039</td>
<td align="char" char=".">0.8404</td>
<td align="char" char=".">0.9576</td>
</tr>
<tr>
<td align="left">CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.8848</td>
<td align="char" char=".">0.7728</td>
<td align="char" char=".">0.8376</td>
<td align="char" char=".">0.9316</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC &#x2b; KSCTriad</td>
<td align="char" char=".">0.9121</td>
<td align="char" char=".">0.8268</td>
<td align="char" char=".">0.8734</td>
<td align="char" char=".">0.9511</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.9054</td>
<td align="char" char=".">0.8155</td>
<td align="char" char=".">0.8518</td>
<td align="char" char=".">0.9589</td>
</tr>
<tr>
<td align="left">KSCTriad &#x2b; CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.9041</td>
<td align="char" char=".">0.8127</td>
<td align="char" char=".">0.8516</td>
<td align="char" char=".">0.9565</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC &#x2b; KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.9157</td>
<td align="char" char=".">0.8351</td>
<td align="char" char=".">0.8701</td>
<td align="char" char=".">0.9618</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad &#x2b; CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.9065</td>
<td align="char" char=".">0.8178</td>
<td align="char" char=".">0.8522</td>
<td align="char" char=".">0.9608</td>
</tr>
<tr>
<td align="left">Our mixed feature</td>
<td align="char" char=".">0.9169</td>
<td align="char" char=".">0.8370</td>
<td align="char" char=".">0.8761</td>
<td align="char" char=".">0.9587</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-1-2">
<title>Performance of Different Classifiers on Cross-Validation</title>
<p>To verify the performance of our used classifier, we compared the classifier used in this study with other classifiers. <xref ref-type="table" rid="T2">Table 2</xref> shows the experimental results. As shown in <xref ref-type="table" rid="T2">Table 2</xref>, the voting of SGD, SMO, and random forest used in our identification system has better performance than other single classifiers. As shown in <xref ref-type="table" rid="T2">Table 2</xref>, our classification model has 0.9169% accuracy and 0.8370 MCC, which are higher than those of other classifiers. It verified that our classification model has better overall performance. According to the number of winning incidences, our classification wins on three indicators and has the highest number of wins. It is shown in <xref ref-type="table" rid="T2">Table 2</xref> that the SE of our classification model was slightly lower than that of random forest. However, the values of ACC, MCC, and SP of our classification model are obviously higher than those of random forest. Therefore, from the overall perspective, our classification model obviously performs better than all other classifiers.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Result of different classifiers on Train-10790.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Classifiers</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">SE</th>
<th align="center">SP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">SGD</td>
<td align="char" char=".">0.8794</td>
<td align="char" char=".">0.7600</td>
<td align="char" char=".">0.8504</td>
<td align="char" char=".">0.9081</td>
</tr>
<tr>
<td align="left">SMO</td>
<td align="char" char=".">0.9038</td>
<td align="char" char=".">0.8106</td>
<td align="char" char=".">0.8594</td>
<td align="char" char=".">0.9478</td>
</tr>
<tr>
<td align="left">Random forest</td>
<td align="char" char=".">0.8850</td>
<td align="char" char=".">0.7699</td>
<td align="char" char=".">0.8830</td>
<td align="char" char=".">0.8869</td>
</tr>
<tr>
<td align="left">Our classification model</td>
<td align="char" char=".">0.9169</td>
<td align="char" char=".">0.8370</td>
<td align="char" char=".">0.8761</td>
<td align="char" char=".">0.9587</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3-2">
<title>Independent-Validation Results of Test-2698</title>
<p>To evaluate the generalization performance of the proposed model, we tested its performance on the Test-2698 dataset. In detail, we trained the model proposed in this study on the Train-10790 dataset and then computed its performance on the test-2698 dataset. The experimental results are shown in <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>. As shown in <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>, the feature extraction method and classifier used in this study have better performance than the other feature extraction methods and classifiers, respectively.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Result of different features on Test-2698.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Features</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">SE</th>
<th align="center">SP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">188D</td>
<td align="char" char=".">0.8926</td>
<td align="char" char=".">0.7869</td>
<td align="char" char=".">0.8593</td>
<td align="char" char=".">0.9259</td>
</tr>
<tr>
<td align="left">APAAC</td>
<td align="char" char=".">0.8357</td>
<td align="char" char=".">0.6892</td>
<td align="char" char=".">0.9533</td>
<td align="char" char=".">0.7139</td>
</tr>
<tr>
<td align="left">KSCTriad</td>
<td align="char" char=".">0.8741</td>
<td align="char" char=".">0.7504</td>
<td align="char" char=".">0.8355</td>
<td align="char" char=".">0.9127</td>
</tr>
<tr>
<td align="left">CKSAAGP</td>
<td align="char" char=".">0.8774</td>
<td align="char" char=".">0.7614</td>
<td align="char" char=".">0.8098</td>
<td align="char" char=".">0.9442</td>
</tr>
<tr>
<td align="left">PAAC</td>
<td align="char" char=".">0.8326</td>
<td align="char" char=".">0.6826</td>
<td align="char" char=".">0.9527</td>
<td align="char" char=".">0.7056</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC</td>
<td align="char" char=".">0.9010</td>
<td align="char" char=".">0.8061</td>
<td align="char" char=".">0.8482</td>
<td align="char" char=".">0.9530</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad</td>
<td align="char" char=".">0.8940</td>
<td align="char" char=".">0.7888</td>
<td align="char" char=".">0.8697</td>
<td align="char" char=".">0.9182</td>
</tr>
<tr>
<td align="left">KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.9055</td>
<td align="char" char=".">0.8155</td>
<td align="char" char=".">0.8540</td>
<td align="char" char=".">0.9573</td>
</tr>
<tr>
<td align="left">CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.8901</td>
<td align="char" char=".">0.7818</td>
<td align="char" char=".">0.8571</td>
<td align="char" char=".">0.9230</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC &#x2b; KSCTriad</td>
<td align="char" char=".">0.9172</td>
<td align="char" char=".">0.8355</td>
<td align="char" char=".">0.8938</td>
<td align="char" char=".">0.9412</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.9130</td>
<td align="char" char=".">0.8287</td>
<td align="char" char=".">0.8729</td>
<td align="char" char=".">0.9532</td>
</tr>
<tr>
<td align="left">KSCTriad &#x2b; CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.9155</td>
<td align="char" char=".">0.8337</td>
<td align="char" char=".">0.8769</td>
<td align="char" char=".">0.9544</td>
</tr>
<tr>
<td align="left">188D &#x2b; APAAC &#x2b; KSCTriad &#x2b; CKSAAGP</td>
<td align="char" char=".">0.9198</td>
<td align="char" char=".">0.8416</td>
<td align="char" char=".">0.8841</td>
<td align="char" char=".">0.9550</td>
</tr>
<tr>
<td align="left">APAAC &#x2b; KSCTriad &#x2b; CKSAAGP &#x2b; PAAC</td>
<td align="char" char=".">0.9134</td>
<td align="char" char=".">0.8300</td>
<td align="char" char=".">0.8693</td>
<td align="char" char=".">0.9574</td>
</tr>
<tr>
<td align="left">Our mixed feature</td>
<td align="char" char=".">0.9246</td>
<td align="char" char=".">0.8502</td>
<td align="char" char=".">0.9034</td>
<td align="char" char=".">0.9466</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Result of different classifiers on Test-2698.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Classifier</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">SE</th>
<th align="center">SP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">SGD</td>
<td align="char" char=".">0.8959</td>
<td align="char" char=".">0.7918</td>
<td align="char" char=".">0.8935</td>
<td align="char" char=".">0.8982</td>
</tr>
<tr>
<td align="left">SMO</td>
<td align="char" char=".">0.9063</td>
<td align="char" char=".">0.8147</td>
<td align="char" char=".">0.8682</td>
<td align="char" char=".">0.9440</td>
</tr>
<tr>
<td align="left">Random forest</td>
<td align="char" char=".">0.8948</td>
<td align="char" char=".">0.7896</td>
<td align="char" char=".">0.8913</td>
<td align="char" char=".">0.8982</td>
</tr>
<tr>
<td align="left">Our classification model</td>
<td align="char" char=".">0.9246</td>
<td align="char" char=".">0.8502</td>
<td align="char" char=".">0.9034</td>
<td align="char" char=".">0.9466</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>Comparison With Other Predictors</title>
<p>To evaluate the performance of the classifier PredMHC, we compared it with ELM-MHC on the same dataset including Train-10790 and Test-2698. The comparison results on the 10-fold cross-validation are shown in <xref ref-type="table" rid="T5">Table 5</xref>. As we can see from <xref ref-type="table" rid="T5">Table 5</xref>, PredMHC has higher score than ELM-MHC on the indicators ACC, MCC, and SP. According to the number of winning incidence, PredMHC has better performance than ELM-MHC. According to ACC and MCC, PredMHC has better overall performance than ELM-MHC. Therefore, PredMHC is superior to the existing methods in the prediction of MHC protein.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Comparison of 10-fold cross-validation with the existing method on all data.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">SE</th>
<th align="center">SP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">ELM-MHC</td>
<td align="char" char=".">0.9166</td>
<td align="char" char=".">0.822</td>
<td align="char" char=".">0.893</td>
<td align="char" char=".">0.908</td>
</tr>
<tr>
<td align="left">Our method</td>
<td align="char" char=".">0.9185</td>
<td align="char" char=".">0.8403</td>
<td align="char" char=".">0.8741</td>
<td align="char" char=".">0.9627</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="conclusion" id="s4">
<title>Conclusion</title>
<p>In this study, we proposed an efficient, reliable, and simple experimental model for predicting the MHC protein based on mixed features. After a large number of comparative experiments, we selected the mixed features of 188D, APAAC, KSCTriad, CKSAAGP, and PAAC, which showed global performance on the 10-fold cross-validation training dataset and independent test dataset. We then used the voting of SGD, SMO, and random forest to build a prediction model which also achieved the best performance on both training and test datasets. In terms of important indicators, our model obtained an MCC of 0.8370 and ACC of 0.9169 in the 10-fold cross-validation based on the Train-10790 dataset and MCC of 0.8502 and ACC of 0.9246 in the independent validation based on the Test-2698 dataset. In conclusion, we believe that our novel model provides an efficient and reliable method to screen MHCs from a large number of protein sequences. In the future, we will pay more attention to deep learning classifiers and evolution strategies (<xref ref-type="bibr" rid="B43">Tahoces et al., 2021</xref>; <xref ref-type="bibr" rid="B44">Tandel et al., 2021</xref>; <xref ref-type="bibr" rid="B45">Tavolara et al., 2021</xref>; <xref ref-type="bibr" rid="B46">Togacar, 2021</xref>; <xref ref-type="bibr" rid="B47">Tsiknakis et al., 2021</xref>; <xref ref-type="bibr" rid="B48">Turki and Taguchi, 2021</xref>; <xref ref-type="bibr" rid="B49">Usman et al., 2021</xref>; <xref ref-type="bibr" rid="B50">Vafaeezadeh et al., 2021</xref>; <xref ref-type="bibr" rid="B51">Wang et al., 2021</xref>; <xref ref-type="bibr" rid="B52">Watanabe et al., 2021</xref>; <xref ref-type="bibr" rid="B55">Yap et al., 2021</xref>; <xref ref-type="bibr" rid="B56">Yildirim et al., 2021</xref>).</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>Conceptualization, YL; data curation, DC; formal analysis, DC; project administration, DC; writing&#x2014;original draft, YL; and writing&#x2014;review and editing, DC.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was supported by the Research Start-up Funding Project of Quzhou University (BSYJ202112 and BSYJ202109), the National Natural Science Foundation of China (61901103 and 61671189), and the Natural Science Foundation of Heilongjiang Province (LH 2019F002).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahmad</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Farooq</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>M. U. G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning Model for Pathogen Classification Using Feature Fusion and Data Augmentation</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>3</issue>), <fpage>466</fpage>&#x2013;<lpage>483</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200707143535</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akbar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ahmad</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hayat</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rehman</surname>
<given-names>A. U.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ali</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>iAtbP-Hyb-EnC: Prediction of Antitubercular Peptides via Heterogeneous Feature Representation and Genetic Algorithm Based Ensemble Learning Model</article-title>. <source>Comput. Biol. Med.</source> <volume>137</volume>, <fpage>104778</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104778</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Al-Qazzaz</surname>
<given-names>N. K.</given-names>
</name>
<name>
<surname>Alyasseri</surname>
<given-names>Z. A. A.</given-names>
</name>
<name>
<surname>Abdulkareem</surname>
<given-names>K. H.</given-names>
</name>
<name>
<surname>Ali</surname>
<given-names>N. S.</given-names>
</name>
<name>
<surname>Al-Mhiqani</surname>
<given-names>M. N.</given-names>
</name>
<name>
<surname>Guger</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>EEG Feature Fusion for Motor Imagery: A New Robust Framework towards Stroke Patients Rehabilitation</article-title>. <source>Comput. Biol. Med.</source> <volume>137</volume>, <fpage>104799</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104799</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alar</surname>
<given-names>H. S.</given-names>
</name>
<name>
<surname>Fernandez</surname>
<given-names>P. L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Accurate and Efficient Mosquito Genus Classification Algorithm Using Candidate-Elimination and Nearest Centroid on Extracted Features of Wingbeat Acoustic Properties</article-title>. <source>Comput. Biol. Med.</source> <volume>139</volume>, <fpage>104973</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104973</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ali</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Akbar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ghulam</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Maher</surname>
<given-names>Z. A.</given-names>
</name>
<name>
<surname>Unar</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Talpur</surname>
<given-names>D. B.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>AFP-CMBPred: Computational Identification of Antifreeze Proteins by Extending Consensus Sequences into Multi-Blocks Evolutionary Information</article-title>. <source>Comput. Biol. Med.</source> <volume>139</volume>, <fpage>105006</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.105006</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alim</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rafay</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Naseem</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>PoGB-pred: Prediction of Antifreeze Proteins Sequences Using Amino Acid Composition with Feature Selection Followed by a Sequential-Based Ensemble Approach</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>3</issue>), <fpage>446</fpage>&#x2013;<lpage>456</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200707141926</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altuvia</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Schueler</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Margalit</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Ranking Potential Binding Peptides to MHC Molecules by a Computational Threading Approach</article-title>. <source>J. Mol. Biol.</source> <volume>249</volume> (<issue>2</issue>), <fpage>244</fpage>&#x2013;<lpage>250</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1995.0293</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altuvia</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sette</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sidney</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Southwood</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Margalit</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>A Structure-Based Algorithm to Predict Potential Binding Peptides to MHC Molecules with Hydrophobic Binding Pockets</article-title>. <source>Hum. Immunol.</source> <volume>58</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1016/s0198-8859(97)00210-3</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Awais</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hussain</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Rasool</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>Y. D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>iTSP-PseAAC: Identifying Tumor Suppressor Proteins by Using Fully Connected Neural Network and PseAAC</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>5</issue>), <fpage>700</fpage>&#x2013;<lpage>709</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666210108094431</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boehm</surname>
<given-names>K. M.</given-names>
</name>
<name>
<surname>Bhinder</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Raja</surname>
<given-names>V. J.</given-names>
</name>
<name>
<surname>Dephoure</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Elemento</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Predicting Peptide Presentation by Major Histocompatibility Complex Class I: an Improved Machine Learning Approach to the Immunopeptidome</article-title>. <source>BMC Bioinformatics</source> <volume>20</volume> (<issue>1</issue>), <fpage>7</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-018-2561-z</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breiman</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Random Forests</article-title>. <source>Mach Learn.</source> <volume>45</volume> (<issue>1</issue>), <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Buriro</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Ahmed</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Baloch</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ahmed</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shoorangiz</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Weddell</surname>
<given-names>S. J.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Classification of Alcoholic EEG Signals Using Wavelet Scattering Transform-Based Features</article-title>. <source>Comput. Biol. Med.</source> <volume>139</volume>, <fpage>104969</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104969</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Burton</surname>
<given-names>W. S.</given-names>
</name>
<name>
<surname>Myers</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hamilton</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Shelburne</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Banks</surname>
<given-names>S. A.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Automatic Tracking of Healthy Joint Kinematics from Stereo-Radiography Sequences</article-title>. <source>Comput. Biol. Med.</source> <volume>139</volume>, <fpage>104945</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104945</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Qian</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Identification and Analysis of Adenine N6-Methylation Sites in the rice Genome</article-title>. <source>Nat. Plants</source> <volume>4</volume> (<issue>8</issue>), <fpage>554</fpage>&#x2013;<lpage>563</lpage>. <pub-id pub-id-type="doi">10.1038/s41477-018-0214-x</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kurgan</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Prediction of Integral Membrane Protein Type by Collocated Hydrophobic Amino Acid Pairs</article-title>. <source>J. Comput. Chem.</source> <volume>30</volume> (<issue>1</issue>), <fpage>163</fpage>&#x2013;<lpage>172</lpage>. <pub-id pub-id-type="doi">10.1002/jcc.21053</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chou</surname>
<given-names>K-C.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Prediction of Protein Cellular Attributes Using Pseudo-amino Acid Composition</article-title>. <source>Proteins Struct. Funct. Bioinformatics</source> <volume>43</volume> (<issue>3</issue>), <fpage>246</fpage>&#x2013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1002/prot.1035</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chou</surname>
<given-names>K.-C.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Using Amphiphilic Pseudo Amino Acid Composition to Predict Enzyme Subfamily Classes</article-title>. <source>Bioinformatics</source> <volume>21</volume> (<issue>1</issue>), <fpage>10</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bth466</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dubchak</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Muchnik</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Holbrook</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S. H.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Prediction of Protein Folding Class Using Global Description of Amino Acid Sequence</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>92</volume> (<issue>19</issue>), <fpage>8700</fpage>&#x2013;<lpage>8704</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.92.19.8700</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Giuseppe</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>James</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Keith</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Guethlein</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Unni</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jim</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>IPD-MHC 2.0: an Improved Inter-species Database for the Study of the Major Histocompatibility Complex</article-title>. <source>Nucleic Acids Res.</source> <volume>45</volume> (<issue>D1</issue>), <fpage>D860</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1050</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hearst</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Dumais</surname>
<given-names>S. T.</given-names>
</name>
<name>
<surname>Osuna</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Support Vector Machines: Training and Applications</article-title>. <source>IEEE Intel. Syst. App.</source> <volume>13</volume> (<issue>4</issue>), <fpage>18</fpage>&#x2013;<lpage>28</lpage>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hopkins</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dutia</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Mcconnell</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>Monoclonal Antibodies to Sheep Lymphocytes. I. Identification of MHC Class II Molecules on Lymphoid Tissue and Changes in the Level of Class II Expression on Lymph-Borne Cells Following Antigen Stimulation <italic>In Vivo</italic>
</article-title>. <source>Immunology</source> <volume>59</volume> (<issue>3</issue>), <fpage>433</fpage> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Predicting MHC Class I Binder: Existing Approaches and a Novel Recurrent Neural Network Solution</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume> (<issue>6</issue>), <fpage>bbab216</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab216</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Karcioglu</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Bulut</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>The WM-Q Multiple Exact String Matching Algorithm for DNA Sequences</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>, <fpage>104656</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104656</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kubiniok</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Marcu</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bichmann</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kuchenbecker</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Schuster</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hamelin</surname>
<given-names>D. J.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Understanding the Constitutive Presentation of MHC Class I Immunopeptidomes in Primary Tissues</article-title>. <source>Iscience</source> <volume>25</volume> (<issue>2</issue>), <fpage>103768</fpage>. <pub-id pub-id-type="doi">10.1016/j.isci.2022.103768</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Niu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>An Improved MHC Identification Method with Extreme Learning Machine Algorithm</article-title>. <source>J. proteome Res.</source> <volume>18</volume> (<issue>3</issue>), <fpage>1392</fpage>&#x2013;<lpage>1401</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jproteome.9b00012</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Flower</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Quantitative Prediction of Mouse Class I MHC Peptide Binding Affinity Using Support Vector Machine Regression (SVR) Models</article-title>. <source>BMC Bioinformatics</source> <volume>7</volume> (<issue>1</issue>), <fpage>182</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-7-182</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lundegaard</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lamberth</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Harndahl</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Buus</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lund</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>NetMHC-3.0: Accurate Web Accessible Predictions of Human, Mouse and Monkey MHC Class I Affinities for Peptides of Length 8-11</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>W509</fpage>&#x2013;<lpage>W512</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn202</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Protein Function Prediction: From Traditional Classifier to Deep Learning</article-title>. <source>Proteomics</source> <volume>19</volume> (<issue>14</issue>), <fpage>e1900119</fpage>. <pub-id pub-id-type="doi">10.1002/pmic.201900119</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Anticancer Peptides Prediction with Deep Representation Learning Features</article-title>. <source>Brief Bioinform</source> <volume>22</volume> (<issue>5</issue>), <fpage>bbab008</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab008</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Convolutional Neural Network Using Dinucleotide One-Hot Encoder for Identifying DNA N6-Methyladenine Sites in the Rice Genome</article-title>. <source>Neurocomputing</source> <volume>422</volume>, <fpage>214</fpage>&#x2013;<lpage>221</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.09.056</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Sub-golgi Protein Localization by Use of Deep Representation Learning Features</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>24</issue>), <fpage>5600</fpage>&#x2013;<lpage>5609</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa1074</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maccari</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Robinson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bontrop</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Otting</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>de Groot</surname>
<given-names>N. G.</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>C. S.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>IPD-MHC: Nomenclature Requirements for the Non-human Major Histocompatibility Complex in the Next-Generation Sequencing Era</article-title>. <source>Immunogenetics</source> <volume>70</volume> (<issue>10</issue>), <fpage>619</fpage>&#x2013;<lpage>623</lpage>. <pub-id pub-id-type="doi">10.1007/s00251-018-1072-4</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mahoney</surname>
<given-names>K. E.</given-names>
</name>
<name>
<surname>Shabanowitz</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hunt</surname>
<given-names>D. F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MHC Phosphopeptides: Promising Targets for Immunotherapy of Cancer and Other Chronic Diseases</article-title>. <source>Mol. Cell Proteomics</source> <volume>20</volume> (<issue>640</issue>), <fpage>100112</fpage>. <pub-id pub-id-type="doi">10.1016/j.mcpro.2021.100112</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marcoux</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Laroche</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hasse</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bellio</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mbarik</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tamagne</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Platelet EVs Contain an Active Proteasome Involved in Protein Processing for Antigen Presentation via MHC-I Molecules</article-title>. <source>Blood J. Am. Soc. Hematol.</source> <volume>138</volume> (<issue>25</issue>), <fpage>2607</fpage>&#x2013;<lpage>2620</lpage>. <pub-id pub-id-type="doi">10.1182/blood.2020009957</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McShan</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Devlin</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Morozov</surname>
<given-names>G. I.</given-names>
</name>
<name>
<surname>Overall</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Moschidi</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Akella</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>TAPBPR Promotes Antigen Loading on MHC-I Molecules Using a Peptide Trap</article-title>. <source>Nat. Commun.</source> <volume>12</volume> (<issue>1</issue>), <fpage>3174</fpage>&#x2013;<lpage>3218</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-23225-6</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Naseer</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hussain</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>Y. D.</given-names>
</name>
<name>
<surname>Rasool</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>NPalmitoylDeep-Pseaac: A Predictor of N-Palmitoylation Sites in Proteins Using Deep Representations of Proteins and PseAAC via Modified 5-Steps Rule</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>2</issue>), <fpage>294</fpage>&#x2013;<lpage>305</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200605142828</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Platt</surname>
<given-names>J. C.</given-names>
</name>
</person-group> (<year>1999</year>).<source>Fast Training of Support Vector Machines Using Sequential Minimal Optimization, Advances in Kernel Methods</source>. <publisher-name>Support Vector Learning</publisher-name> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Roy</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Mazid</surname>
<given-names>M. I.</given-names>
</name>
<name>
<surname>Akhand</surname>
<given-names>R. N.</given-names>
</name>
<name>
<surname>Das</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Marufatuzzahan</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Identification and Host Response Interaction Study of SARS-CoV-2 Encoded miRNA-like Sequences: an In Silico Approach</article-title>. <source>Comput. Biol. Med.</source> <volume>134</volume>, <fpage>104451</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104451</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Safaei</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sundararajan</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Driss</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Boulila</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Shapi&#x27;i</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Systematic Literature Review on Obesity: Understanding the Causes &#x26; Consequences of Obesity and Reviewing Various Machine Learning Approaches Used to Predict Obesity</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>, <fpage>104754</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104754</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saxena</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Siddiqui</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Development of Machine Learning Based Blood-Brain Barrier Permeability Prediction Models Using Physicochemical Properties, MACCS and Substructure Fingerprints</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>6</issue>), <fpage>855</fpage>&#x2013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.2174/1574893616666210203104013</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saxena</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Animesh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fullwood</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mu</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>OnionMHC: A Deep Learning Model for Peptide - HLA-A&#x2a;02:01 Binding Predictions Using Both Structure and Sequence Feature Sets</article-title> <source>J. Micromech. Mol. Phys.</source> <volume>5</volume> (<issue>03</issue>), <fpage>2050009</fpage>. </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shiina</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Yamada</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Aarnink</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Suzuki</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Masuya</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ito</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Discovery of Novel MHC-Class I Alleles and Haplotypes in Filipino Cynomolgus Macaques (Macaca fascicularis) by Pyrosequencing and Sanger Sequencing</article-title>. <source>Immunogenetics</source> <volume>67</volume> (<issue>10</issue>), <fpage>563</fpage>&#x2013;<lpage>578</lpage>. <pub-id pub-id-type="doi">10.1007/s00251-015-0867-9</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tahoces</surname>
<given-names>P. G.</given-names>
</name>
<name>
<surname>Varela</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Carreira</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning Method for Aortic Root Detection</article-title>. <source>Comput. Biol. Med.</source> <volume>135</volume>, <fpage>104533</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104533</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tandel</surname>
<given-names>G. S.</given-names>
</name>
<name>
<surname>Tiwari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kakde</surname>
<given-names>O. G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Performance Optimisation of Deep Learning Models Using Majority Voting Algorithm for Brain Tumour Classification</article-title>. <source>Comput. Biol. Med.</source> <volume>135</volume>, <fpage>104564</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104564</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tavolara</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Gurcan</surname>
<given-names>M. N.</given-names>
</name>
<name>
<surname>Segal</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Niazi</surname>
<given-names>M. K. K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Identification of Difficult to Intubate Patients from Frontal Face Images Using an Ensemble of Deep Learning Models</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>, <fpage>104737</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104737</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Togacar</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Detection of Segmented Uterine Cancer Images by Hotspot Detection Method Using Deep Learning Models, Pigeon-Inspired Optimization, Types-Based Dominant Activation Selection Approaches</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>, <fpage>104659</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104659</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsiknakis</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Theodoropoulos</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Manikis</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ktistakis</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Boutsora</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Berto</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Deep Learning for Diabetic Retinopathy Detection and Classification Based on Fundus Images: A Review</article-title>. <source>Comput. Biol. Med.</source> <volume>135</volume>, <fpage>104599</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104599</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Turki</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Taguchi</surname>
<given-names>Y. h.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Discriminating the Single-Cell Gene Regulatory Networks of Human Pancreatic Islets: A Novel Deep Learning Application</article-title>. <source>Comput. Biol. Med.</source> <volume>132</volume>, <fpage>132</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104257</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Usman</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Khalid</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bashir</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Deep Learning Based Ensemble Learning Method for Epileptic Seizure Prediction</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104710</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vafaeezadeh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Behnam</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hosseinsabet</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gifani</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>A Deep Learning Approach for the Automatic Recognition of Prosthetic Mitral Valve in Echocardiographic Images</article-title>. <source>Comput. Biol. Med.</source> <volume>133</volume>, <fpage>104388</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104388</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ruan</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>DeepFusion-Rbp</surname>
</name>
</person-group> (<year>2021</year>). <article-title>DeepFusion-RBP: Using Deep Learning to Fuse Multiple Features to Identify RNA-Binding Protein Sequences</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>8</issue>), <fpage>1089</fpage>&#x2013;<lpage>1100</lpage>. <pub-id pub-id-type="doi">10.2174/1574893616666210618145121</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Watanabe</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sakaguchi</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Murata</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ishii</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning-Based Hounsfield Unit Value Measurement Method for Bolus Tracking Images in Cerebral Computed Tomography Angiography</article-title>. <source>Comput. Biol. Med.</source> <volume>137</volume>, <fpage>104824</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104824</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Westbrook</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Karl</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Wiseman</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Mate</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Koroleva</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Garcia</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>No Assembly Required: Full-Length MHC Class I Allele Discovery by PacBio Circular Consensus Sequencing</article-title>. <source>Hum. Immunol.</source> <volume>76</volume> (<issue>12</issue>), <fpage>891</fpage>&#x2013;<lpage>896</lpage>. <pub-id pub-id-type="doi">10.1016/j.humimm.2015.03.022</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yan</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Editorial: Feature Representation and Learning Methods with Applications in Protein Secondary Structure</article-title>. <source>Front. Bioeng. Biotechnol.</source> <volume>20219</volume> (<issue>822</issue>). <pub-id pub-id-type="doi">10.3389/fbioe.2021.748722</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yap</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Hachiuma</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Alavi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Br&#xfc;ngel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Cassidy</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Goyal</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Deep Learning in Diabetic Foot Ulcers Detection: A Comprehensive Evaluation</article-title>. <source>Comput. Biol. Med.</source> <volume>135</volume>, <fpage>104596</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104596</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yildirim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bozdag</surname>
<given-names>P. G.</given-names>
</name>
<name>
<surname>Talo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yildirim</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Karabatak</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Acharya</surname>
<given-names>U. R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Learning Model for Automated Kidney Stone Detection Using Coronal CT Images</article-title>. <source>Comput. Biol. Med.</source> <volume>135</volume>, <fpage>104569</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104569</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Prediction of lncRNA-Disease Associations Based on Robust Multi-Label Learning</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>9</issue>), <fpage>1179</fpage>&#x2013;<lpage>1189</lpage>. <pub-id pub-id-type="doi">10.2174/1574893616666210712091221</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Computational Traditional Chinese Medicine Diagnosis: A Literature Survey</article-title>. <source>Comput. Biol. Med.</source> <volume>133</volume>, <fpage>104358</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104358</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>REUR: A Unified Deep Framework for Signet Ring Cell Detection in Low-Resolution Pathological Images</article-title>. <source>Comput. Biol. Med.</source> <volume>136</volume>, <fpage>104711</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104711</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Duan</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>F.-X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MDAPlatform: A Component-Based Platform for Constructing and Assessing miRNA-Disease Association Prediction Methods</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>5</issue>), <fpage>710</fpage>&#x2013;<lpage>721</lpage>. <pub-id pub-id-type="doi">10.2174/1574893616999210120181506</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Qin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Self-supervised CT Super-resolution with Hybrid Model</article-title>. <source>Comput. Biol. Med.</source> <volume>138</volume>, <fpage>104775</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104775</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ju</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Bioluminescent Proteins Prediction with Voting Strategy</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>2</issue>), <fpage>240</fpage>&#x2013;<lpage>251</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200601122328</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>A CNN-Based Multi-Target Fast Classification Method for AR-SSVEP</article-title>. <source>Comput. Biol. Med.</source> <volume>141</volume>, <fpage>105042</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.105042</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pei</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Fuyi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Marquez-Lago</surname>
<given-names>T. T.</given-names>
</name>
<name>
<surname>Andr&#xe9;</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jerico</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>iLearn: an Integrated Platform and Meta-Learner for Feature Engineering, Machine Learning Analysis and Modeling of DNA, RNA and Protein Sequence Data</article-title>. <source>Brief. Bioinform.</source> <volume>21</volume> (<issue>3</issue>), <fpage>1047</fpage>&#x2013;<lpage>1057</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz041</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Fusing Multiple Biological Networks to Effectively Predict miRNA-Disease Associations</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>3</issue>), <fpage>371</fpage>&#x2013;<lpage>384</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200715165335</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>MK-FSVM-SVDD: A Multiple Kernel-Based Fuzzy SVM Model for Predicting DNA-Binding Proteins via Support Vector Data Description</article-title>. <source>Cbio</source> <volume>16</volume> (<issue>2</issue>), <fpage>274</fpage>&#x2013;<lpage>283</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200607173829</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>