<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">773202</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.773202</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>iAIPs: Identifying Anti-Inflammatory Peptides Using Random Forest</article-title>
<alt-title alt-title-type="left-running-head">Zhao et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Polarization in Atlantic Canada</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhao</surname>
<given-names>Dongxu</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1493739/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Teng</surname>
<given-names>Zhixia</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1010496/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Yanjuan</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1465511/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Chen</surname>
<given-names>Dong</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>College of Information and Computer Engineering, Northeast Forestry University, <addr-line>Harbin</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>College of Electrical and Information Engineering, Quzhou University, <addr-line>Quzhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/560593/overview">Juan Wang</ext-link>, Inner Mongolia University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1473670/overview">Haiying Zhang</ext-link>, Xiamen University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/563708/overview">Jiajie Peng</ext-link>, Northwestern Polytechnical University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Zhixia Teng, <email>tengzhixia@nefu.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>30</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>773202</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>08</day>
<month>10</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Zhao, Teng, Li and Chen.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Zhao, Teng, Li and Chen</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Recently, several anti-inflammatory peptides (AIPs) have been found in the process of the inflammatory response, and these peptides have been used to treat some inflammatory and autoimmune diseases. Therefore, identifying AIPs accurately from a given amino acid sequences is critical for the discovery of novel and efficient anti-inflammatory peptide-based therapeutics and the acceleration of their application in therapy. In this paper, a random forest-based model called iAIPs for identifying AIPs is proposed. First, the original samples were encoded with three feature extraction methods, including g-gap dipeptide composition (GDC), dipeptide deviation from the expected mean (DDE), and amino acid composition (AAC). Second, the optimal feature subset is generated by a two-step feature selection method, in which the feature is ranked by the analysis of variance (ANOVA) method, and the optimal feature subset is generated by the incremental feature selection strategy. Finally, the optimal feature subset is inputted into the random forest classifier, and the identification model is constructed. Experiment results showed that iAIPs achieved an AUC value of 0.822 on an independent test dataset, which indicated that our proposed model has better performance than the existing methods. Furthermore, the extraction of features for peptide sequences provides the basis for evolutionary analysis. The study of peptide identification is helpful to understand the diversity of species and analyze the evolutionary history of species.</p>
</abstract>
<kwd-group>
<kwd>anti-inflammatory peptides</kwd>
<kwd>random forest</kwd>
<kwd>feature extraction</kwd>
<kwd>evolutionary information</kwd>
<kwd>evolutionary analysis</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>As a part of the nonspecific immune response, inflammation response usually occurs in response to any type of bodily injury (<xref ref-type="bibr" rid="B12">Ferrero-Miliani et&#x20;al., 2007</xref>). When the inflammatory response occurs in the condition of no obvious infection, or when the response continues despite the resolution of the initial insult, the process may be pathological and leads to chronic inflammation (<xref ref-type="bibr" rid="B43">Patterson et&#x20;al., 2014</xref>). At present, the therapy for inflammatory and autoimmune diseases usually uses nonspecific anti-inflammatory drugs or other immunosuppressants, which may produce some side effects (<xref ref-type="bibr" rid="B53">Tabas and Glass, 2013</xref>; <xref ref-type="bibr" rid="B69">Yu et&#x20;al., 2021</xref>). Several endogenous peptides found in the process of inflammatory response have become anti-inflammatory agents and can be used as new therapies for autoimmune diseases and inflammatory disorders (<xref ref-type="bibr" rid="B17">Gonzalez-Rey et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B70">Yu et&#x20;al., 2020a</xref>). Compared with small-molecule drugs, the therapy based on peptides has minimal toxicity and high specificity under normal conditions, which is a better choice for inflammatory and autoimmune disorders and has been widely used in treatment (<xref ref-type="bibr" rid="B6">de la Fuente-N&#xfa;&#xf1;ez et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B46">Shang et&#x20;al., 2021</xref>).</p>
<p>Due to the biological importance of AIPs, many biochemical experimental methods have been developed for identifying AIPs. However, these biochemical methods usually need a long experimental cycle and have a high experimental cost. In recent years, machine learning has increasingly become the most popular tool in the field of bioinformatics (<xref ref-type="bibr" rid="B79">Zhao et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B33">Liu et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B34">Luo et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B52">Sun et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B78">Zhao et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B27">Jin et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B59">Wang et&#x20;al., 2021a</xref>). Many researchers have tried to adopt machine learning algorithms to identify AIPs only based on peptide amino acid sequence information. In 2017, Gupta et&#x20;al. proposed a predictor of AIPs based on the machine learning method. They constructed the combined features and inputted them in the SVM classifier to construct the prediction model (<xref ref-type="bibr" rid="B19">Gupta et&#x20;al., 2017</xref>).</p>
<p>In 2018, Manavalan et&#x20;al. proposed a novel prediction model called AIPpred. They encoded the original peptide sequence by the dipeptide composition (DPC) feature representation method, and then, they developed a random forest-based model to identify AIPs (<xref ref-type="bibr" rid="B37">Manavalan et&#x20;al., 2018</xref>). AIEpred is a novel prediction model and is proposed by Zhang et&#x20;al. AIEpred encodes peptide sequences based on three feature representations. Based on various feature representations, it constructed many base classifiers, which are the basis of ensemble classifier (<xref ref-type="bibr" rid="B75">Zhang et&#x20;al., 2020a</xref>).</p>
<p>In this paper, we proposed a novel identification model of AIPs for further improving the identification ability. First, we encoded the samples with multiple features consisting of AAC, DDE, and GDC. It has been proven that multiple features can effectively discriminate positive instances from negative ones in various biological problems. Second, we selected the optimal features based on a feature selection strategy, which has achieved better performance in many biological problems. Finally, we used the random forest classifier to construct an identification model based on the optimal features. The experimental result shows that our proposed method in this paper has better performance than the existing methods.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<p>
<xref ref-type="fig" rid="F1">Figure&#x20;1</xref> gives the general framework of iAIPs proposed in this paper. The framework consists of four steps as follows: 1) Dataset preparation&#x2014;It collects the data required for the experiment. 2) Feature extraction&#x2014;It converts the collected sequence data from step 1 into numerical features. 3) Feature selection&#x2014;removes redundant features from a feature set. 4) Prediction model construction. Each step of the framework will be described as follows.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The framework of iAIPs.</p>
</caption>
<graphic xlink:href="fgene-12-773202-g001.tif"/>
</fig>
<sec id="s2-1">
<title>2.1 Dataset Preparation</title>
<p>A high-quality dataset is critical to construct an effective and reliable prediction model. To measure the performance of our model by comparing it with other existing machine learning-based prediction models, we used the dataset with no change proposed in AIPpred (<xref ref-type="bibr" rid="B37">Manavalan et&#x20;al., 2018</xref>). The dataset was first retrieved from the IEDB database (<xref ref-type="bibr" rid="B28">Kim et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B55">Vita et&#x20;al., 2019</xref>), and then the samples with sequence identity &#x3e;80% (<xref ref-type="bibr" rid="B81">Zou et&#x20;al., 2020</xref>) are excluded by using CD-HIT (<xref ref-type="bibr" rid="B24">Huang et&#x20;al., 2010</xref>). The dataset contains 1,678 AIPs and 2,516&#x20;non-AIPs. For this dataset, it is randomly selected as the training dataset, which is inputted into the classifier and used to construct the identification model. The training dataset is also used to measure the cross-validation performance of our model. The remaining dataset is used as an independent dataset, which will be used to evaluate the generalization capability of our identification model. In detail, the training dataset consists of 1,258 AIPs and 1,887&#x20;non-AIPs, and the independent dataset consists of 420 AIPs and 629&#x20;non-AIPs.</p>
</sec>
<sec id="s2-2">
<title>2.2 Feature Extraction Methods</title>
<p>In the process of peptide identification, finding an effective feature extraction method is the most important step (<xref ref-type="bibr" rid="B31">Liu, 2019</xref>; <xref ref-type="bibr" rid="B15">Fu et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B5">Cai et&#x20;al., 2021</xref>). In this study, we tried a variety of feature extraction methods and used the random forest classifier to evaluate the performance of those methods. Finally, we chose three efficient feature extraction methods to encode peptide amino acid sequences, including amino acid composition, dipeptide deviation from expected mean, and g-gap dipeptide composition. The details of each feature extraction method are described as follows.</p>
<sec id="s2-2-1">
<title>2.2.1 Amino Acid Composition</title>
<p>Different peptide sequences consist of different amino acid sequences. AAC tried to count the composition information of peptides. In detail, AAC calculates the frequency of occurrence of each amino acid type (<xref ref-type="bibr" rid="B60">Wei et&#x20;al., 2018a</xref>; <xref ref-type="bibr" rid="B32">Liu et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B40">Ning et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B68">Yang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B77">Zhang and Zou, 2020</xref>; <xref ref-type="bibr" rid="B66">Wu and Yu, 2021</xref>). The computation formula of AAC is as follows:<disp-formula id="equ1">
<mml:math id="m1">
<mml:mrow>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mi mathvariant="normal">C</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>L</mml:mi>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mtext>&#x200a;</mml:mtext>
<mml:mi mathvariant="normal">j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>E</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>L</italic> denotes the length of the peptide, which is the number of characters in the peptide, <italic>AAC</italic> (<italic>j</italic>) denotes the percentage of amino acid j, <italic>N</italic> (<italic>j</italic>) denotes the total number of amino acid <italic>j</italic>. The dimension of AAC is&#x20;20.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Dipeptide Deviation From the Expected Mean</title>
<p>According to the dipeptide composition information, DDE computes deviation frequencies from expected mean values (<xref ref-type="bibr" rid="B45">Saravanan and Gautham, 2015</xref>). The feature vector extracted by DDE is generated by three parameters: theoretical variance (TV), dipeptide composition (DC), and theoretical mean (TM). The formulas of the three parameters are as follows:<disp-formula id="equ2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="normal">j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m3">
<mml:mrow>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> denotes the occurred frequency of dipeptide <italic>j</italic>, and <italic>L</italic> denotes the length of peptide sequences.<disp-formula id="equ3">
<mml:math id="m4">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>N</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>
<italic>C</italic>
<sub>
<italic>j</italic>1</sub> denotes the number of codons that encode for the first amino acid, and <italic>C</italic>
<sub>
<italic>j</italic>2</sub> denotes the number of codons that encode for the second amino acid in the dipeptide <italic>j</italic>. CN denotes the total number of possible codons.<disp-formula id="equ4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>V</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
<p>The formula of DDE(i) is as follows.<disp-formula id="equ5">
<mml:math id="m6">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>E</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>C</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>M</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>V</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
</p>
</sec>
<sec id="s2-2-3">
<title>2.2.3&#x20;G-Gap Dipeptide Composition</title>
<p>GDC is used to measure the correlation of two non-adjacent residues; its dimension is 400 (<xref ref-type="bibr" rid="B61">Wei et&#x20;al., 2018b</xref>). GDC can be represented as follows:<disp-formula id="equ6">
<mml:math id="m7">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>1</mml:mn>
<mml:mi>g</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>2</mml:mn>
<mml:mi>g</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mn>400</mml:mn>
</mml:mrow>
<mml:mi>g</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>where <inline-formula id="inf2">
<mml:math id="m8">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the frequency of v (v&#xa0;&#x3d;&#xa0;1,2, &#x2026;, 400), and it can be calculated as:<disp-formula id="equ7">
<mml:math id="m9">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mn>400</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>where <inline-formula id="inf3">
<mml:math id="m10">
<mml:mrow>
<mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> denotes the number of the v-th g-gap dipeptide in a given peptide. In this study, every peptide has a different length; the minimum length is 5. Therefore, we set the range of g from 1 to 4. For the different values of g, we represent the feature as GDC-gap1, GDC-gap2, GDC-gap3, and GDC-gap4.</p>
</sec>
</sec>
<sec id="s2-3">
<title>2.3 Feature Selection</title>
<p>In the <italic>Feature extraction methods</italic> section, we introduced the feature extraction method used in this paper. However, like other feature representation methods, our feature representation may also produce many noises (<xref ref-type="bibr" rid="B62">Wei et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B58">Wang et&#x20;al., 2020a</xref>; <xref ref-type="bibr" rid="B29">Li et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B54">Tang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B57">Wang et&#x20;al., 2021b</xref>). Recently, many feature selection methods for eliminating noise has been used to solve many bioinformatics problems (<xref ref-type="bibr" rid="B22">He et&#x20;al., 2020</xref>), such as TATA-binding protein prediction (<xref ref-type="bibr" rid="B80">Zou et&#x20;al., 2016</xref>), DNA 4mc site prediction (<xref ref-type="bibr" rid="B36">Manavalan et&#x20;al., 2019</xref>), antihypertensive peptide prediction (<xref ref-type="bibr" rid="B38">Manayalan et&#x20;al., 2019</xref>), drug-induced hepatotoxicity prediction (<xref ref-type="bibr" rid="B50">Su et&#x20;al., 2019</xref>), and enhance-promoter interaction prediction (<xref ref-type="bibr" rid="B23">Hong et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B39">Min et&#x20;al., 2021</xref>).</p>
<p>Likewise, we will use a two-step feature selection method to solve the noise of features. In detail, the feature is first ranked based on the ANOVA score. Then, based on the orderly features, we use the incremental feature selection (IFS) strategy to generate different feature subsets, the feature subset with optimal performance is selected as the optimal feature subset. In the <italic>Result and discussion</italic> section, we will give the experiments about feature extraction, in which we will verify the effectiveness of our feature representation.</p>
<sec id="s2-3-1">
<title>2.3.1 Analysis of Variance</title>
<p>In this work, the feature is first ranked based on the ANOVA score. For every feature, ANOVA calculated the ratio of the variance between groups and the variance within groups, which can test the mean difference between groups effectively (<xref ref-type="bibr" rid="B7">Ding et&#x20;al., 2014</xref>). The score is calculated as follows:<disp-formula id="equ8">
<mml:math id="m11">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>B</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>S</italic> (<italic>t</italic>) is the score of the feature t, <inline-formula id="inf4">
<mml:math id="m12">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>B</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the variance between groups, and <inline-formula id="inf5">
<mml:math id="m13">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the variance within groups. The formula of <inline-formula id="inf6">
<mml:math id="m14">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>B</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf7">
<mml:math id="m15">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>W</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is as follows:<disp-formula id="equ9">
<mml:math id="m16">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>B</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ10">
<mml:math id="m17">
<mml:mrow>
<mml:msubsup>
<mml:mi>S</mml:mi>
<mml:mi>w</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>K</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>m</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>K</italic> denotes the number of groups, and <italic>N</italic> denotes the total number of instances; <inline-formula id="inf8">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denote the value of the <italic>j</italic>-th sample in the <italic>i</italic>-th group of the feature&#x20;<italic>t</italic>.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Incremental Feature Selection</title>
<p>Based on the orderly features, we use the incremental feature selection strategy to generate different feature subsets; the feature subset with optimal performance is selected as the optimal feature subset. In the incremental feature selection method, the feature set is constructed as empty at first, and then the feature vector is added one by one from the ranked feature set. Meanwhile, the new feature set is inputted into a classifier, and then a prediction model is constructed. We evaluate the performance of the model according to some indicators. Finally, the feature subset with the optimal performance is considered as the optimal feature&#x20;set.</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4 Machine Learning Methods</title>
<p>In this paper, we utilized various ensemble learning classification algorithms to develop identification models, which contain random forest (<xref ref-type="bibr" rid="B44">Ru et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B56">Wang et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B2">Ao et&#x20;al., 2021</xref>), AdaBoost, Gradient Boost Decision Tree (<xref ref-type="bibr" rid="B71">Yu et&#x20;al., 2020b</xref>), LightGBM, and XGBoost. In addition, we also tried some traditional machine learning classification algorithms, such as logistic regression and Na&#xef;ve Bayes. The description of these methods is as follows.</p>
<sec id="s2-4-1">
<title>2.4.1 Random Forest</title>
<p>As one of the most powerful ensemble learning methods, random forest was proposed by <xref ref-type="bibr" rid="B4">Breiman (2001)</xref>. Due to its effectiveness, random forest has been widely used in bioinformatics areas. Random forest can solve regression and classification tasks. To solve the problem, random forest uses the random feature selection method to construct hundreds or thousands of decision trees (<xref ref-type="bibr" rid="B1">Akbar et&#x20;al., 2020</xref>). By voting on these decision trees, the final identification result is obtained. The random forest algorithm used in this paper is from WEKA (<xref ref-type="bibr" rid="B20">Hall et&#x20;al., 2008</xref>), and all parameters are default.</p>
</sec>
<sec id="s2-4-2">
<title>2.4.2 AdaBoost</title>
<p>The AdaBoost algorithm is an iterative algorithm, which was proposed by <xref ref-type="bibr" rid="B13">Freund (1990)</xref>. For a benchmark dataset, AdaBoost will train various weak classifiers and combine these weak classifiers by sample weight to construct a stronger final classifier. Among samples, low weights are assigned to easy samples that are classified correctly by the weak learner, while high weights are for the hard or misclassified samples. By constantly adjusting the weight of samples, AdaBoost will focus more on the samples that are classified incorrectly.</p>
</sec>
<sec id="s2-4-3">
<title>2.4.3 Gradient Boost Decision Tree</title>
<p>Similar to AdaBoost, Gradient Boost Decision Tree (GBDT) also combines weak learners to construct a prediction model (<xref ref-type="bibr" rid="B14">Friedman, 2001</xref>). Different from AdaBoost, GBDT will constantly adapt to the new model when the weak learners are learned. In detail, based on the negative gradient information of the loss function of the current model, the new weak classifier is trained. The training result is accumulated into the existing model to improve its performance (<xref ref-type="bibr" rid="B3">Basith et&#x20;al., 2018</xref>).</p>
</sec>
<sec id="s2-4-4">
<title>2.4.4 LightGBM and XGBoost</title>
<p>Both LightGBM and XGBoost are improved algorithms based on GBDT. LightGBM is mainly optimized in three aspects. The histogram algorithm is used to convert continuous features into discrete features, the gradient-based one-side sampling (GOSS) method is used to adjust the sample distribution and reduce the numbers of samples, and the exclusive feature bundling (EFB) is used to merge multiple independent features. XGBoost adds the second-order Taylor expansion and regularization term to the loss function.</p>
</sec>
<sec id="s2-4-5">
<title>2.4.5 Na&#xef;ve Bayes</title>
<p>Na&#xef;ve Bayes is a probabilistic classification algorithm based on Bayes&#x2019; theorem, which assumes that the features are independent of each other. According to this theorem, the probability of a given sample classified into class <italic>k</italic> can be calculated as<disp-formula id="equ11">
<mml:math id="m19">
<mml:mrow>
<mml:mi mathvariant="normal">P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>X</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>where the sample has the expression formula of {X,&#x20;C}.</p>
</sec>
<sec id="s2-4-6">
<title>2.4.6 Other Machine Learning Methods</title>
<p>Other traditional machine learning methods used for performance comparison include J48, logistic, SMO, and SGD. J48 is a decision tree algorithm provided in Weka, which is implemented based on the C4.5 idea. Logistic is a probability-based classification algorithm. Based on linear regression, Logistic introduces sigmoid function to limit the output value to [0,1] interval. SMO and SGD are optimization algorithms provided in Weka. SMO (sequential minimal optimization) is based on support vector machine (SVM), and SGD is based on linear regression.</p>
</sec>
</sec>
<sec id="s2-5">
<title>2.5 Performance Evaluation</title>
<p>To measure the performance of our proposed model, we chose four commonly used measurements: SN, SP, ACC, and MCC (<xref ref-type="bibr" rid="B26">Jiang et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B64">Wei et&#x20;al., 2017a</xref>; <xref ref-type="bibr" rid="B9">Ding et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B49">Shen et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B25">Huang et&#x20;al., 2020</xref>). These measurements are calculated as follows.<disp-formula id="equ12">
<mml:math id="m20">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ13">
<mml:math id="m21">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ14">
<mml:math id="m22">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="equ15">
<mml:math id="m23">
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>where <italic>FP</italic>, <italic>FN</italic>, <italic>TN</italic>, and <italic>TP</italic> show the number of false-positive, false-negative, true-negative, and true-positive, respectively. These are widely used in bioinformatics studies, such as protein fold recognition (<xref ref-type="bibr" rid="B48">Shao et&#x20;al., 2021</xref>), DNA-binding protein prediction (<xref ref-type="bibr" rid="B63">Wei et&#x20;al., 2017b</xref>), protein&#x2013;protein interaction prediction (<xref ref-type="bibr" rid="B65">Wei et&#x20;al., 2017c</xref>), and drug&#x2013;target interaction identification (<xref ref-type="bibr" rid="B10">Ding et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B8">Ding and JijunGuo, 2020</xref>).</p>
<p>Furthermore, we also used the receiver operating characteristic (ROC) curve (<xref ref-type="bibr" rid="B21">Hanley and McNeil, 1982</xref>; <xref ref-type="bibr" rid="B16">Fushing and Turnbull, 1996</xref>) to evaluate the performance of our proposed model. ROC computes the true-positive rate and low false-positive rate by setting various possible thresholds (<xref ref-type="bibr" rid="B18">Gribskov and Robinson, 1996</xref>). The area under the ROC curve (AUC) also shows the performance of the proposed model, which is more accurate in the aspect of evaluating the performance of the prediction model constructed by an imbalanced dataset.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and Discussion</title>
<p>To verify the effectiveness of our proposed model, we will measure the performance of our model from different perspectives. The detailed process of these experiments is presented as follows.</p>
<sec id="s3-1">
<title>3.1 Performance of Different Features</title>
<p>In this study, we use a variety of feature extraction methods and their combinations to encode peptide sequences. At first, we measure the effectiveness of single features. The comparison results of the fivefold cross-validation on the training dataset are shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Performance comparison of various single features.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feature</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Amino acid composition (AAC)</td>
<td align="char" char=".">0.529</td>
<td align="char" char=".">0.845</td>
<td align="char" char=".">0.719</td>
<td align="char" char=".">0.398</td>
<td align="char" char=".">0.760</td>
</tr>
<tr>
<td align="left">Dipeptide deviation for the expected mean (DDE)</td>
<td align="char" char=".">0.589</td>
<td align="char" char=".">0.854</td>
<td align="char" char=".">0.748</td>
<td align="char" char=".">0.464</td>
<td align="char" char=".">0.784</td>
</tr>
<tr>
<td align="left">G-gap dipeptide composition (GDC)-gap1</td>
<td align="char" char=".">0.456</td>
<td align="char" char=".">0.862</td>
<td align="char" char=".">0.700</td>
<td align="char" char=".">0.353</td>
<td align="char" char=".">0.764</td>
</tr>
<tr>
<td align="left">GDC-gap2</td>
<td align="char" char=".">0.466</td>
<td align="char" char=".">0.852</td>
<td align="char" char=".">0.697</td>
<td align="char" char=".">0.348</td>
<td align="char" char=".">0.751</td>
</tr>
<tr>
<td align="left">GDC-gap3</td>
<td align="char" char=".">0.454</td>
<td align="char" char=".">0.869</td>
<td align="char" char=".">0.703</td>
<td align="char" char=".">0.361</td>
<td align="char" char=".">0.741</td>
</tr>
<tr>
<td align="left">GDC-gap4</td>
<td align="char" char=".">0.449</td>
<td align="char" char=".">0.853</td>
<td align="char" char=".">0.692</td>
<td align="char" char=".">0.335</td>
<td align="char" char=".">0.733</td>
</tr>
<tr>
<td align="left">CKSAAGP</td>
<td align="char" char=".">0.477</td>
<td align="char" char=".">0.861</td>
<td align="char" char=".">0.707</td>
<td align="char" char=".">0.371</td>
<td align="char" char=".">0.732</td>
</tr>
<tr>
<td align="left">CTriad</td>
<td align="char" char=".">0.215</td>
<td align="char" char=".">0.897</td>
<td align="char" char=".">0.624</td>
<td align="char" char=".">0.155</td>
<td align="char" char=".">0.668</td>
</tr>
<tr>
<td align="left">GAAC</td>
<td align="char" char=".">0.533</td>
<td align="char" char=".">0.750</td>
<td align="char" char=".">0.663</td>
<td align="char" char=".">0.288</td>
<td align="char" char=".">0.679</td>
</tr>
<tr>
<td align="left">GDPC</td>
<td align="char" char=".">0.525</td>
<td align="char" char=".">0.826</td>
<td align="char" char=".">0.706</td>
<td align="char" char=".">0.370</td>
<td align="char" char=".">0.727</td>
</tr>
<tr>
<td align="left">GTPC</td>
<td align="char" char=".">0.470</td>
<td align="char" char=".">0.855</td>
<td align="char" char=".">0.701</td>
<td align="char" char=".">0.357</td>
<td align="char" char=".">0.742</td>
</tr>
<tr>
<td align="left">TPC</td>
<td align="char" char=".">0.304</td>
<td align="char" char=".">0.910</td>
<td align="char" char=".">0.668</td>
<td align="char" char=".">0.277</td>
<td align="char" char=".">0.739</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="table" rid="T1">Table&#x20;1</xref> shows that DDE is much better than other features according to the indicators of AUC, MCC, ACC, SP, and SN. In detail, the AUC value reaches 0.784, which is 2%&#x2013;11.6% higher than other features. Based on the indicator of AUC, the features of DDE, GDC-gap1, and AAC have the best performance.</p>
<p>To achieve better performance, we further test the performance of multiple features on the basis of DDE, GDC, and AAC. In detail, the GDC feature adopts four different parameters, that is, gap1, gap2, gap3, and gap4. The corresponding feature is GDC-gap1, GDC-gap2, GDC-gap3, and GDC-gap4. The performance comparison of the fivefold cross-validation on the training dataset is shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Performance comparison of various combined features of fivefold cross-validation on the training dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feature</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AAC&#x2b;DDE</td>
<td align="char" char=".">0.582</td>
<td align="char" char=".">0.857</td>
<td align="char" char=".">0.747</td>
<td align="char" char=".">0.461</td>
<td align="char" char=".">0.784</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap1</td>
<td align="char" char=".">0.483</td>
<td align="char" char=".">0.870</td>
<td align="char" char=".">0.715</td>
<td align="char" char=".">0.388</td>
<td align="char" char=".">0.770</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap2</td>
<td align="char" char=".">0.453</td>
<td align="char" char=".">0.871</td>
<td align="char" char=".">0.704</td>
<td align="char" char=".">0.363</td>
<td align="char" char=".">0.773</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap3</td>
<td align="char" char=".">0.435</td>
<td align="char" char=".">0.866</td>
<td align="char" char=".">0.694</td>
<td align="char" char=".">0.339</td>
<td align="char" char=".">0.759</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap4</td>
<td align="char" char=".">0.447</td>
<td align="char" char=".">0.873</td>
<td align="char" char=".">0.703</td>
<td align="char" char=".">0.360</td>
<td align="char" char=".">0.760</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap1</td>
<td align="char" char=".">0.586</td>
<td align="char" char=".">0.858</td>
<td align="char" char=".">0.749</td>
<td align="char" char=".">0.466</td>
<td align="char" char=".">0.790</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap2</td>
<td align="char" char=".">0.588</td>
<td align="char" char=".">0.854</td>
<td align="char" char=".">0.748</td>
<td align="char" char=".">0.464</td>
<td align="char" char=".">0.791</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap3</td>
<td align="char" char=".">0.583</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.749</td>
<td align="char" char=".">0.466</td>
<td align="char" char=".">0.785</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap4</td>
<td align="char" char=".">0.587</td>
<td align="char" char=".">0.851</td>
<td align="char" char=".">0.746</td>
<td align="char" char=".">0.459</td>
<td align="char" char=".">0.784</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap1</td>
<td align="char" char=".">0.585</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.750</td>
<td align="char" char=".">0.468</td>
<td align="char" char=".">0.794</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap2</td>
<td align="char" char=".">0.584</td>
<td align="char" char=".">0.852</td>
<td align="char" char=".">0.745</td>
<td align="char" char=".">0.457</td>
<td align="char" char=".">0.790</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap3</td>
<td align="char" char=".">0.593</td>
<td align="char" char=".">0.857</td>
<td align="char" char=".">0.751</td>
<td align="char" char=".">0.471</td>
<td align="char" char=".">0.784</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap4</td>
<td align="char" char=".">0.587</td>
<td align="char" char=".">0.855</td>
<td align="char" char=".">0.748</td>
<td align="char" char=".">0.464</td>
<td align="char" char=".">0.785</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to <xref ref-type="table" rid="T2">Table&#x20;2</xref>, the multiple features of AAC&#xa0;&#x2b;&#xa0;DDE&#xa0;&#x2b;&#xa0;GDC-gap1 has the best performance. Its value of SN, SP, ACC, MCC, and AUC are 0.585, 0.860, 0.750, 0.468, and 0.794, respectively.</p>
<p>To verify the performance of these combined features, we tested them on the independent test set. <xref ref-type="table" rid="T3">Table&#x20;3</xref> shows the experimental results on the independent dataset. The results show that the combined features of AAC &#x2b; DDE &#x2b; GDC-gap1 have the best performance on the independent dataset.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Performance comparison of various combined features on the independent dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Feature</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AAC&#x2b;DDE</td>
<td align="char" char=".">0.564</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.742</td>
<td align="char" char=".">0.450</td>
<td align="char" char=".">0.808</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap1</td>
<td align="char" char=".">0.488</td>
<td align="char" char=".">0.884</td>
<td align="char" char=".">0.725</td>
<td align="char" char=".">0.413</td>
<td align="char" char=".">0.799</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap2</td>
<td align="char" char=".">0.455</td>
<td align="char" char=".">0.878</td>
<td align="char" char=".">0.708</td>
<td align="char" char=".">0.373</td>
<td align="char" char=".">0.787</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap3</td>
<td align="char" char=".">0.448</td>
<td align="char" char=".">0.881</td>
<td align="char" char=".">0.707</td>
<td align="char" char=".">0.371</td>
<td align="char" char=".">0.795</td>
</tr>
<tr>
<td align="left">AAC&#x2b;GDC-gap4</td>
<td align="char" char=".">0.462</td>
<td align="char" char=".">0.865</td>
<td align="char" char=".">0.704</td>
<td align="char" char=".">0.362</td>
<td align="char" char=".">0.783</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap1</td>
<td align="char" char=".">0.569</td>
<td align="char" char=".">0.857</td>
<td align="char" char=".">0.742</td>
<td align="char" char=".">0.450</td>
<td align="char" char=".">0.812</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap2</td>
<td align="char" char=".">0.560</td>
<td align="char" char=".">0.854</td>
<td align="char" char=".">0.736</td>
<td align="char" char=".">0.437</td>
<td align="char" char=".">0.805</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap3</td>
<td align="char" char=".">0.576</td>
<td align="char" char=".">0.857</td>
<td align="char" char=".">0.745</td>
<td align="char" char=".">0.456</td>
<td align="char" char=".">0.808</td>
</tr>
<tr>
<td align="left">DDE&#x2b;GDC-gap4</td>
<td align="char" char=".">0.569</td>
<td align="char" char=".">0.857</td>
<td align="char" char=".">0.742</td>
<td align="char" char=".">0.450</td>
<td align="char" char=".">0.801</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap1</td>
<td align="char" char=".">0.56</td>
<td align="char" char=".">0.859</td>
<td align="char" char=".">0.739</td>
<td align="char" char=".">0.443</td>
<td align="char" char=".">0.806</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap2</td>
<td align="char" char=".">0.557</td>
<td align="char" char=".">0.855</td>
<td align="char" char=".">0.736</td>
<td align="char" char=".">0.437</td>
<td align="char" char=".">0.805</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap3</td>
<td align="char" char=".">0.552</td>
<td align="char" char=".">0.855</td>
<td align="char" char=".">0.734</td>
<td align="char" char=".">0.433</td>
<td align="char" char=".">0.806</td>
</tr>
<tr>
<td align="left">AAC&#x2b;DDE&#x2b;GDC-gap4</td>
<td align="char" char=".">0.567</td>
<td align="char" char=".">0.859</td>
<td align="char" char=".">0.742</td>
<td align="char" char=".">0.450</td>
<td align="char" char=".">0.801</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>3.2 Performance of Different Classifiers</title>
<p>In this study, we chose the random forest algorithm to construct the classifier. To verify the effectiveness of the random forest classifier, we compared its performance with other classifiers. We chose several ensemble classifiers that are similar to the random forest classifier, including AdaBoost, GBDT, LightGBM, and XGBoost. In addition, we also chose some machine learning classifiers, including J48, Logistic, SMO, SGD, and Na&#xef;ve Bayes.</p>
<p>Based on the best feature combination, which is obtained from previous experiments, we constructed different identification models using different classifiers. The performance of these classifiers on the training dataset is shown in <xref ref-type="table" rid="T4">Table&#x20;4</xref>.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Performance of various classifiers utilizing AAC-DDE-GDC-gap1 feature and fivefold cross-validation on the training dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Classifier</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Random forest</td>
<td align="char" char=".">0.585</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.750</td>
<td align="char" char=".">0.468</td>
<td align="char" char=".">0.794</td>
</tr>
<tr>
<td align="left">AdaBoost</td>
<td align="char" char=".">0.579</td>
<td align="char" char=".">0.743</td>
<td align="char" char=".">0.678</td>
<td align="char" char=".">0.324</td>
<td align="char" char=".">0.661</td>
</tr>
<tr>
<td align="left">Gradient Boost Decision Tree (GBDT)</td>
<td align="char" char=".">0.583</td>
<td align="char" char=".">0.788</td>
<td align="char" char=".">0.706</td>
<td align="char" char=".">0.379</td>
<td align="char" char=".">0.686</td>
</tr>
<tr>
<td align="left">LightGBM</td>
<td align="char" char=".">0.564</td>
<td align="char" char=".">0.754</td>
<td align="char" char=".">0.678</td>
<td align="char" char=".">0.321</td>
<td align="char" char=".">0.659</td>
</tr>
<tr>
<td align="left">XGBoost</td>
<td align="char" char=".">0.576</td>
<td align="char" char=".">0.757</td>
<td align="char" char=".">0.684</td>
<td align="char" char=".">0.336</td>
<td align="char" char=".">0.666</td>
</tr>
<tr>
<td align="left">J48</td>
<td align="char" char=".">0.552</td>
<td align="char" char=".">0.737</td>
<td align="char" char=".">0.663</td>
<td align="char" char=".">0.292</td>
<td align="char" char=".">0.647</td>
</tr>
<tr>
<td align="left">Logistic</td>
<td align="char" char=".">0.497</td>
<td align="char" char=".">0.677</td>
<td align="char" char=".">0.605</td>
<td align="char" char=".">0.175</td>
<td align="char" char=".">0.624</td>
</tr>
<tr>
<td align="left">Sequential minimal optimization (SMO)</td>
<td align="char" char=".">0.476</td>
<td align="char" char=".">0.725</td>
<td align="char" char=".">0.626</td>
<td align="char" char=".">0.206</td>
<td align="char" char=".">0.601</td>
</tr>
<tr>
<td align="left">SGD</td>
<td align="char" char=".">0.491</td>
<td align="char" char=".">0.689</td>
<td align="char" char=".">0.610</td>
<td align="char" char=".">0.182</td>
<td align="char" char=".">0.590</td>
</tr>
<tr>
<td align="left">Na&#xef;ve Bayes</td>
<td align="char" char=".">0.483</td>
<td align="char" char=".">0.684</td>
<td align="char" char=".">0.603</td>
<td align="char" char=".">0.168</td>
<td align="char" char=".">0.604</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results in <xref ref-type="table" rid="T4">Table&#x20;4</xref> show that the performance of the random forest classifier is the best, and its AUC value is 10.8%&#x2013;20.4% higher than other classifiers. To further compare the generalization ability of these classifiers, we test those models on the independent dataset. <xref ref-type="table" rid="T5">Table&#x20;5</xref> shows the experimental results. The results showed that the random forest classifier is also better than other classifiers on the independent dataset.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Performance of various classifiers based on AAC-DDE-GDC-gap1 feature on the independent dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Classifier</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Random forest</td>
<td align="char" char=".">0.560</td>
<td align="char" char=".">0.859</td>
<td align="char" char=".">0.739</td>
<td align="char" char=".">0.443</td>
<td align="char" char=".">0.806</td>
</tr>
<tr>
<td align="left">AdaBoost</td>
<td align="char" char=".">0.607</td>
<td align="char" char=".">0.809</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.426</td>
<td align="char" char=".">0.708</td>
</tr>
<tr>
<td align="left">GBDT</td>
<td align="char" char=".">0.640</td>
<td align="char" char=".">0.798</td>
<td align="char" char=".">0.735</td>
<td align="char" char=".">0.443</td>
<td align="char" char=".">0.719</td>
</tr>
<tr>
<td align="left">LightGBM</td>
<td align="char" char=".">0.538</td>
<td align="char" char=".">0.859</td>
<td align="char" char=".">0.730</td>
<td align="char" char=".">0.424</td>
<td align="char" char=".">0.698</td>
</tr>
<tr>
<td align="left">XGBoost</td>
<td align="char" char=".">0.579</td>
<td align="char" char=".">0.847</td>
<td align="char" char=".">0.740</td>
<td align="char" char=".">0.446</td>
<td align="char" char=".">0.713</td>
</tr>
<tr>
<td align="left">J48</td>
<td align="char" char=".">0.524</td>
<td align="char" char=".">0.738</td>
<td align="char" char=".">0.652</td>
<td align="char" char=".">0.266</td>
<td align="char" char=".">0.621</td>
</tr>
<tr>
<td align="left">Logistic</td>
<td align="char" char=".">0.498</td>
<td align="char" char=".">0.658</td>
<td align="char" char=".">0.594</td>
<td align="char" char=".">0.156</td>
<td align="char" char=".">0.615</td>
</tr>
<tr>
<td align="left">SMO</td>
<td align="char" char=".">0.442</td>
<td align="char" char=".">0.701</td>
<td align="char" char=".">0.598</td>
<td align="char" char=".">0.147</td>
<td align="char" char=".">0.572</td>
</tr>
<tr>
<td align="left">SGD</td>
<td align="char" char=".">0.493</td>
<td align="char" char=".">0.679</td>
<td align="char" char=".">0.604</td>
<td align="char" char=".">0.173</td>
<td align="char" char=".">0.586</td>
</tr>
<tr>
<td align="left">Na&#xef;ve Bayes</td>
<td align="char" char=".">0.486</td>
<td align="char" char=".">0.676</td>
<td align="char" char=".">0.600</td>
<td align="char" char=".">0.162</td>
<td align="char" char=".">0.602</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 The Analysis of Feature Selection</title>
<p>In the extracted features, some feature vectors may be noisy or redundant. To further improve the identification performance, we try to find optimal features by feature selection methods in this section. In this paper, the two-step feature selection strategy is used as the feature selection strategy to eliminate noise. In detail, we first used the ANOVA method to rank feature vectors, and then we used the IFS strategy to filter the optimal feature&#x20;set.</p>
<p>The comparison of performance before and after dimensionality reduction is shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. All indicators of the selected features have higher values than the original ones. The results suggest that the optimal feature set can improve the overall performance of our identification model and our fewer selected features can still accurately describe&#x20;AIPs.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Comparison of identification performance before and after dimensionality reduction.</p>
</caption>
<graphic xlink:href="fgene-12-773202-g002.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Comparison With Existing Methods</title>
<p>Independent dataset test plays an important role in testing the generalization ability of the identification model. Therefore, the independent dataset was used to measure our identification model; the performance of our identification model was compared with existing methods, which contains AntiInflam (<xref ref-type="bibr" rid="B12">Ferrero-Miliani et&#x20;al., 2007</xref>), AIPpred, and AIEpred. <xref ref-type="table" rid="T6">Table&#x20;6</xref> shows the detailed results of the different methods for identifying AIPs, where the results are ranked according to&#x20;AUC.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Performance of different identification models on the independent dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">SN</th>
<th align="center">SP</th>
<th align="center">ACC</th>
<th align="center">MCC</th>
<th align="center">AUC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AntiInflam (LA)</td>
<td align="char" char=".">0.258</td>
<td align="char" char=".">0.892</td>
<td align="char" char=".">0.638</td>
<td align="char" char=".">0.197</td>
<td align="char" char=".">0.647</td>
</tr>
<tr>
<td align="left">AntiInflam (MA)</td>
<td align="char" char=".">0.786</td>
<td align="char" char=".">0.417</td>
<td align="char" char=".">0.565</td>
<td align="char" char=".">0.210</td>
<td align="char" char=".">0.706</td>
</tr>
<tr>
<td align="left">AIEpred</td>
<td align="char" char=".">0.555</td>
<td align="char" char=".">0.899</td>
<td align="char" char=".">0.762</td>
<td align="char" char=".">0.495</td>
<td align="char" char=".">0.767</td>
</tr>
<tr>
<td align="left">AIPpred</td>
<td align="char" char=".">0.741</td>
<td align="char" char=".">0.746</td>
<td align="char" char=".">0.744</td>
<td align="char" char=".">0.479</td>
<td align="char" char=".">0.813</td>
</tr>
<tr>
<td align="left">iAIPs (our work)</td>
<td align="char" char=".">0.567</td>
<td align="char" char=".">0.874</td>
<td align="char" char=".">0.751</td>
<td align="char" char=".">0.471</td>
<td align="char" char=".">0.822</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="T6">Table&#x20;6</xref>, the value of our proposed identification model iAIPs in SN, SP, ACC, AUC, and MCC are 0.567, 0.874, 0.751, 0.822, and 0.471, respectively. Furthermore, the same independent dataset-based experimental results showed that the ACC of iAIPs was 0.007&#x2013;0.186 higher than that of AntiInflam and AIPpred, which is similar to AIEpred. Moreover, according to AUC, our performance is better than the other methods, which is 0.009&#x2013;0.175 higher than the others. The results indicate that our method has better performance than other existing prediction models.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>In this paper, an identifying AIP model based on peptide sequence is proposed. We tried various features and their combinations, utilized various commonly used ensemble learning classification algorithms and the two-step feature selection strategy. After trying a large number of experiments, we finally constructed an effective AIP prediction model. By conducting a large number of experiments on the training dataset and independent dataset, we verified that our proposed prediction model iAIPs could efficiently identify AIPs from the newly synthesized and discovered peptide sequences, which is better than the existing AIP prediction models.</p>
<p>In the future, the optimization of the feature representation method is a research direction. Especially, the research on a new feature representation method that can adaptively encode peptide sequences is of great significance. Furthermore, other optimization methods and computational intelligence models will be considered for identifying anti-inflammatory peptides. Deep learning (<xref ref-type="bibr" rid="B35">Lv et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B73">Zeng et&#x20;al., 2020a</xref>; <xref ref-type="bibr" rid="B74">Zeng et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B76">Zhang et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B11">Du et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B42">Pang and Liu, 2020</xref>), unsupervised learning (<xref ref-type="bibr" rid="B72">Zeng et&#x20;al., 2020c</xref>), and ensemble learning (<xref ref-type="bibr" rid="B51">Sultana et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B67">Zhong et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B30">Li et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B41">Niu et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B47">Shao and Liu, 2021</xref>) will be employed when the dataset is large enough.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. These data can be found here: <ext-link ext-link-type="uri" xlink:href="http://www.thegleelab.org/AIPpred/">http://www.thegleelab.org/AIPpred/</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>DZ and ZT conceptualized the study. DZ and YL formulated the methodology. DZ validated the study and wrote the original draft. DC and YL reviewed and edited the manuscript. ZT supervised the study and acquired the funding. All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work is supported by the Fundamental Research Funds for the Central Universities (2572018BH05, 2572017CB33), the National Natural Science Foundation of China (61901103, 28961671189), and the Natural Science Foundation of Heilongjiang Province (LH2019F002).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Akbar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ateeq Ur</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Maqsood</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mohammad</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>cACP: Classifying Anticancer Peptides Using Discriminative Intelligent Model via Chou&#x2019;s 5-step Rules and General Pseudo Components</article-title>. <source>Chemometrics Intell. Lab. Syst.</source> <volume>196</volume>, <fpage>103912</fpage>. <pub-id pub-id-type="doi">10.1016/j.chemolab.2019.103912</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <source>RFhy-m2G: Identification of RNA N2-Methylguanosine Modification Sites Based on Random forest and Hybrid Features</source>. <publisher-loc>Methods (San Diego, Calif.)</publisher-loc>: <publisher-name>Elsevier</publisher-name>. </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Basith</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Manavalan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>iGHBP: Computational Identification of Growth Hormone Binding Proteins from Sequences Using Extremely Randomised Tree</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>16</volume>, <fpage>412</fpage>&#x2013;<lpage>420</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2018.10.007</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breiman</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Random Forests</article-title>. <source>Machine Learn.</source> <volume>45</volume> (<issue>1</issue>), <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/a:1010933404324</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cai</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>ITP-pred: an Interpretable Method for Predicting, Therapeutic Peptides with Fused Features Low-Dimension Representation</article-title>. <source>Brief Bioinform</source> <volume>22</volume> (<issue>4</issue>), <fpage>bbaa367</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa367</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de la Fuente-N&#xfa;&#xf1;ez</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Silva</surname>
<given-names>O. N.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>T. K.</given-names>
</name>
<name>
<surname>Franco</surname>
<given-names>O. L.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Antimicrobial Peptides: Role in Human Disease and Potential as Immunotherapies</article-title>. <source>Pharmacol. Ther.</source> <volume>178</volume>, <fpage>132</fpage>&#x2013;<lpage>140</lpage>. <pub-id pub-id-type="doi">10.1016/j.pharmthera.2017.04.002</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>P.-M.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Identification of Bacteriophage Virion Proteins by the ANOVA Feature Selection and Analysis</article-title>. <source>Mol. Biosyst.</source> <volume>10</volume> (<issue>8</issue>), <fpage>2229</fpage>&#x2013;<lpage>2235</lpage>. <pub-id pub-id-type="doi">10.1039/c4mb00316k</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y. T.</given-names>
</name>
<name>
<surname>JijunGuo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Drug-Target Interactions via Dual Laplacian Regularized Least Squares with Multiple Kernel Fusion</article-title>. <source>Knowledge-Based Syst.</source>, <fpage>204</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106254</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Identification of Drug-Side Effect Association via Multiple Information Integration with Centered Kernel Alignment</article-title>. <source>Neurocomputing</source> <volume>325</volume>, <fpage>211</fpage>&#x2013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2018.10.028</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Drug-Target Interactions via Fuzzy Bipartite Local Model</article-title>. <source>Neural Comput. Applic</source> <volume>32</volume>, <fpage>10303</fpage>&#x2013;<lpage>10319</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-019-04569-z</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Uversky</surname>
<given-names>V. N.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Classification of Chromosomal DNA Sequences Using Hybrid Deep Learning Architectures</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume> (<issue>10</issue>), <fpage>1130</fpage>&#x2013;<lpage>1136</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666200224095531</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ferrero-Miliani</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nielsen</surname>
<given-names>O. H.</given-names>
</name>
<name>
<surname>Andersen</surname>
<given-names>P. S.</given-names>
</name>
<name>
<surname>Girardin</surname>
<given-names>S. E.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Chronic Inflammation: Importance of NOD2 and NALP3 in Interleukin-1beta Generation</article-title>. <source>Clin. Exp. Immunol.</source> <volume>147</volume> (<issue>2</issue>), <fpage>227</fpage>&#x2013;<lpage>235</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2249.2006.03261.x</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freund</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Boosting a Weak Learning Algorithm by Majority</article-title>. <source>Inf. Comput.</source> <volume>121</volume> (<issue>2</issue>), <fpage>256</fpage>&#x2013;<lpage>285</lpage>. <pub-id pub-id-type="doi">10.1006/inco.1995.1136</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Friedman</surname>
<given-names>J.&#x20;H.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Greedy Function Approximation: A Gradient Boosting Machine</article-title>. <source>Ann. Stat.</source> <volume>29</volume> (<issue>5</issue>), <fpage>1189</fpage>&#x2013;<lpage>1232</lpage>. <pub-id pub-id-type="doi">10.1214/aos/1013203451</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>StackCPPred: a Stacking and Pairwise Energy Content-Based Prediction of Cell-Penetrating Peptides and Their Uptake Efficiency</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>10</issue>), <fpage>3028</fpage>&#x2013;<lpage>3034</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa131</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fushing</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Turnbull</surname>
<given-names>B. W.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Nonparametric and Semiparametric Estimation of the Receiver Operating Characteristic Curve</article-title>. <source>Ann. Stat.</source> <volume>24</volume> (<issue>1</issue>), <fpage>25</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1214/aos/1033066197</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gonzalez-Rey</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Anderson</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Delgado</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Emerging Roles of Vasoactive Intestinal Peptide: a New Approach for Autoimmune Therapy</article-title>. <source>Ann. Rheum. Dis.</source> <volume>66</volume> (<issue>3</issue>), <fpage>iii70</fpage>&#x2013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1136/ard.2007.078519</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gribskov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Robinson</surname>
<given-names>N. L.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Use of Receiver Operating Characteristic (ROC) Analysis to Evaluate Sequence Matching</article-title>. <source>Comput. Chem.</source> <volume>20</volume> (<issue>1</issue>), <fpage>25</fpage>&#x2013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1016/s0097-8485(96)80004-0</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gupta</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>A. K.</given-names>
</name>
<name>
<surname>Shastri</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Madhu</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>V. K.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Prediction of Anti-inflammatory Proteins/peptides: an Insilico Approach</article-title>. <source>J.&#x20;Transl Med.</source> <volume>15</volume> (<issue>1</issue>), <fpage>7</fpage>. <pub-id pub-id-type="doi">10.1186/s12967-016-1103-6</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hall</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Eibe</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Geoffrey</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bernhard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Peter</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Witten</surname>
<given-names>I. H.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>The WEKA Data Mining Software: An Update</article-title>. <source>ACM SIGKDD Explorations Newsl.</source> <volume>11</volume> (<issue>1</issue>), <fpage>10</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1145/1656274.1656278</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hanley</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>McNeil</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>1982</year>). <article-title>The Meaning and Use of the Area under a Receiver Operating Characteristic (ROC) Curve</article-title>. <source>Radiology</source> <volume>143</volume> (<issue>1</issue>), <fpage>29</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1148/radiology.143.1.7063747</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fei</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Quan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Hui</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>MRMD2.0: A Python Tool for Machine Learning with Feature Ranking and Reduction</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume> (<issue>10</issue>), <fpage>1213</fpage>&#x2013;<lpage>1221</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200503030350</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identifying Enhancer-Promoter Interactions with Neural Network Based on Pre-trained DNA Vectors and Attention Mechanism</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>4</issue>), <fpage>1037</fpage>&#x2013;<lpage>1043</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz694</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Niu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>CD-HIT Suite: a Web Server for Clustering and Comparing Biological Sequences</article-title>. <source>Bioinformatics</source> <volume>26</volume> (<issue>5</issue>), <fpage>680</fpage>&#x2013;<lpage>682</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq003</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Prediction of Transcription Factors Binding Events Based on Epigenetic Modifications in Different Human Cells</article-title>. <source>Epigenomics</source> <volume>12</volume> (<issue>16</issue>), <fpage>1443</fpage>&#x2013;<lpage>1456</lpage>. <pub-id pub-id-type="doi">10.2217/epi-2019-0321</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Predicting Human microRNA-Disease Associations Based on Support Vector Machine</article-title>. <source>Ijdmb</source> <volume>8</volume> (<issue>3</issue>), <fpage>282</fpage>&#x2013;<lpage>293</lpage>. <pub-id pub-id-type="doi">10.1504/ijdmb.2013.056078</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Application of Deep Learning Methods in Biological Networks</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume> (<issue>2</issue>), <fpage>1902</fpage>&#x2013;<lpage>1917</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa043</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ponomarenko</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Tamang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Greenbaum</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Immune Epitope Database Analysis Resource</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>W525</fpage>&#x2013;<lpage>W530</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks438</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Pu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>DeepATT: a Hybrid Category Attention Neural Network for Identifying Functional Effects of DNA Sequences</article-title>. <source>Brief Bioinform</source> <volume>21</volume>, <fpage>8</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa159</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>EP3: An Ensemble Predictor that Accurately Identifies Type III Secreted Effectors</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume> (<issue>2</issue>), <fpage>1918</fpage>&#x2013;<lpage>1928</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa008</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>BioSeq-Analysis: a Platform for DNA, RNA and Protein Sequence Analysis Based on Machine Learning Approaches</article-title>. <source>Brief. Bioinform.</source> <volume>20</volume> (<issue>4</issue>), <fpage>1280</fpage>&#x2013;<lpage>1294</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbx165</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>BioSeq-Analysis2.0: an Updated Platform for Analyzing DNA, RNA and Protein Sequences at Sequence Level and Residue Level Based on Machine Learning Approaches</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>(<issue>20</issue>): p. <fpage>e127</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz740</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yalin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Guohua</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yadong</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Deep Learning Approach for Filtering Structural Variants in Short Read Sequencing Data</article-title>. <source>Brief Bioinform</source> <volume>22</volume> (<issue>4</issue>). <pub-id pub-id-type="doi">10.1093/bib/bbaa370</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Luo</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Methylation States of DNA Regions for Illumina Methylation BeadChip</article-title>. <source>BMC Genomics</source> <volume>21</volume>(Suppl<issue>. 1</issue>): p. <fpage>672</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-019-6019-0</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Protein Function Prediction: From Traditional Classifier to Deep Learning</article-title>. <source>Proteomics</source> <volume>19</volume> (<issue>14</issue>), <fpage>e1900119</fpage>. <pub-id pub-id-type="doi">10.1002/pmic.201900119</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manavalan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Basith</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Meta-4mCpred: A Sequence-Based Meta-Predictor for Accurate DNA 4mC Site Prediction Using Effective Feature Representation</article-title>. <source>Mol. Ther. - Nucleic Acids</source> <volume>16</volume>, <fpage>733</fpage>&#x2013;<lpage>744</lpage>. <pub-id pub-id-type="doi">10.1016/j.omtn.2019.04.019</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manavalan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>T. H.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>M. O.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>AIPpred: Sequence-Based Prediction of Anti-inflammatory Peptides Using Random Forest</article-title>. <source>Front. Pharmacol.</source> <volume>9</volume>, <fpage>276</fpage>. <pub-id pub-id-type="doi">10.3389/fphar.2018.00276</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manayalan</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Shaherin</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Tae Hwan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Leyi</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Gwang</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>mAHTPred: a Sequence-Based Meta-Predictor for Improving the Prediction of Anti-hypertensive Peptides Using Effective Feature Representation</article-title>. <source>Bioinformatics</source> <volume>35</volume> (<issue>16</issue>), <fpage>2757</fpage>&#x2013;<lpage>2765</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty1047</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Min</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Predicting Enhancer-Promoter Interactions by Deep Learning and Matching Heuristic</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume>. <pub-id pub-id-type="doi">10.1093/bib/bbaa254</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ning</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>An In Silico Immunogenicity Analysis for PbHRH: An Antiangiogenic Peptibody by Fusing HRH Peptide and Human IgG1 Fc Fragment</article-title>. <source>Cbio</source> <volume>15</volume> (<issue>6</issue>), <fpage>547</fpage>&#x2013;<lpage>553</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666190730104348</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>sgRNACNN: Identifying sgRNA On-Target Activity in Four Crops Using Ensembles of Convolutional Neural Networks</article-title>. <source>Plant Mol. Biol.</source> <volume>105</volume> (<issue>4-5</issue>), <fpage>483</fpage>&#x2013;<lpage>495</lpage>. <pub-id pub-id-type="doi">10.1007/s11103-020-01102-y</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group>, (<year>2020</year>). <article-title>SelfAT-Fold: Protein Fold Recognition Based on Residue-Based and Motif-Based Self-Attention Networks</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>1</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2020.3031888</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Patterson</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Nibbs</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>McInnes</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Siebert</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Protein Kinase Inhibitors in the Treatment of Inflammatory and Autoimmune Diseases</article-title>. <source>Clin. Exp. Immunol.</source> <volume>176</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1111/cei.12248</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ru</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Incorporating Distance-Based Top-N-Gram and Random Forest to Identify Electron Transport Proteins</article-title>. <source>J.&#x20;Proteome Res.</source> <volume>18</volume> (<issue>7</issue>), <fpage>2931</fpage>&#x2013;<lpage>2939</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jproteome.9b00250</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saravanan</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Gautham</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Harnessing Computational Biology for Exact Linear B-Cell Epitope Prediction: A Novel Amino Acid Composition-Based Feature Descriptor</article-title>. <source>OMICS: A J.&#x20;Integr. Biol.</source> <volume>19</volume> (<issue>10</issue>), <fpage>648</fpage>&#x2013;<lpage>658</lpage>. <pub-id pub-id-type="doi">10.1089/omi.2015.0095</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Prediction of Drug-Target Interactions Based on Multi-Layer Network Representation Learning</article-title>. <source>Neurocomputing</source> <volume>434</volume>, <fpage>80</fpage>&#x2013;<lpage>89</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.12.068</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>ProtFold-DFG: Protein Fold Recognition by Combining Directed Fusion Graph and PageRank Algorithm</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume>, <fpage>32</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa192</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>FoldRec-C2C: Protein Fold Recognition by Combining Cluster-To-Cluster Model and Protein Similarity Network</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume>, <fpage>32</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa144</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Identification of Protein Subcellular Localization via Integrating Evolutionary and Physicochemical Information into Chou&#x27;s General PseAAC</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>462</volume>, <fpage>230</fpage>&#x2013;<lpage>239</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2018.11.012</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Developing a Multi-Dose Computational Model for Drug-Induced Hepatotoxicity Prediction Based on Toxicogenomics Data</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>16</volume> (<issue>4</issue>), <fpage>1231</fpage>&#x2013;<lpage>1239</lpage>. <pub-id pub-id-type="doi">10.1109/tcbb.2018.2858756</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sultana</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>K. P.</given-names>
</name>
<name>
<surname>Verma</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Sequential Ensemble Model for Communicable Disease Forecasting</article-title>. <source>Cbio</source> <volume>15</volume> (<issue>4</issue>), <fpage>309</fpage>&#x2013;<lpage>317</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666191202153824</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lei</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Quan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Guohua</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>BP4RNAseq: a Babysitter Package for Retrospective and Newly Generated RNA-Seq Data Analyses Using Both Alignment-Based and Alignment-free Quantification Method</article-title>. <source>Bioinformatics</source> <volume>37</volume> (<issue>9</issue>), <fpage>1319</fpage>&#x2013;<lpage>1321</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa832</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tabas</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Glass</surname>
<given-names>C. K.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Anti-inflammatory Therapy in Chronic Disease: Challenges and Opportunities</article-title>. <source>Science</source> <volume>339</volume> (<issue>6116</issue>), <fpage>166</fpage>&#x2013;<lpage>172</lpage>. <pub-id pub-id-type="doi">10.1126/science.1230720</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>Y.-J.</given-names>
</name>
<name>
<surname>Pang</surname>
<given-names>Y.-H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Idp-Seq2Seq</surname>
</name>
</person-group> (<year>2020</year>). <article-title>IDP-Seq2Seq: Identification of Intrinsically Disordered Regions Based on Sequence to Sequence Learning</article-title>. <source>Bioinformaitcs</source> <volume>36</volume> (<issue>21</issue>), <fpage>5177</fpage>&#x2013;<lpage>5186</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa667</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vita</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mahajan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Overton</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>Dhanda</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Martini</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cantrell</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>The Immune Epitope Database (IEDB): 2018 Update</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume> (<issue>D1</issue>), <fpage>D339</fpage>&#x2013;<lpage>D343</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1006</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Its2vec: Fungal Species Identification Using Sequence Embedding and Random Forest Classification</article-title>. <source>Biomed. Res. Int.</source> <volume>2020</volume>, <fpage>1</fpage>&#x2013;<lpage>11</lpage>. <pub-id pub-id-type="doi">10.1155/2020/2468789</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>H. D.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Identify RNA-Associated Subcellular Localizations Based on Multi-Label Learning Using Chou&#x27;s 5-steps Rule</article-title>. <source>BMC Genomics</source> <volume>22</volume>(<issue>56</issue>): p. <fpage>1</fpage>-1.<pub-id pub-id-type="doi">10.1186/s12864-020-07347-7</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Membrane Protein Types via Multivariate Information Fusion with Hilbert-Schmidt Independence Criterion</article-title>. <source>Neurocomputing</source> <volume>383</volume>, <fpage>257</fpage>&#x2013;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.11.103</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jian</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guohua</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>The Stacking Strategy-Based Hybrid Framework for Identifying Non-coding RNAs</article-title>. <source>Brief Bioinform</source> <volume>22</volume> (<issue>5</issue>), <fpage>32</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab023</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>ACPred-FL: a Sequence-Based Predictor Using Effective Feature Representation to Improve the Prediction of Anti-cancer Peptides</article-title>. <source>Bioinformatics</source> <volume>34</volume> (<issue>23</issue>), <fpage>4007</fpage>&#x2013;<lpage>4016</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty451</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jie</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fuyi</surname>
<given-names>Li.</given-names>
</name>
<name>
<surname>Jiangning</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ran</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Quan</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Comparative Analysis and Prediction of Quorum-Sensing Peptides Using Feature Representation Learning and Machine Learning Algorithms</source>, <volume>21</volume>. <publisher-name>Brief Bioinform</publisher-name>, <fpage>106</fpage>&#x2013;<lpage>119</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bby107</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Improved and Promising Identification of Human MicroRNAs by Incorporating a High-Quality Negative Set</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>11</volume> (<issue>1</issue>), <fpage>192</fpage>&#x2013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.1109/tcbb.2013.146</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Local-DPP: An Improved DNA-Binding Protein Prediction Method by Exploring Local Evolutionary Information</article-title>. <source>Inf. Sci.</source> <volume>384</volume>, <fpage>135</fpage>&#x2013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2016.06.026</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>K. K.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Novel Hierarchical Selective Ensemble Classifier with Bioinformatics Application</article-title>. <source>Artif. Intelligence Med.</source> <volume>83</volume>, <fpage>82</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2017.02.005</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Improved Prediction of Protein-Protein Interactions Using Novel Negative Samples, Features, and an Ensemble Classifier</article-title>. <source>Artif. Intelligence Med.</source> <volume>83</volume>, <fpage>67</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2017.03.001</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <source>EPSOL: Sequence-Based Protein Solubility Prediction Using Multidimensional Embedding</source>. <publisher-loc>Oxford, England)</publisher-loc>: <publisher-name>Bioinformatics</publisher-name>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab463</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Predicting Disease-Associated Circular RNAs Using Deep Forests Combined with Positive-Unlabeled Learning Methods</article-title>. <source>Brief. Bioinformatics</source> <volume>21</volume> (<issue>4</issue>), <fpage>1425</fpage>&#x2013;<lpage>1436</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz080</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Identification of Cancerlectins by Using Cascade Linear Discriminant Analysis and Optimal G-gap Tripeptide Composition</article-title>. <source>Cbio</source> <volume>15</volume> (<issue>6</issue>), <fpage>528</fpage>&#x2013;<lpage>537</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666190730103156</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Predicting Therapeutic Drugs for Hepatocellular Carcinoma Based on Tissue-specific Pathways</article-title>. <source>Plos Comput. Biol.</source> <volume>17</volume> (<issue>2</issue>), <fpage>e1008696</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1008696</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Exploring Drug Treatment Patterns Based on the Action of Drug and Multilayer Network Model</article-title>. <source>Ijms</source> <volume>21</volume> (<issue>14</issue>), <fpage>5014</fpage>. <pub-id pub-id-type="doi">10.3390/ijms21145014</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Jianguo</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Mingming</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Qing</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Exploiting XGBoost for Predicting Enhancer-Promoter Interactions</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume> (<issue>9</issue>), <fpage>1036</fpage>&#x2013;<lpage>1045</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666200120103948</pub-id> </citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yen</surname>
<given-names>G. G.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Consensus Community-Based Particle Swarm Optimization for Dynamic Community Detection</article-title>. <source>IEEE Trans. Cybern.</source> <volume>50</volume> (<issue>6</issue>), <fpage>2502</fpage>&#x2013;<lpage>2513</lpage>. <pub-id pub-id-type="doi">10.1109/tcyb.2019.2938895</pub-id> </citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yinglai</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Yuying</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Linyuan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xiaoping</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rodriguez-Paton</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Deep Collaborative Filtering for Prediction of Disease Genes</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinform.</source> <volume>17</volume> (<issue>5</issue>), <fpage>1639</fpage>&#x2013;<lpage>1647</lpage>. <pub-id pub-id-type="doi">10.1109/tcbb.2019.2907536</pub-id> </citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Target Identification Among Known Drugs by Deep Learning from Heterogeneous Networks</article-title>. <source>Chem. Sci.</source> <volume>11</volume> (<issue>7</issue>), <fpage>1775</fpage>&#x2013;<lpage>1797</lpage>. <pub-id pub-id-type="doi">10.1039/c9sc04336e</pub-id> </citation>
</ref>
<ref id="B75">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zehua</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lianrong</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Jijun</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Fei</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2020</year>). <source>AIEpred: An Ensemble Predictive Model of Classifier Chain to Identify Anti-inflammatory Peptides</source>, <volume>18</volume>. <publisher-name>IEEE/ACM Trans Comput Biol Bioinform</publisher-name>, <fpage>1831</fpage>&#x2013;<lpage>1840</lpage>. <pub-id pub-id-type="doi">10.1109/tcbb.2020.2968419</pub-id> </citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jianrong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Siyu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Meiqin</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Dongrui</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Min</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Review of the Applications of Deep Learning in Bioinformatics</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume> (<issue>8</issue>), <fpage>898</fpage>&#x2013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200711165743</pub-id> </citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y. P.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>PPTPP: A Novel Therapeutic Peptide Prediction Method Using Physicochemical Property Encoding and Adaptive Feature Representation Learning</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>13</issue>), <fpage>3982</fpage>&#x2013;<lpage>3987</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa275</pub-id> </citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Jiao</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>ECFS-DEA: an Ensemble Classifier-Based Feature Selection for Differential Expression Analysis on Expression Profiles</article-title>. <source>BMC Bioinformatics</source> <volume>21</volume> (<issue>1</issue>), <fpage>43</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3388-y</pub-id> </citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Methods of MicroRNA Promoter Prediction and Transcription Factor Mediated Regulatory Network</article-title>. <source>Biomed. Res. Int.</source> <volume>2017</volume>, <fpage>7049406</fpage>. <pub-id pub-id-type="doi">10.1155/2017/7049406</pub-id> </citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ju</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Pretata: Predicting TATA Binding Proteins with Novel Features and Dimensionality Reduction Strategy</article-title>. <source>BMC Syst. Biol.</source> <volume>10</volume>, <fpage>114</fpage>. <pub-id pub-id-type="doi">10.1186/s12918-016-0353-5</pub-id> </citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Gang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xingpeng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xiangrong</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xiangxiang</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Sequence Clustering in Bioinformatics: an Empirical Study</article-title>. <source>Brief. Bioinform.</source> <volume>21</volume> (<issue>1</issue>), <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bby090</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>