<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Pharmacol.</journal-id>
<journal-title>Frontiers in Pharmacology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Pharmacol.</abbrev-journal-title>
<issn pub-type="epub">1663-9812</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">771808</article-id>
<article-id pub-id-type="doi">10.3389/fphar.2021.771808</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Pharmacology</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>DrugHybrid_BS: Using Hybrid Feature Combined With Bagging-SVM to Predict Potentially Druggable Proteins</article-title>
<alt-title alt-title-type="left-running-head">Gong et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">DrugHybrid_BS</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Gong</surname>
<given-names>Yuxin</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1469410/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Liao</surname>
<given-names>Bo</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/570131/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Peng</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zou</surname>
<given-names>Quan</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/531759/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>School of Mathematics and Statistics, Hainan Normal University, <addr-line>Haikou</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Key Laboratory of Computational Science and Application of Hainan Province, <addr-line>Haikou</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Key Laboratory of Data Science and Smart Education, Hainan Normal University, Ministry of Education, <addr-line>Haikou</addr-line>, <country>China</country>
</aff>
<aff id="aff4">
<label>
<sup>4</sup>
</label>Yangtze Delta Region Institute (Quzhou), University of Electronic Science and Technology of China, <addr-line>Quzhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/567625/overview">Xiujuan Lei</ext-link>, Shaanxi Normal University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/585503/overview">Jiawei Luo</ext-link>, Hunan University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/498015/overview">Chaoyang Zhang</ext-link>, University of Southern Mississippi, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Bo Liao, <email>dragonbw@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Experimental Pharmacology and Drug Discovery, a section of the journal Frontiers in Pharmacology</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>30</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>771808</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Gong, Liao, Wang and Zou.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Gong, Liao, Wang and Zou</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Drug targets are biological macromolecules or biomolecule structures capable of specifically binding a therapeutic effect with a particular drug or regulating physiological functions. Due to the important value and role of drug targets in recent years, the prediction of potential drug targets has become a research hotspot. The key to the research and development of modern new drugs is first to identify potential drug targets. In this paper, a new predictor, DrugHybrid_BS, is developed based on hybrid features and Bagging-SVM to identify potentially druggable proteins. This method combines the three features of monoDiKGap (k &#x3d; 2), cross-covariance, and grouped amino acid composition. It removes redundant features and analyses key features through MRMD and MRMD2.0. The cross-validation results show that 96.9944% of the potentially druggable proteins can be accurately identified, and the accuracy of the independent test set has reached 96.5665%. This all means that DrugHybrid_BS has the potential to become a useful predictive tool for druggable proteins. In addition, the hybrid key features can identify 80.0343% of the potentially druggable proteins combined with Bagging-SVM, which indicates the significance of this part of the features for research.</p>
</abstract>
<kwd-group>
<kwd>monoDiKGap</kwd>
<kwd>CC</kwd>
<kwd>GAAC</kwd>
<kwd>bagging</kwd>
<kwd>support vector machine</kwd>
</kwd-group>
<contract-num rid="cn001">61863010 11926205&#x20;11926412 61873076</contract-num>
<contract-num rid="cn002">119MS036 120RC588</contract-num>
<contract-num rid="cn003">2020YFB2104400</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">Natural Science Foundation of Hainan Province<named-content content-type="fundref-id">10.13039/501100004761</named-content>
</contract-sponsor>
<contract-sponsor id="cn003">National Key Research and Development Program of China<named-content content-type="fundref-id">10.13039/501100012166</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Drug targets refer to the binding sites of drugs in the body. To date, there are approximately 130 protein families as therapeutic drug targets, which usually include enzymes (<xref ref-type="bibr" rid="B29">Liu et&#x20;al., 2019a</xref>; <xref ref-type="bibr" rid="B33">Meng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B61">Xu et&#x20;al., 2021a</xref>; <xref ref-type="bibr" rid="B52">Wang et&#x20;al., 2021</xref>), G protein-coupled receptors (<xref ref-type="bibr" rid="B40">Ru et&#x20;al., 2020</xref>), ion channels and transporters (<xref ref-type="bibr" rid="B14">Han et&#x20;al., 2019</xref>), nuclear hormone receptors, etc (<xref ref-type="bibr" rid="B25">Li and Lai, 2007</xref>). These drug targets are of great significance for disease treatment and drug research and development (<xref ref-type="bibr" rid="B5">Ding et&#x20;al., 2019a</xref>; <xref ref-type="bibr" rid="B6">Ding et&#x20;al., 2019b</xref>; <xref ref-type="bibr" rid="B45">Shi et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B4">Ding et&#x20;al., 2020a</xref>; <xref ref-type="bibr" rid="B48">Wang et&#x20;al., 2020a</xref>; <xref ref-type="bibr" rid="B7">Ding et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B44">Shang et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B74">Zhuang et&#x20;al., 2021</xref>). However, the discovery and development of modern drugs is usually a time-consuming and laborious process. It is estimated that it takes an average of 10&#x2013;15&#xa0;years to bring a drug to the market, which costs approximately US $2,558 million (<xref ref-type="bibr" rid="B71">Zhong et&#x20;al., 2018</xref>). Therefore, predicting whether a protein can potentially be used as a drug target has significant value in disease treatment and reducing the time and cost of drug development, which greatly accelerates the drug development process for the protein (<xref ref-type="bibr" rid="B49">Wang et&#x20;al., 2020b</xref>; <xref ref-type="bibr" rid="B64">Yu et&#x20;al., 2021</xref>).</p>
<p>The discovery of drug targets has attracted extensive attention in both academia and the pharmaceutical industry. The commonly used methods for drug target prediction can be roughly divided into three types. The first type is to analyse known drug targets at the genome level based on sequence homology and to find potential drug targets from protein families (<xref ref-type="bibr" rid="B16">Hopkins and Groom, 2002</xref>; <xref ref-type="bibr" rid="B41">Russ and Lampel, 2005</xref>; <xref ref-type="bibr" rid="B34">Munir et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B1">Ao et&#x20;al., 2021</xref>). Not all members of the same protein family can be used as therapeutic drug targets. The second type predicts whether the new target is druggable based on several chemical properties, molecular drug similarity, and target properties (<xref ref-type="bibr" rid="B11">Gayvert et&#x20;al., 2016</xref>). This method is usually limited by experimental cost. The third type is discovering drug targets based on protein structure, which predicts the protein&#x2019;s drug properties by searching for the binding site and binding affinity of the target protein (<xref ref-type="bibr" rid="B42">Salmaso and Moro, 2018</xref>). However, this method has limitations because the three-dimensional structure of most proteins is not easy to obtain.</p>
<p>With the advent of the genome era, revolutionary changes have taken place in the field of drug research and development. Many computing methods were used for effective drug target prediction. To better find potential drug targets and provide new options for drug redirection, Cheng et&#x20;al. (<xref ref-type="bibr" rid="B2">Cheng et&#x20;al., 2021</xref>) established the GraphMS model. They fused heterogeneous graph information using mutual information in the heterogeneous graph to obtain effective node information and substructure information. The experimental results show that the area under the receiver operating characteristic curve (AUROC) was 0.959, and the area under the precision-recall curve (AUPR) was 0.847. Dezs&#x151; et&#x20;al. (<xref ref-type="bibr" rid="B3">Dezs&#x151; and Ceccarelli, 2020</xref>) developed a machine learning model for tumour drug targets. A variety of protein features, including features from sequences, features that characterize protein functions, and network features from protein-protein interaction networks, were included in the model. It has achieved high accuracy on the drug target of independent clinical trial drug targets, with an area under the curve of 0.89. In order to establish a high-quality environment-specific metabolic model that can be used for drug target prediction, Pacheco et&#x20;al. (<xref ref-type="bibr" rid="B37">Pacheco et&#x20;al., 2019</xref>) developed a metabolic model FASTCORMICS RNA-seq workflow (rFASTCORMICS) based on RNA-seq data. The genes and response characteristics of 13 different types of cancer were extracted. At the same time, 17 new colon cancer candidate drugs were predicted, of which 3 drugs were verified <italic>in&#x20;vitro</italic> in colon cancer cell lines. Ji et&#x20;al. (<xref ref-type="bibr" rid="B20">Ji et&#x20;al., 2019</xref>) proposed a DTINet method based on network propagation, starting from the diffusion component analysis of potential drug targets and disease networks. The DTINet performed well under the receiver operating characteristic curve (AUROC &#x3d; 0.86&#x20;&#xb1; 0.008). To achieve the rapid identification of novel targets, Li et&#x20;al. (<xref ref-type="bibr" rid="B25">Li and Lai, 2007</xref>) constructed a simple model extraction characteristics from known drug target protein sequences. Using this model, drug targets and nondrug targets can be distinguished with 84% accuracy. Jamali et&#x20;al. (<xref ref-type="bibr" rid="B19">Jamali et&#x20;al., 2016</xref>) based on the protein features derived from 443 sequences, the accuracy of predicting drug targets through neural network models reached 89.98%.</p>
<p>This paper selected three feature extraction methods: monoDiKGap (k &#x3d; 2), cross covariance (CC) and grouped amino acid composition (GAAC) (<xref ref-type="bibr" rid="B77">Zuo et&#x20;al., 2017</xref>). The three individual features were mixed in different combinations through the hybrid feature method. The MRMD was used to remove redundant hybrid features, and the integrated method bagging was used to improve the classification performance of potentially druggable proteins. We performed the importance analysis on the best feature combination and selected the key features that distinguish potentially druggable proteins. The results show that the hybrid features of the three feature extraction methods can predict the potentially druggable proteins well by the integrated method bagging, and can correctly predict 96.9944% of the druggable target proteins. This model was conducive to better promotion of drug development. Furthermore, the potential drug targets screened out can provide references for new drug targets.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<p>This paper mainly studied the following parts, and the step flow chart was shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>:<list list-type="simple">
<list-item>
<p>1. Establishment of dataset.</p>
</list-item>
<list-item>
<p>2. Use three single feature extraction methods, monoDiKGap, Cross Covariance, and Grouped Amino Acid Composition, to represent the features of dataset.</p>
</list-item>
<list-item>
<p>3. Combine three single feature methods to obtain hybrid features.</p>
</list-item>
<list-item>
<p>4. The MRMD was used to remove redundant features, and the MRMD2.0 obtained key features.</p>
</list-item>
<list-item>
<p>5. The feature subset predicted the potentially druggable proteins through the optimized Bagging-SVM&#x20;model.</p>
</list-item>
</list>
</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Flow chart of DrugHybrid_BS model <bold>(A)</bold> Process the referenced dataset <bold>(B)</bold> Three single feature representation methods were used to extract features <bold>(C)</bold> Combine three single feature representation methods and select the best hybrid feature <bold>(D)</bold> Use MRMD to remove redundant features and MRMD2.0 to obtain key features <bold>(E)</bold> Feature subsets were used to predict potentially druggable proteins through the optimized Bagging-SVM model <bold>(F)</bold> Evaluate model prediction effects based on performance indicators.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g001.tif"/>
</fig>
<p>This research was carried out under the software python 3.7.4. By comparing the new method DrugHybrid_BS with other machine learning models, the study found that the classification effect of DrugHybrid_BS was better, which was helpful for the prediction of potentially druggable proteins.</p>
<sec id="s2-1">
<title>2.1 Dataset Construction</title>
<p>This paper cited the dataset proposed by Lin et&#x20;al. (<xref ref-type="bibr" rid="B28">Lin et&#x20;al., 2019</xref>), in which the drug target dataset was downloaded from the DrugBank (<xref ref-type="bibr" rid="B59">Wishart et&#x20;al., 2006</xref>) database. In the original dataset, 1,224 druggable protein sequences were selected as the positive sample set, and 1,319&#x20;non-druggable proteins were selected as the negative sample set. We further processed the dataset by removing the protein sequences containing non-standard amino acid characters &#x201c;B", &#x201c;J", &#x201c;O", &#x201c;U", &#x201c;X" and &#x201c;Z". For the remaining sequences, the CD-Hit program (<xref ref-type="bibr" rid="B10">Fu et&#x20;al., 2012</xref>) was used to set a critical value of 60% sequence identity to delete highly similar sequences to avoid overfitting caused by homologous deviation and noise in training (<xref ref-type="bibr" rid="B76">Zou et&#x20;al., 2020</xref>).</p>
<p>The processed dataset was represented by D, which is the combination of <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2b;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2212;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>:<disp-formula id="e1">
<mml:math id="m3">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2b;</mml:mo>
</mml:msup>
<mml:mo>&#x222a;</mml:mo>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2212;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2b;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents potentially druggable protein samples and <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mo>&#x2212;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents non-druggable protein samples. The positive sample set contained 1,050 protein sequences, and the negative sample set concluded contained 1,279 protein sequences. <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> showed the sample distribution of the dataset.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Sample distribution of dataset.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g002.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>2.2 Feature Representation</title>
<sec id="s2-2-1">
<title>2.2.1 monoDiKGap</title>
<p>The monoDiKGap feature is a variant of the kmer feature extraction method in the PyFeat package. Kmer, as our common feature extraction method, is also called k-tuples (<xref ref-type="bibr" rid="B30">Liu et&#x20;al., 2019b</xref>; <xref ref-type="bibr" rid="B32">Lv et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B35">Niu et&#x20;al., 2021a</xref>). MonoDiKGap refers to the combination of subsequences with KGap used to describe the sequence. While monoDiKGap generates all feature sets, it can also use the AdaBoost (<xref ref-type="bibr" rid="B72">Zhu et&#x20;al., 2006</xref>) classification model to reduce redundant features to generate the optimal feature set. The generated optimal feature set will not only reduce the feature dimension but also ensure a good prediction. In this study, we set KGap to 2. At this time, the monoDiKGap feature can be expressed as:<disp-formula id="e2">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>V</mml:mi>
<mml:mrow>
<mml:mi mathvariant="italic">KGap</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mn>8000</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>1</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mn>8000</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf5">
<mml:math id="m7">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>8000</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the frequency of the <italic>i</italic>th feature calculated when the feature was shaped like <inline-formula id="inf6">
<mml:math id="m8">
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>X</mml:mi>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, and the generated feature at this time was like <inline-formula id="inf7">
<mml:math id="m9">
<mml:mrow>
<mml:mo>"</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>A</mml:mi>
<mml:mi>A</mml:mi>
<mml:mo>"</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. <inline-formula id="inf8">
<mml:math id="m10">
<mml:mrow>
<mml:msubsup>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>8000</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the frequency of the <italic>i</italic>th feature calculated when the feature was shaped like <inline-formula id="inf9">
<mml:math id="m11">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mo>&#x2014;</mml:mo>
<mml:mo>&#x2014;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>X</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, the generated feature was like <inline-formula id="inf10">
<mml:math id="m12">
<mml:mrow>
<mml:mo>"</mml:mo>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mo>&#x2014;</mml:mo>
<mml:mo>&#x2014;</mml:mo>
</mml:mrow>
</mml:msub>
<mml:mi>A</mml:mi>
<mml:mi>A</mml:mi>
<mml:mo>"</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and X represents twenty natural amino acids. Therefore, the total feature set generated by this feature extraction method has a total of 16,000 features, which AdaBoost automatically optimizes to generate 466 feature subsets with more discriminative capabilities.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Cross Covariance (CC)</title>
<p>CC is the correlation between two different attributes separated by lag (<xref ref-type="bibr" rid="B13">Guo et&#x20;al., 2008</xref>). For this study, the CC variable described the average interaction between two fragments with different physical and chemical properties separated by lag fragments. Suppose that the protein sequence P has L residues, <inline-formula id="inf11">
<mml:math id="m13">
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mn>...</mml:mn>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>. where <inline-formula id="inf12">
<mml:math id="m14">
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>E</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>F</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>G</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>H</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>I</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>K</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>M</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>P</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>Q</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>S</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>T</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>W</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the amino acid at position <inline-formula id="inf13">
<mml:math id="m15">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> in the sequence. Then, for each protein sequence, there is a physical and chemical information matrix of the following <inline-formula id="inf14">
<mml:math id="m16">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> size, which can be expressed as:<disp-formula id="e3">
<mml:math id="m17">
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>11</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>12</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>13</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>21</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>...</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
<mml:mtd>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>22</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>...</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
<mml:mtd>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>23</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mn>...</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf15">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>L</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> stands for the hydrophobicity values, hydrophilicity values and side chain mass of amino acid <inline-formula id="inf16">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, respectively.</p>
<p>CC converts protein sequences of different lengths into feature vectors of the same length. The calculation formula of the CC feature representation method is as follows:<disp-formula id="e4">
<mml:math id="m20">
<mml:mrow>
<mml:mi mathvariant="italic">CC</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>g</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#xaf;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>
<inline-formula id="inf17">
<mml:math id="m21">
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>g</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, here <inline-formula id="inf18">
<mml:math id="m22">
<mml:mrow>
<mml:mi>lg</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> was the default. Because CC was an asymmetric vector, under this physical and chemical characteristic condition, the feature dimension of the CC vector was twelve.</p>
</sec>
<sec id="s2-2-3">
<title>2.2.3 Grouped Amino Acid Composition (GAAC)</title>
<p>In the GAAC code, twenty amino acid types are divided into five categories based on their physical and chemical properties (<xref ref-type="bibr" rid="B24">Lee et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B69">Zheng et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B70">Zheng et&#x20;al., 2021</xref>). These five categories include the aliphatic group <inline-formula id="inf19">
<mml:math id="m23">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="normal">GAVLMI</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, aromatic group <inline-formula id="inf20">
<mml:math id="m24">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="normal">FYW</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, positive charge group <inline-formula id="inf21">
<mml:math id="m25">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mn>3</mml:mn>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="normal">KRH</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, negative charged group <inline-formula id="inf22">
<mml:math id="m26">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="normal">DE</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, and uncharged group <inline-formula id="inf23">
<mml:math id="m27">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
<mml:mo>:</mml:mo>
<mml:mi mathvariant="normal">STCPNQ</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The GAAC descriptor refers to the frequency of each amino acid group, which is calculated as follows:<disp-formula id="e5">
<mml:math id="m28">
<mml:mrow>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m29">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf24">
<mml:math id="m30">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the number of amino acids in group <inline-formula id="inf25">
<mml:math id="m31">
<mml:mi>g</mml:mi>
</mml:math>
</inline-formula>, <inline-formula id="inf26">
<mml:math id="m32">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the number of amino acid types <inline-formula id="inf27">
<mml:math id="m33">
<mml:mi>t</mml:mi>
</mml:math>
</inline-formula>, and N is the length of the protein sequence.</p>
<p>As an example, for the sequence <inline-formula id="inf28">
<mml:math id="m34">
<mml:mrow>
<mml:mo>"</mml:mo>
<mml:mi mathvariant="normal">EAHGAFLMDKPSMFNERV</mml:mi>
<mml:mo>"</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, the amount of occurrences of character &#x201c;E" was 2, the amount of occurrences of character &#x201c;A" was 2, the amount of occurrences of character &#x201c;H" was 1, the amount of occurrences of character &#x201c;G" was 1, etc. The length of the sequence was 18, <inline-formula id="inf29">
<mml:math id="m35">
<mml:mrow>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>7</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>4</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mn>5</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>. Therefore, the GAAC feature of this sequence was expressed as <inline-formula id="inf30">
<mml:math id="m36">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mn>7</mml:mn>
<mml:mrow>
<mml:mn>18</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>9</mml:mn>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mn>3</mml:mn>
<mml:mrow>
<mml:mn>18</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mn>3</mml:mn>
<mml:mrow>
<mml:mn>18</mml:mn>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mn>3</mml:mn>
<mml:mrow>
<mml:mn>18</mml:mn>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
</sec>
</sec>
<sec id="s2-3">
<title>2.3 Machine Learning Algorithm</title>
<p>In this study, predicting druggable proteins was a typical binary classification problem. To better explore prediction models and analysis features, we mainly used four machine learning algorithms for prediction tasks, namely, support vector machine, K-nearest neighbour, bagging integrated learning, and random forest.</p>
<sec id="s2-3-1">
<title>2.3.1&#x20;K-Nearest Neighbour (KNN)</title>
<p>The k-nearest neighbour algorithm is a classic machine learning algorithm (<xref ref-type="bibr" rid="B27">Liao and Vemuri, 2002</xref>; <xref ref-type="bibr" rid="B43">Samanthula et&#x20;al., 2014</xref>). The principle of the k-nearest neighbour algorithm is straightforward: a sample in the feature space will always find the k data closest to it, that is, the nearest sample in the feature space. If most of the k data belong to a specific category, the sample also belongs to this category. In this study, the default parameters of the prediction model were selected, and the value of k was&#x20;3.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Support Vector Machine (SVM)</title>
<p>Although the support vector machine has only a short development history of more than 20&#x20;years. It shows strong energy in classification problems (<xref ref-type="bibr" rid="B8">Ding et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B55">Wei et&#x20;al., 2018a</xref>; <xref ref-type="bibr" rid="B51">Wang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B50">Wang et&#x20;al., 2020c</xref>; <xref ref-type="bibr" rid="B18">Huo et&#x20;al., 2020</xref>). It has become the mainstream technology of machine learning from the end of the 20th century to the beginning of the 21st century, applied to many fields (<xref ref-type="bibr" rid="B21">Jiang et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B62">Xu et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B67">Zhang et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B53">Wei et&#x20;al., 2019a</xref>; <xref ref-type="bibr" rid="B31">Liu et&#x20;al., 2021</xref>). The support vector machine uses the maximum classification interval to determine the optimal partitioning hyperplane to obtain good generalization. For the binary classification problem in this study, when we obtain a feature dataset containing category information:<disp-formula id="e7">
<mml:math id="m37">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mn>...</mml:mn>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where n was the number of samples, the feature dimension of each sample was d, and the samples were divided into positive categories ( <inline-formula id="inf31">
<mml:math id="m38">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> represents druggable protein) and negative categories ( <inline-formula id="inf32">
<mml:math id="m39">
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> represents non-druggable protein). Our goal was to find the optimal hyperplane to maximize the sample interval between the positive class and the negative&#x20;class.</p>
<p>We used <inline-formula id="inf33">
<mml:math id="m40">
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:mi>x</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> to represent the partitioning hyperplane and used the geometric margin to find the optimal partitioning hyperplane. The geometric interval was numerically equal to the distance from the sample point to the partition hyperplane. The distance from the positive sample point <inline-formula id="inf34">
<mml:math id="m41">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to the partition hyperplane was <inline-formula id="inf35">
<mml:math id="m42">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, and the distance from the negative sample <inline-formula id="inf36">
<mml:math id="m43">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to the partition hyperplane was <inline-formula id="inf37">
<mml:math id="m44">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula id="inf38">
<mml:math id="m45">
<mml:mi>&#x3c9;</mml:mi>
</mml:math>
</inline-formula> was the normal vector of the partition hyperplane and <inline-formula id="inf39">
<mml:math id="m46">
<mml:mi>b</mml:mi>
</mml:math>
</inline-formula> was the intercept. Therefore, the distance from any sample <inline-formula id="inf40">
<mml:math id="m47">
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> to the partition hyperplane can be uniformly expressed as <inline-formula id="inf41">
<mml:math id="m48">
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msup>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>T</mml:mi>
</mml:msup>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>b</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x2016;</mml:mo>
<mml:mi>&#x3c9;</mml:mi>
<mml:mo>&#x2016;</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula>. To solve the optimization problem of linear separable support vector machines. John C. Platt proposed the sequential minimal optimization algorithm (<xref ref-type="bibr" rid="B38">Platt, 1998</xref>) in 1998. The algorithm decomposed the large convex quadratic programming (QP) problem to be solved in the training process of support vector machines into a series of minimum possible QP problems, avoided time-consuming internal iterative optimization, and improves computational efficiency.</p>
<p>In addition, the kernel function is a unique feature of the support vector model. For the same dataset, different kernel function choices will have different prediction effects. Appropriate kernel functions can improve prediction performance. The commonly used functions include the linear, Gaussian, and polynomial kernel functions. In this study, a linear kernel function was selected as the kernel of the support vector machine by comparing different kernel functions.</p>
</sec>
<sec id="s2-3-3">
<title>2.3.3 Bagging</title>
<p>Bagging is one of the common ensemble learning models (<xref ref-type="bibr" rid="B9">Dudoit and Fridlyand, 2003</xref>; <xref ref-type="bibr" rid="B23">Jin et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B22">Jin et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B60">Wu and Yu, 2021</xref>). The ensemble learning model uses a series of weak learners (also called basic models) for learning and integrates the results of each weak learner to obtain a better learning effect than individual learners.</p>
<p>The bagging algorithm uses the simplest combination strategy to obtain the integration model. For the classification problem, the majority voting method is adopted. Each weak learner has one vote, and the final prediction result is generated according to the votes of all weak learners. The process of the bagging method is as follows: suppose we have a training set containing N samples and randomly put back the data to form a new training set. Because there is a way to put back sampling, a sample may be selected multiple times, or a sample may not be selected once. Hence, the size of the sampled data samples is the same as that of the original training data samples, but they contain different data. In this way, after T groups of data are extracted, T weak learners trained by different training sets can be obtained at the end of training. According to the prediction results of T weak learners, the most voting method is adopted to obtain a more accurate and reasonable prediction&#x20;model.</p>
</sec>
<sec id="s2-3-4">
<title>2.3.4 Random Forest(RF)</title>
<p>Random forest is a representative bagging algorithm based on decision trees. Because random forest has good performance in regression and classification prediction, it has attracted great attention. It has been widely used in many practical problems, such as genome data analysis and disease risk prediction. When making classification prediction, each decision tree will make classification judgment on the data according to the characteristics of the data. Through the majority voting method, the category with the most votes is the prediction result of the random forest.</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4 Feature Selection</title>
<p>In the feature extraction section, we introduced three feature representation methods. The optimal feature subset of the dataset sample generated by the monoDiKGap method had 466 features. The CC feature representation method generated 12 features, and the GAAC feature method generated five features. Different feature extraction methods were combined to obtain hybrid features. However, the hybrid of features may lead to feature redundancy and affect the predictive effect of potentially druggable proteins. Therefore, we used MRMD and MRMD2.0 to select features and used fewer features to distinguish between potentially druggable and non-druggable proteins better.</p>
<p>In this study, the MRMD (<xref ref-type="bibr" rid="B39">Quan et&#x20;al., 2016</xref>) was used to remove redundant features in hybrid features. The MRMD will leave the optimal feature subset after automatic feature selection. The main principle of this method is to use the Euclidean distance, cosine distance, and the Tanimoto coefficient to calculate the redundancy between features and use the Pearson correlation coefficient to calculate the correlation between dataset features and class labels to generate feature subsets with low redundancy and strong correlation automatically. When we analyse the hybrid feature subset that can accurately predict potentially druggable proteins, we also need to analyse the importance of different features. MRMD2.0 (<xref ref-type="bibr" rid="B15">He et&#x20;al., 2021</xref>) combined seven algorithms, such as ANOVA, MIC, LASSO, mRMR, and chi-square test, through the PageRank strategy algorithm to rank different algorithm lists to form a directed graph, and each feature obtained a score. According to the ranking information, we analyse the importance of features and obtain key features that influence the prediction of potentially druggable proteins.</p>
</sec>
<sec id="s2-5">
<title>2.5 Performance Evaluation</title>
<p>To intuitively measure the quality of the model, we evaluated the predictive effect of the model. This study used common evaluation indicators, including TP rate (TPR), FP rate (FPR), precision (<xref ref-type="bibr" rid="B47">Su et&#x20;al., 2018</xref>), F-score (<xref ref-type="bibr" rid="B46">Sokolova et&#x20;al., 2006</xref>), and accuracy (ACC) (<xref ref-type="bibr" rid="B57">Wei et&#x20;al., 2017a</xref>; <xref ref-type="bibr" rid="B58">Wei et&#x20;al., 2017b</xref>; <xref ref-type="bibr" rid="B56">Wei et&#x20;al., 2018b</xref>; <xref ref-type="bibr" rid="B54">Wei et&#x20;al., 2019b</xref>; <xref ref-type="bibr" rid="B17">Huang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B26">Liang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B66">Zhang et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B63">Xu et&#x20;al., 2021b</xref>; <xref ref-type="bibr" rid="B73">Zhu et&#x20;al., 2021</xref>). The calculation method of each measurement index was as follows:<disp-formula id="e9">
<mml:math id="m49">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable columnalign="left">
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">TPR</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">FPR</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">FP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">FP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">TN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">precision</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FP</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">recall</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi mathvariant="italic">score</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#x2217;</mml:mo>
<mml:mi mathvariant="italic">precision</mml:mi>
<mml:mo>&#x2217;</mml:mo>
<mml:mi mathvariant="italic">recall</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">precision</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">recall</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mi mathvariant="italic">ACC</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">TN</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">TN</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>Here, TP represents the classification number of correct positive samples, and TN represents the classification number of correct negative samples. FP represents the classification number of false positive samples. FN represents the classification number of false negative samples. In addition, this study also used 5-fold cross-validation to predict and evaluate the&#x20;model.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and Discussion</title>
<sec id="s3-1">
<title>3.1 Performance of Single Feature Extraction Methods</title>
<p>Because the monoDiKGap feature extraction method gradually increases with the value of KGap, the number of corresponding generated feature vectors increases exponentially. In this study, the total feature set generated by monoDiKGap(k &#x3d; 2) has 16,000 features, but in fact many small fragments appear very rarely, and some even appear 0 or 1 times. At this time, a large number of feature vectors composed of 0 or one also have no meaning already. In order to avoid high-dimensional feature vectors introducing dimensional disasters for subsequent machine learning algorithms, resulting in a significant decline in predictive classification performance. Therefore, this study used AdaBoost to automatically generate a more discriminative 466-dimensional feature subset, and compared the ACC values of the full feature set and the feature subset of monoDiKGap(k &#x3d; 2) under different classifiers, as shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Comparison the ACC values of the full feature set and 466-dimensional feature subset extracted by monoDiKGap(k &#x3d; 2) under different classifiers.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g003.tif"/>
</fig>
<p>In this paper, three single feature extraction methods, monoDiKGap(k &#x3d; 2), CC and GAAC, were used to represent the features of the dataset. Three single feature representation methods extracted 466-dimensional, 12-dimensional, and 5-dimensional features. The prediction performance of each extraction method under SVM, KNN, and RF was shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. The data in <xref ref-type="table" rid="T1">Table&#x20;1</xref> showed that the accuracy of the monoDiKGap(k &#x3d; 2) feature representation method in predicting potentially druggable proteins through the SVM classification algorithm was higher than that of KNN and RF. The model can accurately predict 96.608% of the potentially druggable proteins. At this time, the TPR value reached 0.965, the FPR value reached 0.033, the F-score reached 0.962, and the ROC curve area was 0.966. The GAAC feature extraction method had an accuracy of 77.2864% in predicting potentially druggable proteins under SVM, which was 1.20 and 2.40% higher than that of the RF and KNN classification models, respectively. The accuracy of the CC feature extraction method to predict proteins through the SVM feature representation method was only 1.07% lower than that of the KNN algorithm. Therefore, considering the performance evaluation of the three feature representation methods under different classifiers, the SVM classification algorithm was more suitable for accurately predicting potentially druggable proteins.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Compare the results of different feature methods under different classifiers.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">Classifier</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="left">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">monoDiKGap (k &#x3d; 2)</td>
<td align="center">SVM</td>
<td align="char" char=".">96.608</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.033</td>
<td align="char" char=".">0.960</td>
<td align="char" char=".">0.962</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="center">KNN</td>
<td align="char" char=".">58.437</td>
<td align="char" char=".">0.083</td>
<td align="char" char=".">0.004</td>
<td align="char" char=".">0.946</td>
<td align="char" char=".">0.152</td>
<td align="char" char=".">0.628</td>
</tr>
<tr>
<td align="center">RF</td>
<td align="char" char=".">85.272</td>
<td align="char" char=".">0.788</td>
<td align="char" char=".">0.094</td>
<td align="char" char=".">0.873</td>
<td align="char" char=".">0.828</td>
<td align="char" char=".">0.928</td>
</tr>
<tr>
<td rowspan="3" align="left">CC</td>
<td align="center">SVM</td>
<td align="char" char=".">57.364</td>
<td align="char" char=".">0.243</td>
<td align="char" char=".">0.155</td>
<td align="char" char=".">0.563</td>
<td align="char" char=".">0.339</td>
<td align="char" char=".">0.544</td>
</tr>
<tr>
<td align="center">KNN</td>
<td align="char" char=".">58.437</td>
<td align="char" char=".">0.625</td>
<td align="char" char=".">0.449</td>
<td align="char" char=".">0.533</td>
<td align="char" char=".">0.575</td>
<td align="char" char=".">0.599</td>
</tr>
<tr>
<td align="center">RF</td>
<td align="char" char=".">63.718</td>
<td align="char" char=".">0.569</td>
<td align="char" char=".">0.306</td>
<td align="char" char=".">0.604</td>
<td align="char" char=".">0.586</td>
<td align="char" char=".">0.679</td>
</tr>
<tr>
<td rowspan="3" align="left">GAAC</td>
<td align="center">SVM</td>
<td align="char" char=".">77.286</td>
<td align="char" char=".">0.768</td>
<td align="char" char=".">0.223</td>
<td align="char" char=".">0.739</td>
<td align="char" char=".">0.753</td>
<td align="char" char=".">0.772</td>
</tr>
<tr>
<td align="center">KNN</td>
<td align="char" char=".">74.882</td>
<td align="char" char=".">0.745</td>
<td align="char" char=".">0.248</td>
<td align="char" char=".">0.712</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.807</td>
</tr>
<tr>
<td align="center">RF</td>
<td align="char" char=".">76.084</td>
<td align="char" char=".">0.729</td>
<td align="char" char=".">0.213</td>
<td align="char" char=".">0.738</td>
<td align="char" char=".">0.733</td>
<td align="char" char=".">0.850</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>3.2 Performance of Hybrid Feature Representation Methods</title>
<p>To explore the prediction performance of hybrid features, we combined the above three feature representation methods and obtained new feature vectors of different combinations. After the combination of three single feature extraction methods, four new feature vectors were obtained: monoDiKGap &#x2b; CC, monoDiKGap &#x2b; GAAC, CC &#x2b; GAAC and monoDiKGap &#x2b; CC &#x2b; GAAC. <xref ref-type="table" rid="T2">Table&#x20;2</xref> showed the evaluation performance of different combinations of hybrid features using the SVM classification algorithm. <xref ref-type="table" rid="T2">Table&#x20;2</xref> indicated that compared with the single feature representation method, the hybrid feature showed higher performance. The accuracy of the combination of monoDiKGap and other feature representation methods was more than 96%. In addition, the prediction performance of the CC &#x2b; GAAC feature combination was also higher than that of the single feature representation method. Importantly, we found that the combination of monoDiKGap, CC, and GAAC features showed the best prediction performance, and the hybrid feature could accurately predict 96.6509% of potentially druggable proteins.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Performance comparison of different feature combinations under SVM classifiers.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="center">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">monoDiKGap &#x2b; CC</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.967</td>
<td align="char" char=".">0.034</td>
<td align="char" char=".">0.959</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.967</td>
</tr>
<tr>
<td align="left">monoDiKGap &#x2b; GAAC</td>
<td align="char" char=".">96.350</td>
<td align="char" char=".">0.958</td>
<td align="char" char=".">0.032</td>
<td align="char" char=".">0.961</td>
<td align="char" char=".">0.959</td>
<td align="char" char=".">0.963</td>
</tr>
<tr>
<td align="left">CC &#x2b; GAAC</td>
<td align="char" char=".">78.360</td>
<td align="char" char=".">0.770</td>
<td align="char" char=".">0.206</td>
<td align="char" char=".">0.755</td>
<td align="char" char=".">0.801</td>
<td align="char" char=".">0.782</td>
</tr>
<tr>
<td align="left">monoDiKGap &#x2b; CC &#x2b; GAAC</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.961</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 Kernel and Parameters of Support Vector Machine</title>
<p>The kernel function is an important feature of support vector machines. The kernel function choice of the support vector machine affects the prediction performance of the model. For the monoDiKGap, CC, and GAAC hybrid features to represent the dataset features, we used different kernel functions and 5-fold cross-validation to select the appropriate kernel function. We compared the performance of the linear kernel function, quadratic polynomial kernel function, and radial basis kernel function. The ROC curves of different kernel functions were shown in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref>. The ROC values were 0.966, 0.955, and 0.846. The evaluation indicators of the three kernel functions were shown in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. We can see that the prediction effect of the hybrid feature using the linear kernel function was better than the quadratic kernel function and the radial basis function. At this time, the three kernel functions predicted 96.6509, 95.5346, and 85.745% of the potentially druggable proteins, respectively. Therefore, this paper chose a linear kernel function as the kernel of the support vector machine.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>ROC curves of support vector machines in different kernel functions.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g004.tif"/>
</fig>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Performance comparison of hybrid features under different kernel functions.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Kernel function</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="center">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">liner kernel</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.961</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">polynomial kernel</td>
<td align="char" char=".">95.535</td>
<td align="char" char=".">0.953</td>
<td align="char" char=".">0.043</td>
<td align="char" char=".">0.948</td>
<td align="char" char=".">0.951</td>
<td align="char" char=".">0.955</td>
</tr>
<tr>
<td align="left">RBF</td>
<td align="char" char=".">85.745</td>
<td align="char" char=".">0.730</td>
<td align="char" char=".">0.038</td>
<td align="char" char=".">0.940</td>
<td align="char" char=".">0.822</td>
<td align="char" char=".">0.846</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For the linear kernel of the support vector machine, the penalty parameter C is an important parameter. The larger the value of C is, the easier it is to overfit, while the smaller the value of C is, the easier it is to underfit. The most commonly used C values are 1, 10, 100, and 1,000. We selected the appropriate C value with the help of grid search. <xref ref-type="table" rid="T4">Table&#x20;4</xref> showed the prediction performance of different penalty parameters. When the C value was 1, the support vector machine classification algorithm achieved a better prediction effect and shortened the running&#x20;time.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Performance comparison of hybrid features with different penalty parameter C values under linear kernel.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">C Values</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="center">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.961</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">10</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.960</td>
<td align="char" char=".">0.030</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">100</td>
<td align="char" char=".">96.608</td>
<td align="char" char=".">0.960</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.962</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">1,000</td>
<td align="char" char=".">96.608</td>
<td align="char" char=".">0.960</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.962</td>
<td align="char" char=".">0.966</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-4">
<title>3.4 Hybrid Feature Selection</title>
<p>The best hybrid features are 483-dimensional features mixed by the monoDiKGap, CC, and GAAC feature representation methods. These features may contain redundancy and affect the performance. Since the monoDiKGap feature extraction method automatically generated the optimal feature subset, we also need to remove redundant features from the CC and GAAC feature extraction methods. We used MRMD to filter the feature sets extracted by CC and GAAC and generated the optimal feature subset with low redundancy and strong correlation. Finally, we combined the feature subsets to obtain the filtered new hybrid features. These hybrid features not only reduced the feature dimensions but also had more expressiveness (<xref ref-type="table" rid="T5">Table&#x20;5</xref>).</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Comparison of classification performance of hybrid features before and after using MRMD feature selection.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Number of feature</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="center">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">483</td>
<td align="char" char=".">96.651</td>
<td align="char" char=".">0.961</td>
<td align="char" char=".">0.029</td>
<td align="char" char=".">0.965</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">472</td>
<td align="char" char=".">96.694</td>
<td align="char" char=".">0.959</td>
<td align="char" char=".">0.027</td>
<td align="char" char=".">0.967</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-5">
<title>3.5 Bagging Algorithm and Comparison With Other Algorithms</title>
<p>The expressive ability of a single support vector machine classification model may be limited so that the bagging ensemble algorithm based on a support vector machine has room for improvement. Compared with a single model, the bagging integration method can enhance the expressive ability of the model and reduce the error. When it is difficult for a single model to correctly distinguish the two types of data, the ensemble algorithm can often improve the model&#x2019;s prediction performance by constructing multiple independent base models.</p>
<p>In this study, a support vector machine with a penalty coefficient of one and a linear kernel function was used as the basic model, and the number of optimal basic models was selected to construct a Bagging-SVM classification algorithm. The hybrid features of monoDiKGap, CC, and GAAC, which removed the cumbersome features, were shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref> under the Bagging-SVM classification algorithm where the number of base models was 1&#x2013;20. The accuracy of combining hybrid features and Bagging-SVM to predict potentially druggable proteins was basically more than 96.73%, and the highest prediction accuracy was 96.9944% when the number of base models was&#x20;12.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The accuracy of hybrid features in predicting potential druggable proteins under the Bagging-SVM classification algorithm where the number of base models was 1&#x2013;20.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g005.tif"/>
</fig>
<p>Based on the hybrid features of monoDiKGap, CC, GAAC, and Bagging-SVM, a new predictive model, DrugHybrid_BS, was constructed. To further explore the prediction model, we evaluated the performance of SVM, RF, and KNN using the same hybrid feature set. <xref ref-type="table" rid="T6">Table&#x20;6</xref> showed that the DrugHybrid_BS model can better predict potentially druggable proteins. At this time, the TPR value reached 0.970, the F-score reached 0.967, and the AUC value reached 0.992. In addition, <xref ref-type="table" rid="T6">Table&#x20;6</xref> showed the prediction performance comparison between the DrugHybrid_BS model and the previous model when using the same dataset as Lin et&#x20;al. (<xref ref-type="bibr" rid="B28">Lin et&#x20;al., 2019</xref>) and Jamali et&#x20;al. (<xref ref-type="bibr" rid="B19">Jamali et&#x20;al., 2016</xref>). The study found that the accuracy of the original data set using the DrugHybrid_BS model reached 100%, which shows that the original data does have redundancy, and it also reflects the significance of the initial data preprocessing in this article.</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>Comparison of prediction performance with other algorithms.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">ACC(%)</th>
<th align="center">TPR</th>
<th align="center">FPR</th>
<th align="center">Precision</th>
<th align="center">F-score</th>
<th align="center">auROC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">DrugHybrid_BS(This paper)</td>
<td align="char" char=".">96.994</td>
<td align="char" char=".">0.970</td>
<td align="char" char=".">0.030</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.967</td>
<td align="char" char=".">0.992</td>
</tr>
<tr>
<td align="left">DrugHybrid_KNN</td>
<td align="char" char=".">58.652</td>
<td align="char" char=".">0.587</td>
<td align="char" char=".">0.502</td>
<td align="char" char=".">0.729</td>
<td align="char" char=".">0.473</td>
<td align="char" char=".">0.625</td>
</tr>
<tr>
<td align="left">DrugHybrid_SVM</td>
<td align="char" char=".">96.694</td>
<td align="char" char=".">0.959</td>
<td align="char" char=".">0.027</td>
<td align="char" char=".">0.967</td>
<td align="char" char=".">0.963</td>
<td align="char" char=".">0.966</td>
</tr>
<tr>
<td align="left">DrugHybrid_RF</td>
<td align="char" char=".">87.763</td>
<td align="char" char=".">0.834</td>
<td align="char" char=".">0.087</td>
<td align="char" char=".">0.888</td>
<td align="char" char=".">0.860</td>
<td align="char" char=".">0.949</td>
</tr>
<tr>
<td align="left">DrugHybrid_BS(Original dataset)</td>
<td align="char" char=".">100</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.000</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">Jamali et&#x20;al. (<xref ref-type="bibr" rid="B19">Jamali et&#x20;al., 2016</xref>) (Original dataset)</td>
<td align="char" char=".">89.78</td>
<td align="char" char=".">0.901</td>
<td align="char" char=".">0.106</td>
<td align="char" char=".">0.901</td>
<td align="char" char=".">0.901</td>
<td align="char" char=".">0.959</td>
</tr>
<tr>
<td align="left">Lin et&#x20;al. (<xref ref-type="bibr" rid="B28">Lin et&#x20;al., 2019</xref>) (Original dataset)</td>
<td align="char" char=".">93.78</td>
<td align="char" char=".">0.928</td>
<td align="char" char=".">0.056</td>
<td align="char" char=".">0.942</td>
<td align="char" char=".">0.936</td>
<td align="char" char=".">0.978</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-6">
<title>3.6 Independent Test Set</title>
<p>The accuracy of the classification and prediction model in predicting the training set cannot well reflect the future performance of the prediction model. To effectively judge the performance of a predictive model, we divided 80% of the dataset as the training set and 20% as the test set. The detailed information was shown in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. The independent test set using the DrugHybrid_BS model can accurately predict 96.5665% of the potentially druggable protein. The TPR value was 0.948, the FPR value was 0.02, the precision value was 0.975, and the AUC value was&#x20;0.990.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Details of the training set and independent test&#x20;set.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g006.tif"/>
</fig>
</sec>
<sec id="s3-7">
<title>3.7 Feature Importance Analysis</title>
<p>From the DrugHybrid_BS model, we obtained the following: after combining the single feature representation methods, the hybrid features of monoDiKGap, CC and GAAC combined with Bagging-SVM can improve the accuracy of predicting druggable proteins. This part further explored the features that play a key role in the DrugHybrid_BS model, that is, the importance of these features.</p>
<p>First, we used the MRMD2.0 to sort the feature sets extracted by three single feature representation methods and simultaneously obtained the relationship between the number of features and the accuracy of predicting potential druggable proteins (<xref ref-type="fig" rid="F7">Figures 7A&#x2013;C</xref>). <xref ref-type="fig" rid="F7">Figure&#x20;7A</xref> showed that when the number of features of the CC feature extraction method was more than eight, the accuracy rate reached more than 60% and continued to grow. Therefore, we selected the top eight features as the key features of the CC feature representation method. <xref ref-type="fig" rid="F7">Figure&#x20;7B</xref> showed the GAAC feature representation method. When the number of features was two, the accuracy rate reached more than 70%, and the accuracy rate continued to increase as the number increased. Therefore, we selected the top two features as the key features of the GAAC feature extraction method. <xref ref-type="fig" rid="F7">Figure&#x20;7C</xref> showed the monoDiKGap feature extraction method. When the number of features was twenty-six, the accuracy of predicting potentially druggable proteins was significantly improved, and then the accuracy increased steadily as the number of features increased. Therefore, we chose the top twenty-six features as the key features of the monoDiKGap feature extraction method. Second, we combined the key features of the single feature extraction methods to obtain the hybrid key features. The detailed information was shown in <xref ref-type="table" rid="T7">Table&#x20;7</xref>. Finally, the number of base models suitable for the hybrid key features was selected through the Bagging-SVM classification&#x20;model.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>The relationship between the number of features extracted by the three methods and the accuracy of predicting potentially druggable proteins <bold>(A)</bold> CC <bold>(B)</bold> GAAC, and <bold>(C)</bold> monoDiKGap.</p>
</caption>
<graphic xlink:href="fphar-12-771808-g007.tif"/>
</fig>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Key feature details of each feature representation method.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th colspan="9" align="center">Key features</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">GAAC</td>
<td colspan="3" align="left">Aromatic group</td>
<td colspan="6" align="left">Uncharge group</td>
</tr>
<tr>
<td rowspan="3" align="left">CC</td>
<td colspan="3" align="left">(mass, hydrophobicity,1)</td>
<td colspan="3" align="left">(mass, hydrophilicity,1)</td>
<td colspan="3" align="left">(hydrophilicity, mass,1)</td>
</tr>
<tr>
<td colspan="3" align="left">(mass, hydrophobicity,2)</td>
<td colspan="3" align="left">(hydrophobicity, mass,2)</td>
<td colspan="3" align="left">(hydrophobicity, mass,1)</td>
</tr>
<tr>
<td colspan="3" align="left">(hydrophilicity, mass,2)</td>
<td colspan="6" align="left">(hydrophilicity, hydrophobicity,1)</td>
</tr>
<tr>
<td rowspan="4" align="left">monoDiKGap</td>
<td align="left">C_ _NQ</td>
<td align="left">C_ _RT</td>
<td colspan="2" align="left">E_ DT</td>
<td align="left">W_ _PR</td>
<td colspan="2" align="left">E_ _VW</td>
<td align="left">T_ _IL</td>
<td align="left">T_ _PN</td>
</tr>
<tr>
<td align="left">I_RH</td>
<td align="left">Q_ _SA</td>
<td colspan="2" align="left">K_ _IY</td>
<td align="left">L_ _HY</td>
<td colspan="2" align="left">N_TD</td>
<td align="left">T_YK</td>
<td align="left">E_ _DI</td>
</tr>
<tr>
<td align="left">Y_ _LI</td>
<td align="left">R_ _MH</td>
<td colspan="2" align="left">T_ _YY</td>
<td align="left">N_DD</td>
<td colspan="2" align="left">P_RQ</td>
<td align="left">R_ _CT</td>
<td align="left">S_ _GL</td>
</tr>
<tr>
<td align="left">E_VC</td>
<td align="left">P_NY</td>
<td colspan="2" align="left">D_KK</td>
<td align="left">N_PK</td>
<td colspan="2" align="left">F_ _LK</td>
<td align="left">&#x2014;</td>
<td align="left">&#x2014;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>After research, we obtained that the hybrid key features can accurately predict 80.0343% of the potentially druggable proteins under the bagging algorithm based on the integration of fifteen SVMs. These hybrid key features combined with Bagging-SVM have achieved good prediction results, which fully demonstrated the importance of this part of the feature for the new method DrugHybrid_BS for predicting potentially druggable proteins.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>Research on potentially druggable proteins is of great significance in the field of drug development and disease treatment. However, identifying potentially druggable proteins is the first step in research. This research focused on combining hybrid features and Bagging-SVM to predict potentially druggable proteins. The hybrid features included three feature extraction methods: monoDiKGap, CC, and GAAC, which were based on sequence information, physiochemical properties, and correlation. Through the three single feature representation methods of monoDiKGap, CC, GAAC, and the comparison of combined feature prediction, it was found that the hybrid features of monoDiKGap, CC, and GAAC can accurately predict 96.9944% of the potentially druggable proteins under Bagging-SVM. In addition, the accuracy of the independent test set using the new method DrugHybrid_BS reached 96.5665%. Therefore, the DrugHybrid_BS model used in this study could be a powerful method to study potentially druggable proteins and provide a reference value for other studies. In the future, we will try more deep learning techniques (<xref ref-type="bibr" rid="B75">Zou et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B12">Guo et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B65">Zeng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B36">Niu et&#x20;al., 2021b</xref>; <xref ref-type="bibr" rid="B68">Zhang et&#x20;al., 2021</xref>) for this problem.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s10">Supplementary Material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>Conceptualization, BL and QZ; data collection or analysis, YG and PW; validation, YG; writing&#x2014;original draft preparation, YG; writing&#x2014;review and editing, YG. and QZ All authors have read and agreed to the published version of the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was supported by the National Nature Science Foundation of China (Grant Nos 61863010, 11926205, 11926412, and 61873076), National Key R&#x26;D Program of China (No.2020YFB2104400), Natural Science Foundation of Hainan, China(Grant Nos. 119MS036 and 120RC588), Hainan Normal University 2020 Graduate Student Innovation Research Project (hsyx 2020-40) and the Special Science Foundation of Quzhou (2020D003).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fphar.2021.771808/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fphar.2021.771808/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.xlsx" id="SM1" mimetype="application/xlsx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ao</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>RFhy-m2G: Identification of RNA N2-Methylguanosine Modification Sites Based on Random Forest and Hybrid Features</article-title>. <source>Methods</source>. <pub-id pub-id-type="doi">10.1016/j.ymeth.2021.05.016</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheng</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Drug Target Prediction Using Graph Representation Learning via Substructures Contrast</article-title>, <source>Appl. Sci.</source>, <volume>11</volume>, <fpage>3239</fpage>. <pub-id pub-id-type="doi">10.3390/app11073239</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dezs&#x151;</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ceccarelli</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Machine Learning Prediction of Oncology Drug Targets Based on Protein and Network Properties</article-title>. <source>BMC Bioinformatics</source> <volume>21</volume>, <fpage>104</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3442-9</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jijun</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020a</year>). <article-title>Identification of Drug-Target Interactions via Dual Laplacian Regularized Least Squares with Multiple Kernel Fusion</article-title>. <source>Knowledge-Based Syst.</source> <volume>204</volume>, <fpage>106254</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2020.106254</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2019a</year>). <article-title>Identification of Drug-Side Effect Association via Semisupervised Model and Multiple Kernel Learning</article-title>. <source>IEEE J.&#x20;Biomed. Health Inform.</source> <volume>23</volume>, <fpage>2619</fpage>&#x2013;<lpage>2632</lpage>. <pub-id pub-id-type="doi">10.1109/jbhi.2018.2883834</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2019b</year>). <article-title>Identification of Drug-Side Effect Association via Multiple Information Integration with Centered Kernel Alignment</article-title>. <source>Neurocomputing</source> <volume>325</volume>, <fpage>211</fpage>&#x2013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2018.10.028</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020b</year>). <article-title>Identification of Drug-Target Interactions via Fuzzy Bipartite Local Model</article-title>. <source>Neural Comput. Applic</source> <volume>32</volume>, <fpage>10303</fpage>&#x2013;<lpage>10319</lpage>. <pub-id pub-id-type="doi">10.1007/s00521-019-04569-z</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Identification of Drug-Target Interactions via Multiple Information Integration</article-title>. <source>Inf. Sci.</source> <volume>418-419</volume>, <fpage>546</fpage>&#x2013;<lpage>560</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2017.08.045</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dudoit</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fridlyand</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Bagging to Improve the Accuracy of a Clustering Procedure</article-title>. <source>Bioinformatics</source> <volume>19</volume>, <fpage>1090</fpage>&#x2013;<lpage>1099</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btg038</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Niu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>CD-HIT: Accelerated for Clustering the Next-Generation Sequencing Data</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>3150</fpage>&#x2013;<lpage>3152</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts565</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gayvert</surname>
<given-names>K. M.</given-names>
</name>
<name>
<surname>Madhukar</surname>
<given-names>N. S.</given-names>
</name>
<name>
<surname>Elemento</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Data-Driven Approach to Predicting Successes and Failures of Clinical Trials</article-title>. <source>Cell Chem Biol</source> <volume>23</volume>, <fpage>1294</fpage>&#x2013;<lpage>1301</lpage>. <pub-id pub-id-type="doi">10.1016/j.chembiol.2016.07.023</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A Deep Convolutional Neural Network to Improve the Prediction of Protein Secondary Structure</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>767</fpage>&#x2013;<lpage>777</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666200120103050</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Using Support Vector Machine Combined with Auto Covariance to Predict Protein-Protein Interactions from Protein Sequences</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>3025</fpage>&#x2013;<lpage>3030</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn159</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Predicting Ion Channels Genes and Their Types with Machine Learning Techniques</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <fpage>399</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00399</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>HuiDing</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MRMD2.0: A Python Tool for Machine Learning with Feature Ranking and Reduction</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>1213</fpage>&#x2013;<lpage>1221</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200503030350</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hopkins</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Groom</surname>
<given-names>C. R.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>The Druggable Genome</article-title>. <source>Nat. Rev. Drug Discov.</source> <volume>1</volume>, <fpage>727</fpage>&#x2013;<lpage>730</lpage>. <pub-id pub-id-type="doi">10.1038/nrd892</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Prediction of Transcription Factors Binding Events Based on Epigenetic Modifications in Different Human Cells</article-title>. <source>Epigenomics</source> <volume>12</volume>, <fpage>1443</fpage>&#x2013;<lpage>1456</lpage>. <pub-id pub-id-type="doi">10.2217/epi-2019-0321</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>SGL-SVM: A Novel Method for Tumor Classification via Support Vector Machine with Sparse Group Lasso</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>486</volume>, <fpage>110098</fpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2019.110098</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jamali</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Ferdousi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Razzaghi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Safdari</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ebrahimie</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>DrugMiner: Comparative Analysis of Machine Learning Algorithms for Prediction of Potential Druggable Proteins</article-title>. <source>Drug Discov. Today</source> <volume>21</volume>, <fpage>718</fpage>&#x2013;<lpage>724</lpage>. <pub-id pub-id-type="doi">10.1016/j.drudis.2016.01.007</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ji</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Freudenberg</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Agarwal</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Integrating Biological Networks for Drug Target Prediction and Prioritization</article-title>. <source>Methods Mol. Biol.</source> <volume>1903</volume>, <fpage>203</fpage>&#x2013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-8955-3_12</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Predicting Human microRNA-Disease Associations Based on Support Vector Machine</article-title>. <source>Int. J.&#x20;Data Min Bioinform</source> <volume>8</volume>, <fpage>282</fpage>&#x2013;<lpage>293</lpage>. <pub-id pub-id-type="doi">10.1504/ijdmb.2013.056078</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Cui</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Free-form Tumor Synthesis in Computed Tomography Images via Richer Generative Adversarial Network</article-title>. <source>Knowledge-Based Syst.</source> <volume>218</volume>, <fpage>106753</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2021.106753</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Pham</surname>
<given-names>T. D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>DUNet: A Deformable Network for Retinal Vessel Segmentation</article-title>. <source>Knowledge-Based Syst.</source> <volume>178</volume>, <fpage>149</fpage>&#x2013;<lpage>162</lpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2019.04.025</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>T. Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>Z. Q.</given-names>
</name>
<name>
<surname>Hsieh</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Breta&#xf1;a</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>C. T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Exploiting Maximal Dependence Decomposition to Identify Conserved Motifs from a Group of Aligned Signal Sequences</article-title>. <source>Bioinformatics</source> <volume>27</volume>, <fpage>1780</fpage>&#x2013;<lpage>1787</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btr291</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Lai</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Prediction of Potential Drug Targets Based on Simple Sequence Properties</article-title>. <source>BMC Bioinformatics</source> <volume>8</volume>, <fpage>353</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-8-353</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mo</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>A Machine Learning Approach for Tracing Tumor Original Sites with Gene Expression Profiles</article-title>. <source>Front. Bioeng. Biotechnol.</source> <volume>8</volume>, <fpage>607126</fpage>. <pub-id pub-id-type="doi">10.3389/fbioe.2020.607126</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Vemuri</surname>
<given-names>V. R.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Use of K-Nearest Neighbor Classifier for Intrusion Detection</article-title>. <source>Comput. Security</source> <volume>21</volume>, <fpage>439</fpage>&#x2013;<lpage>448</lpage>. <pub-id pub-id-type="doi">10.1016/s0167-4048(02)00514-x</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Accurate Prediction of Potential Druggable Proteins Based on Genetic Algorithm and Bagging-SVM Ensemble Classifier</article-title>. <source>Artif. Intell. Med.</source> <volume>98</volume>, <fpage>35</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2019.07.005</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zuo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Function Determinants of TET Proteins: the Arrangements of Sequence Motifs with Specific Codes</article-title>. <source>Brief Bioinform</source> <volume>20</volume>, <fpage>1826</fpage>&#x2013;<lpage>1835</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bby053</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>BioSeq-Analysis2.0: an Updated Platform for Analyzing DNA, RNA and Protein Sequences at Sequence Level and Residue Level Based on Machine Learning Approaches</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>e127</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz740</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Classification and Gene Selection of Triple-Negative Breast Cancer Subtype Embedding Gene Connectivity Matrix in Deep Neural Network</article-title>. <source>Brief. Bioinform.</source> <volume>22</volume>, <fpage>bbaa395</fpage>. <comment>LID - bbaa395 [pii] LID - 10.1093/bib/bbaa395 [doi]</comment>. <pub-id pub-id-type="doi">10.1093/bib/bbaa395</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lv</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z. M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>J.&#x20;X.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Evaluation of Different Computational Methods on 5-methylcytosine Sites Identification</article-title>. <source>Brief Bioinform</source> <volume>21</volume>, <fpage>982</fpage>&#x2013;<lpage>995</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz048</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>CWLy-SVM: A Support Vector Machine-Based Tool for Identifying Cell wall Lytic Enzymes</article-title>. <source>Comput. Biol. Chem.</source> <volume>87</volume>, <fpage>107304</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2020.107304</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Munir</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>S. I.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>K. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Proteome Mining for the Identification of Putative Drug Targets for Human Pathogen <italic>Clostridium tetani</italic>
</article-title>. <source>Curr. Bioinformatics</source> <volume>14</volume>, <fpage>532</fpage>&#x2013;<lpage>540</lpage>. <pub-id pub-id-type="doi">10.2174/1574893613666181114095736</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>sgRNACNN: Identifying sgRNA On-Target Activity in Four Crops Using Ensembles of Convolutional Neural Networks</article-title>. <source>Plant Mol. Biol.</source> <volume>105</volume>, <fpage>483</fpage>&#x2013;<lpage>495</lpage>. <pub-id pub-id-type="doi">10.1007/s11103-020-01102-y</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Niu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>rBPDL:Predicting RNA-Binding Proteins Using Deep Learning</article-title>. <source>IEEE J.&#x20;Biomed. Health Inform.</source> <volume>25</volume>, <fpage>3668</fpage>&#x2013;<lpage>3676</lpage>. <pub-id pub-id-type="doi">10.1109/jbhi.2021.3069259</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pacheco</surname>
<given-names>M. P.</given-names>
</name>
<name>
<surname>Bintener</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ternes</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kulms</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Haan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Letellier</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Identifying and Targeting Cancer-specific Metabolism with Network-Based Drug Target Prediction</article-title>. <source>EBioMedicine</source> <volume>43</volume>, <fpage>98</fpage>&#x2013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1016/j.ebiom.2019.04.046</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Platt</surname>
<given-names>J.&#x20;C.</given-names>
</name>
</person-group> (<year>1998</year>). <source>Sequential Minimal Optimization: A Fast Algorithm for Training Support Vector Machines</source>. <publisher-name>Technical Report MSR-TR-98-14</publisher-name>. </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Quan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Novel Features Ranking Metric with Application to Scalable Visual and Bioinformatics Data Classification</article-title>. <source>Neurocomputing</source> <volume>173</volume>, <fpage>346</fpage>&#x2013;<lpage>354</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2014.12.123</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ru</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Exploration of the Correlation between GPCRs and Drugs Based on a Learning to Rank Algorithm</article-title>. <source>Comput. Biol. Med.</source> <volume>119</volume>, <fpage>103660</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2020.103660</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Russ</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Lampel</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>The Druggable Genome: an Update</article-title>. <source>Drug Discov. Today</source> <volume>10</volume>, <fpage>1607</fpage>&#x2013;<lpage>1610</lpage>. <pub-id pub-id-type="doi">10.1016/s1359-6446(05)03666-4</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salmaso</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Moro</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Bridging Molecular Docking to Molecular Dynamics in Exploring Ligand-Protein Recognition Process: An Overview</article-title>. <source>Front. Pharmacol.</source> <volume>9</volume>, <fpage>923</fpage>. <pub-id pub-id-type="doi">10.3389/fphar.2018.00923</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Samanthula</surname>
<given-names>B. K.</given-names>
</name>
<name>
<surname>Elmehdwi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>K-Nearest Neighbor Classification over Semantically Secure Encrypted Relational Data</article-title>. <source>IEEE Trans. Knowledge Data Eng.</source> <volume>27</volume>, <fpage>1261</fpage>&#x2013;<lpage>1273</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2014.2364027</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Prediction of Drug-Target Interactions Based on Multi-Layer Network Representation Learning</article-title>. <source>Neurocomputing</source> <volume>434</volume>, <fpage>80</fpage>&#x2013;<lpage>89</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2020.12.068</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shi</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Predicting Drug-Target Interactions Using Lasso with Random forest Based on Evolutionary Information and Chemical Structure</article-title>. <source>Genomics</source> <volume>111</volume>, <fpage>1839</fpage>&#x2013;<lpage>1852</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2018.12.007</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sokolova</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Japkowicz</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Szpakowicz</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2006</year>). <source>Beyond Accuracy, F-Score and ROC: A Family of Discriminant Measures for Performance Evaluation</source>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Developing a Multi-Dose Computational Model for Drug-Induced Hepatotoxicity Prediction Based on Toxicogenomics Data</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinform</source> <volume>16</volume>, <fpage>1231</fpage>&#x2013;<lpage>1239</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2018.2858756</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Predicting Drug-Target Interactions via FM-DNN Learning</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>68</fpage>&#x2013;<lpage>76</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666190227160538</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Drug Target Interaction Prediction Based on LINE-RF Learning</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>750</fpage>&#x2013;<lpage>757</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666191227092453</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020c</year>). <article-title>Identification of Membrane Protein Types via Multivariate Information Fusion with Hilbert-Schmidt Independence Criterion</article-title>. <source>Neurocomputing</source> <volume>383</volume>, <fpage>257</fpage>&#x2013;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.11.103</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Lv</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Pancreatic Cancer Biomarker Detection by Two Support Vector Strategies for Recursive Feature Elimination</article-title>. <source>Biomark Med.</source> <volume>13</volume>, <fpage>105</fpage>&#x2013;<lpage>121</lpage>. <pub-id pub-id-type="doi">10.2217/bmm-2018-0273</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zuo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Modular Arrangements of Sequence Motifs Determine the Functional Diversity of KDM Proteins</article-title>. <source>Brief. Bioinformatics</source> <volume>22</volume>, <fpage>bbaa215</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa215</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Luan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nagai</surname>
<given-names>L. A. E.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Exploring Sequence-Based Features for the Improved Prediction of DNA N4-Methylcytosine Sites in Multiple Species</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>1326</fpage>&#x2013;<lpage>1333</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty824</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fast Prediction of Protein Methylation Sites Using a Sequence-Based Feature Selection Technique</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinform</source> <volume>16</volume>, <fpage>1264</fpage>&#x2013;<lpage>1273</lpage>. <pub-id pub-id-type="doi">10.1109/tcbb.2017.2670558</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Prediction of Human Protein Subcellular Localization Using Deep Learning</article-title>. <source>J.&#x20;Parallel Distributed Comput.</source> <volume>117</volume>, <fpage>212</fpage>&#x2013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1016/j.jpdc.2017.08.009</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>ACPred-FL: a Sequence-Based Predictor Using Effective Feature Representation to Improve the Prediction of Anti-cancer Peptides</article-title>. <source>Bioinformatics</source> <volume>34</volume>, <fpage>4007</fpage>&#x2013;<lpage>4016</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty451</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Local-DPP: An Improved DNA-Binding Protein Prediction Method by Exploring Local Evolutionary Information</article-title>. <source>Inf. Sci.</source> <volume>384</volume>, <fpage>135</fpage>&#x2013;<lpage>144</lpage>. <pub-id pub-id-type="doi">10.1016/j.ins.2016.06.026</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Improved Prediction of Protein-Protein Interactions Using Novel Negative Samples, Features, and an Ensemble Classifier</article-title>. <source>Artif. Intell. Med.</source> <volume>83</volume>, <fpage>67</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1016/j.artmed.2017.03.001</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wishart</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Knox</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Shrivastava</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hassanali</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Stothard</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>DrugBank: a Comprehensive Resource for In Silico Drug Discovery and Exploration</article-title>. <source>Nucleic Acids Res.</source> <volume>34</volume>, <fpage>D668</fpage>&#x2013;<lpage>D672</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkj067</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>EPSOL: Sequence-Based Protein Solubility Prediction Using Multidimensional Embedding</article-title>. <source>Bioinformatics (Oxford, England)</source>, <fpage>btab463</fpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab463</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zuo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Multi-substrate Selectivity Based on Key Loops and Non-homologous Domains: New Insight into ALKBH Family</article-title>. <source>Cell Mol Life Sci</source> <volume>78</volume>, <fpage>129</fpage>&#x2013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1007/s00018-020-03594-9</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>SeqSVM: A Sequence-Based Support Vector Machine Method for Identifying Antioxidant Proteins</article-title>. <source>Int. J.&#x20;Mol. Sci.</source> <volume>19</volume>. <pub-id pub-id-type="doi">10.3390/ijms19061773</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Xue</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>DLpTCR: an Ensemble Deep Learning Framework for Predicting Immunogenic Peptide Recognized by T&#x20;Cell Receptor</article-title>. <source>Brief Bioinform</source> <volume>22</volume>, <fpage>bbab335</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab335</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Predicting Therapeutic Drugs for Hepatocellular Carcinoma Based on Tissue-specific Pathways</article-title>. <source>Plos Comput. Biol.</source> <volume>17</volume>, <fpage>e1008696</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1008696</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Predicting Disease-Associated Circular RNAs Using Deep Forests Combined with Positive-Unlabeled Learning Methods</article-title>. <source>Brief Bioinform</source> <volume>21</volume>, <fpage>1425</fpage>&#x2013;<lpage>1436</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz080</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Z. C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>iPromoter-5mC: A Novel Fusion Decision Predictor for the Identification of 5-Methylcytosine Sites in Genome-wide DNA Promoters</article-title>. <source>Front Cel Dev Biol</source> <volume>8</volume>, <fpage>614</fpage>. <pub-id pub-id-type="doi">10.3389/fcell.2020.00614</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Sa</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Discriminating Ramos and Jurkat Cells with Image Textures from Diffraction Imaging Flow Cytometry Based on a Support Vector Machine</article-title>. <source>Curr. Bioinformatics</source> <volume>11</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.2174/1574893611666160608102537</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Review of the Applications of Deep Learning in Bioinformatics</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>898</fpage>&#x2013;<lpage>911</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615999200711165743</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>RAACBook: a Web Server of Reduced Amino Acid Alphabet for Sequence-dependent Inference by Using Chou&#x27;s Five-step Rule</article-title>. <source>Database (Oxford)</source> <volume>2019</volume>, <fpage>baz131</fpage>. <pub-id pub-id-type="doi">10.1093/database/baz131</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zuo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>RaacLogo: a New Sequence Logo Generator by Using Reduced Amino Acid Clusters</article-title>. <source>Brief. Bioinformatics</source> <volume>22</volume>, <fpage>bbaa096</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa096</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhong</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Artificial Intelligence in Drug Design</article-title>. <source>Sci. China Life Sci.</source> <volume>61</volume>, <fpage>1191</fpage>&#x2013;<lpage>1204</lpage>. <pub-id pub-id-type="doi">10.1007/s11427-018-9342-2</pub-id> </citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Arbor</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hastie</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Multi-class AdaBoost</article-title>. <source>Stat. Its Interf.</source> <volume>2</volume>, <fpage>349</fpage>&#x2013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.4310/SII.2009.v2.n3.a8</pub-id> </citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Xiang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Akutsu</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Computational Identification of Eukaryotic Promoters Based on Cascaded Deep Capsule Neural Networks</article-title>. <source>Brief Bioinform</source> <volume>22</volume>, <fpage>bbaa299</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa299</pub-id> </citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhuang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Identifying Breast Cancer-Induced Gene Perturbations and its Application in Guiding Drug Repurposing</article-title>. <source>Curr. Bioinformatics</source> <volume>15</volume>, <fpage>1075</fpage>&#x2013;<lpage>1089</lpage>. <pub-id pub-id-type="doi">10.2174/1574893615666200203104214</pub-id> </citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Xing</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Gene2vec: Gene Subsequence Embedding for Prediction of Mammalian N 6-methyladenosine Sites from mRNA</article-title>. <source>RNA</source> <volume>25</volume>, <fpage>205</fpage>&#x2013;<lpage>218</lpage>. <pub-id pub-id-type="doi">10.1261/rna.069112.118</pub-id> </citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Sequence Clustering in Bioinformatics: an Empirical Study</article-title>. <source>Brief. Bioinform.</source> <volume>21</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bby090</pub-id> </citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zuo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>PseKRAAC: a Flexible Web Server for Generating Pseudo K-Tuple Reduced Amino Acids Composition</article-title>. <source>Bioinformatics</source> <volume>33</volume>, <fpage>122</fpage>&#x2013;<lpage>124</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btw564</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>