<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">752732</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.752732</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Deep Learning and XGBoost-Based Method for Predicting Protein-Protein Interaction Sites</article-title>
<alt-title alt-title-type="left-running-head">Wang et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Protein-Protein Interaction Site Prediction</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Pan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1459502/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Zhang</surname>
<given-names>Guiyang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1503111/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yu</surname>
<given-names>Zu-Guo</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Huang</surname>
<given-names>Guohua</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/801622/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>School of Electrical Engineering, Shaoyang University, <addr-line>Shaoyang</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Key Laboratory of Intelligent Computing and Information Processing of Ministry of Education and Hunan Key Laboratory for Computation and Simulation in Science and Engineering, Xiangtan University, <addr-line>Xiangtan</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/664933/overview">Lei Wang</ext-link>, Changsha University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1442135/overview">Qi Dai</ext-link>, Zhejiang Sci-Tech University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1442498/overview">Zhaohui Qi</ext-link>, Hunan Normal University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Guohua Huang, <email>guohuahhn@163.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>26</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>752732</elocation-id>
<history>
<date date-type="received">
<day>03</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Wang, Zhang, Yu and Huang.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Wang, Zhang, Yu and Huang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Knowledge about protein-protein interactions is beneficial in understanding cellular mechanisms. Protein-protein interactions are usually determined according to their protein-protein interaction sites. Due to the limitations of current techniques, it is still a challenging task to detect protein-protein interaction sites. In this article, we presented a method based on deep learning and XGBoost (called DeepPPISP-XGB) for predicting protein-protein interaction sites. The deep learning model served as a feature extractor to remove redundant information from protein sequences. The Extreme Gradient Boosting algorithm was used to construct a classifier for predicting protein-protein interaction sites. The DeepPPISP-XGB achieved the following results: area under the receiver operating characteristic curve of 0.681, a recall of 0.624, and area under the precision-recall curve of 0.339, being competitive with the state-of-the-art methods. We also validated the positive role of global features in predicting protein-protein interaction&#x20;sites.</p>
</abstract>
<kwd-group>
<kwd>protein-protein interaction</kwd>
<kwd>deep learning</kwd>
<kwd>machine learning</kwd>
<kwd>extreme gradient boosting</kwd>
<kwd>protein functions</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>Proteins are one of the most important components of the cell, and also are the principal undertaker of the activities of life. The functions of proteins are manifested mainly by interacting with various molecules such as DNA/RNA, proteins, or other ligands (<xref ref-type="bibr" rid="B27">Dias and Kolaczkowski, 2017</xref>). The protein-protein interaction (PPI) plays a key role in the cellular process such as signal transduction, transport, and metabolism (<xref ref-type="bibr" rid="B58">Li et&#x20;al., 2019</xref>) and also is involved in the pathogenesis of diseases such as Alzheimer&#x2019;s cervical cancer, bacterial infection, and prion diseases (<xref ref-type="bibr" rid="B19">Cohen and Prusiner, 1998</xref>; <xref ref-type="bibr" rid="B79">Selkoe, 1998</xref>; <xref ref-type="bibr" rid="B62">Loregian et&#x20;al., 2002</xref>). Therefore, knowledge of PPI is critical for understanding the molecular mechanisms hidden in the phenomenon of life (<xref ref-type="bibr" rid="B20">Das and Chakrabarti, 2021</xref>). Many experimentally verified or computationally predicted PPIs have been hosted for scientific research in public databases such as the Human Protein Reference Database (<xref ref-type="bibr" rid="B50">Keshava Prasad et&#x20;al., 2009</xref>), STRING (<xref ref-type="bibr" rid="B87">Von Mering et&#x20;al., 2005</xref>), the database of interacting proteins (<xref ref-type="bibr" rid="B77">Salwinski et&#x20;al., 2004</xref>), and the protein interaction database (<xref ref-type="bibr" rid="B49">Kerrien et&#x20;al., 2007</xref>). The protein-protein interaction site (PPIS) is defined as surface residues where proteins interact with each other (<xref ref-type="bibr" rid="B2">Aumentado-Armstrong et&#x20;al., 2015</xref>). The identification of PPIS is the premise for determining PPI (<xref ref-type="bibr" rid="B91">Wang et&#x20;al., 2019</xref>). The knowledge about PPIS holds vast potential to infer cell regulatory mechanisms, locate drug targets, identify structures and functions of protein complexes (<xref ref-type="bibr" rid="B26">Deng et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B69">Orii and Ganapathiraju, 2012</xref>), and uncover disease pathogenesis (<xref ref-type="bibr" rid="B54">Kuzmanov and Emili, 2013</xref>). Drug discovery and development are also closely associated with PPIS (<xref ref-type="bibr" rid="B83">Sperandio, 2012</xref>; <xref ref-type="bibr" rid="B71">Petta et&#x20;al., 2016</xref>). Therefore, identifying PPIS is of great importance in the field of molecule biology.</p>
<p>It is not only costly but also time-consuming and labor-intensive to identify PPIS by experimental methods such as alanine scanning mutagenesis and crystallographic complex determination (<xref ref-type="bibr" rid="B2">Aumentado-Armstrong et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B52">Kr&#xef;&#xbf; &#xbd;ger and Gohlke, 2010</xref>; <xref ref-type="bibr" rid="B8">Bradshaw et&#x20;al., 2011</xref>). Since Jones and Thornton pioneered a computational method for predicting and analyzing PPIS in 1997 (<xref ref-type="bibr" rid="B45">Jones and Thornton, 1997</xref>; <xref ref-type="bibr" rid="B46">Jones and Thornton, 1997</xref>), more than thirty other computational methods have been developed (<xref ref-type="bibr" rid="B104">Zhou and Shan, 2001</xref>; <xref ref-type="bibr" rid="B33">Fernandez-Recio et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B66">Neuvirth et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B7">Bradford and Westhead, 2005</xref>; <xref ref-type="bibr" rid="B14">Chen and Zhou, 2005</xref>; <xref ref-type="bibr" rid="B18">Chung et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B60">Liang et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B70">Patel et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B57">Li et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B68">Ofran and Rost, 2007</xref>; <xref ref-type="bibr" rid="B72">Porollo and Meller, 2007</xref>; <xref ref-type="bibr" rid="B73">Qin and Zhou, 2007</xref>; <xref ref-type="bibr" rid="B85">Tjong et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B16">Chen and Jeong, 2009</xref>; <xref ref-type="bibr" rid="B29">Doszt&#xe1;nyi et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B30">Du et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B32">Engelen et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B81">&#x160;iki&#x107; et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B34">Fiorucci and Zacharias, 2010</xref>; <xref ref-type="bibr" rid="B65">Murakami and Mizuguchi, 2010</xref>; <xref ref-type="bibr" rid="B80">Shoemaker et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B78">Segura et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B97">Xue et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B102">Zhang et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B13">Chen et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B47">Jordan et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B55">La and Kihara, 2012</xref>; <xref ref-type="bibr" rid="B56">Li et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B74">Qiu and Wang, 2012</xref>; <xref ref-type="bibr" rid="B98">Zellner et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B4">Bendell et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B22">de Moraes et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B82">Singh et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B89">Wang et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B2">Aumentado-Armstrong et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B3">Bagchi, 2015</xref>; <xref ref-type="bibr" rid="B21">Dayal et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B63">Maheshwari and Brylinski, 2015</xref>; <xref ref-type="bibr" rid="B28">Dick and Green, 2016</xref>; <xref ref-type="bibr" rid="B43">Jia et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B53">Kuo and Li, 2016</xref>; <xref ref-type="bibr" rid="B95">Wei et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B40">Hou et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B103">Zhao et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B37">Guo et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B67">Northey et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B93">Wang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B88">Wang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B101">Zhang and Kurgan, 2019</xref>; <xref ref-type="bibr" rid="B101">Zhang and Kurgan, 2019</xref>; <xref ref-type="bibr" rid="B25">Deng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B59">Li, 2020</xref>; <xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B105">Zhu et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B94">Wang et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B92">Wang et&#x20;al., 2021</xref>). Due to their efficiency, computational methods are becoming essentially complementary to experimental methods. Most computational methods for identifying PPIS are based on machine learning algorithms where the prediction performance depends heavily on learning algorithms and feature extractions. The learning algorithms used for PPIS prediction generally include conditional random fields (<xref ref-type="bibr" rid="B57">Li et&#x20;al., 2007</xref>), support vector machines (<xref ref-type="bibr" rid="B7">Bradford and Westhead, 2005</xref>), random forest (<xref ref-type="bibr" rid="B16">Chen and Jeong, 2009</xref>), XGBoost (<xref ref-type="bibr" rid="B25">Deng et&#x20;al., 2020</xref>), logistic regression (<xref ref-type="bibr" rid="B101">Zhang and Kurgan, 2019</xref>), Bayes method (<xref ref-type="bibr" rid="B65">Murakami and Mizuguchi, 2010</xref>), and artificial neural networks (<xref ref-type="bibr" rid="B82">Singh et&#x20;al., 2014</xref>). These learning algorithms are not suitable for enough large number of training samples. Recently, deep learning algorithms have been developed that have achieved significant superiority over traditional learning algorithms, especially in many difficult cases such as image classification (<xref ref-type="bibr" rid="B51">Krizhevsky et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B39">He et&#x20;al., 2016</xref>) and protein structure prediction (<xref ref-type="bibr" rid="B11">Callaway, 2020</xref>). Features used for PPIS prediction generally include evolutionary information (<xref ref-type="bibr" rid="B10">Caffrey et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B12">Carl et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B17">Choi et&#x20;al., 2009</xref>), secondary structure (<xref ref-type="bibr" rid="B36">Guharoy and Chakrabarti, 2007</xref>; <xref ref-type="bibr" rid="B68">Ofran and Rost, 2007</xref>; <xref ref-type="bibr" rid="B56">Li et&#x20;al., 2012</xref>) and physicochemical, biophysical and statistical features such as accessible surface area (<xref ref-type="bibr" rid="B23">de Vries and Bonvin, 2008</xref>; <xref ref-type="bibr" rid="B40">Hou et&#x20;al., 2017</xref>) and backbone flexibility (<xref ref-type="bibr" rid="B4">Bendell et&#x20;al., 2014</xref>). According to its source, features are divided into sequence-based, structure-based, and hybrid features, which are a combination of sequence and structure features (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>). The sequence-based feature is cheaper to calculate but does not contain any information from structures that might be responsible for protein functions. The structures of most proteins are not available, while structural information generally obtained by computational prediction contain noise, which sometimes heavily effected subsequent discrimination. Information from neighboring residues of interaction sites is important to determine protein-protein interaction sites. In addition, there exists binding signals far from interaction sites. <xref ref-type="bibr" rid="B99">Zeng et&#x20;al. (2020)</xref> demonstrated that inclusion of global features increased the performance of predicting protein-protein interaction sites. Both the local and the global features were obtained by non-linear degeneration. That is to say, during the transformation from proteins to features, information is lost. In addition, the local and the global features also contained noise. The deep learning-based encoder answers these issues above. Inspired by this, we used the DeepPPISP proposed by <xref ref-type="bibr" rid="B99">Zeng et&#x20;al. (2020)</xref> to refine features of protein-protein interaction sites, Extreme Gradient Boosting (XGBoost) to learn a classifier for unknown PPIS prediction.</p>
</sec>
<sec id="s2">
<title>Datasets</title>
<p>For a fair comparison with other state-of-the-art methods, we used the same three datasets as in the literature (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>). These datasets are named respectively Dset_186, Dset_72 (<xref ref-type="bibr" rid="B65">Murakami and Mizuguchi, 2010</xref>), and Dset_164 (<xref ref-type="bibr" rid="B82">Singh et&#x20;al., 2014</xref>). The procedure of collecting them is briefly described as follows. All the data originated from the PDB database (<xref ref-type="bibr" rid="B5">Berman et&#x20;al., 2000</xref>). Dset_186, Dset_72 and Dset_164 consisted of 186, 72, and 164&#x20;non-repetitive protein sequences with the resolution less than 3.0&#xa0;&#xc5;, respectively. In each dataset, sequence homology between any two sequences was less than 25%. Three datasets were integrated, containing in total 422 protein sequences. Two proteins had no definition of secondary structure of proteins (DSSP) file without which their features cannot be computed. Thus these two protein sequences were removed by <xref ref-type="bibr" rid="B99">Zeng et&#x20;al. (2020)</xref>. Finally, the remaining 420 protein sequences were&#x20;used.</p>
<p>Protein-protein interaction binding sites are determined by the absolute solvent accessibility of amino acids. If the absolute solvent accessibility was less than 1&#xa0;&#xc5;<sup>2</sup>, the amino acid was considered to be a binding site, and otherwise it was a non-interaction site. There were 5,517, 6,096, and 1,923 binding sites, as well as 30,702, 27,585, and 16,217&#x20;non-interaction sites in the Dset_186, Dset_164, and Dset_72 datasets respectively. 83.3% of the protein sequences were randomly selected as the training set and 16.7% of the protein sequences as the testing set. The training set was further divided into two parts: 90% of the training set was used for training and 10% was used for verification. Finally, 300 protein sequences were used for training (containing 65,869 amino acid residues), 50 protein sequences for verification (containing 7,319 amino acid residues), and 70 protein sequences for independent testing (containing 11,791 amino acid residues) (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>).</p>
</sec>
<sec sec-type="methods" id="s3">
<title>Methods</title>
<p>The proposed method called DeepPPISP-XGB consisted of three main steps: extracting features, training a classifier, and predicting PPIS (<xref ref-type="fig" rid="F1">Figure&#x20;1A</xref>). The DeepPPISP was a deep learning model proposed by Zeng et&#x20;al. (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>) for PPIS (<xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>). Here, we used it as an encoder of amino acid sequences, because the deep learning algorithms have a powerful ability to represent objects. We trained the DeepPPISP model with the training set. The input of the first fully connected layer in the trained DeepPPISP was used as a representation of the input. The XGBoost classifier was trained by the preprocessing features of the encoder. For unknown protein sequences which have secondary structure, raw protein sequence, and position-specific scoring matrix feature, the trained DeepPPISP extracted preprocessing features firstly and then the trained XGBoost classifier predicted&#x20;PPIS.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The architecture of DeepPPISP-XGB model. <bold>(A)</bold> Illustration of the DeepPPISP-XGB workflow, which consists of three modules: extracting feature, training classifier, predicting PPIS. <bold>(B)</bold> The architecture of DeepPPISP model, which contain embedding layer, different scale convolutions, fully connected layers and output layer.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g001.tif"/>
</fig>
<sec id="s3-1">
<title>DeepPPISP</title>
<p>As shown in <xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>, the DeepPPISP proposed by Zeng et&#x20;al. (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>) for PPIS prediction had three types of input: position-specific scoring matrix (PSSM), secondary structure, and raw protein sequences. The PSSM is an excellent feature extractor for protein sequences and thus have widely been applied to problems in the field of computational biology, such as predicting protein post-translational modification (<xref ref-type="bibr" rid="B42">Huang et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B41">Huang et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B24">Dehzangi et&#x20;al., 2017</xref>), membrane type (<xref ref-type="bibr" rid="B90">Wang et&#x20;al., 2019</xref>), protein-RNA binding site (<xref ref-type="bibr" rid="B61">Liu et&#x20;al., 2021</xref>), and structure (<xref ref-type="bibr" rid="B38">Guo et&#x20;al., 2021</xref>). The quality of PSSM features is closely associated with the underlying multiple sequence alignments. Although there are many multiple sequence alignment algorithms including HIMMER (<xref ref-type="bibr" rid="B31">Eddy, 2011</xref>; <xref ref-type="bibr" rid="B96">Wheeler and Eddy, 2013</xref>) (<xref ref-type="bibr" rid="B44">Johnson et&#x20;al., 2010</xref>) and Hhbilits (<xref ref-type="bibr" rid="B75">Remmert et&#x20;al., 2012</xref>), PSI-BLAST (<xref ref-type="bibr" rid="B1">Altschul et&#x20;al., 1997</xref>) is still a popular multiple sequence alignment and homology search algorithm. Here, PSI-BLAST was used to search NCBI&#x2019;s non-redundant (NR) sequence database with three iterations and an E-value threshold of&#x20;0.001.</p>
<p>Many protein-protein interfaces are related to secondary structures (<xref ref-type="bibr" rid="B84">Taechalertpaisarn et&#x20;al., 2019</xref>). Information about protein secondary structure is helpful to predict PPIS. The DSSP program (<xref ref-type="bibr" rid="B86">Touw et&#x20;al., 2015</xref>) was used to generate nine state secondary structures: <italic>&#x3b1;</italic>-helix, 3<sub>10</sub>- helix, &#x3c0;-helix, &#x3b2;-bridge, &#x3b2;-strand, &#x3b2;-turn, bend, loop or irregular, and no secondary structure. Therefore, each amino acid residue corresponded to a 9-dimensional vector. The primary protein sequence is valuable information and thus is essential to predict protein properties. One-hot encoding was used to encode the protein sequences. There are 20 kinds of common amino acids in the protein sequences, so each amino acid residue corresponds to a 20-dimensional 0/1 vector. The protein-protein interaction is closely associated with neighboring residues of interaction sites. The local feature of interaction sites contributes to the identification of PPIS. The sliding window method was used to collect the neighboring residues of the interaction sites. The size of the sliding window was seven. For example, if the interaction site was at position i, residues at position i-3, i-2, i-1, i, i&#x2b;1, i&#x2b;2, and i&#x2b;3 were separated. Because each residue corresponds to a 20-dimensional PSSM feature, a 9-dimensional secondary structure feature, and a 20-dimensional one-hot feature vector, a window of seven amino acid residues was encoded into a 343-dimensional vector which was called the local feature.</p>
<p>Protein-protein interaction is not only linked to the local information of interacting sites, but also to global information. <xref ref-type="bibr" rid="B99">Zeng et&#x20;al. (2020)</xref> demonstrated that the inclusion of global information improved the performance of predicting PPIS. A 500-residue peptide was used to represent the global feature of PPIS. If the number of amino acid residues in the protein sequence was less than 500, it was padded with a 0. Each peptide corresponds to a 500&#x2a;49-dimensional vector called a global feature.</p>
<p>The local and the global features were fed into the DeepPPISP (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>). The DeepPPISP was made up of one embedding layer, three different scale convolutions, two fully connected layers, and an output layer (<xref ref-type="fig" rid="F1">Figure&#x20;1B</xref>). For more detail, readers can refer to the reference (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>).</p>
<p>Both the local features or global features would contain a certain degree of noise. The dimension is large, especially for global features. The DeepPPISP was used to extract a more informative representation. The DeepPPISP was trained on the training data in a supervised manner. The local and global features were fed into the trained DeepPPISP, and the input to the first fully connected layer was the abstract representation of the raw features. Compared with the raw features, the abstract representation was of low dimension and had low&#x20;noise.</p>
</sec>
<sec id="s3-2">
<title>XGBoost Algorithm</title>
<p>The XGBoost proposed by <xref ref-type="bibr" rid="B15">Chen and Guestrin, 2016</xref> belongs to Gradient Boosting Decision Tree (GBDT) (<xref ref-type="bibr" rid="B48">Ke et&#x20;al., 2017</xref>), and both are tree boosting algorithms. Compared with traditional tree boosting, the XGBoost used a theoretically justified weighted quantile sketch for approximate learning, a novel sparsity aware algorithm for handling sparse data, and an effective cache-aware block structure for out-of-core tree learning (<xref ref-type="bibr" rid="B15">Chen and Guestrin, 2016</xref>). In addition, the XGBoost performed faster as it exploited parallel and distributed computing. The XGBoost has such a significant superiority that it has widely been used in many areas including machine learning and data mining challenges.</p>
<p>The XGBoost is an addition model. At each iteration, the XGBoost learns a new tree that fits the residual between the predicted result of the previous trees and the true values of the training samples.</p>
<p>Assume that <inline-formula id="inf1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>&#x7c;</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>&#x2208;</mml:mi>
<mml:msup>
<mml:mi>R</mml:mi>
<mml:mi>m</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>&#x2208;</mml:mi>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> denotes a training set, where m and n represented the numbers of features and samples, respectively. At the t-th iteration, the aim of the XGBoost is to learn a function <inline-formula id="inf2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> so that<disp-formula id="e1">
<mml:math id="m3">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf3">
<mml:math id="m4">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> is the fitting value of the previous t&#x2212;1 trees for the i-th sample. To search for <inline-formula id="inf4">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the loss function with the regularization was used as the objective function:<disp-formula id="e2">
<mml:math id="m6">
<mml:mrow>
<mml:mi mathvariant="bold-italic">obj</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;</mml:mi>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:munderover>
<mml:mi mathvariant="bold-italic">l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msubsup>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:munderover>
<mml:mi mathvariant="bold-italic">&#x3a9;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:munderover>
<mml:mi mathvariant="bold-italic">l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3a9;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">constant</mml:mi>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf5">
<mml:math id="m7">
<mml:mi>l</mml:mi>
</mml:math>
</inline-formula> was the loss function which was generally defined as<disp-formula id="e3">
<mml:math id="m8">
<mml:mrow>
<mml:mi mathvariant="bold-italic">l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msubsup>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
<inline-formula id="inf6">
<mml:math id="m9">
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>t</mml:mi>
</mml:munderover>
<mml:mtext>&#x3a9;</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> denotes the regularization. The loss function <inline-formula id="inf7">
<mml:math id="m10">
<mml:mi>l</mml:mi>
</mml:math>
</inline-formula> was approximated by the second-order Taylor series, namely<disp-formula id="e4">
<mml:math id="m11">
<mml:mrow>
<mml:mi mathvariant="bold-italic">l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2248;</mml:mo>
<mml:mi mathvariant="bold-italic">l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi mathvariant="bold-italic">y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mrow>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:msub>
<mml:mi mathvariant="bold-italic">h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf8">
<mml:math id="m12">
<mml:mrow>
<mml:msub>
<mml:mi>g</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2202;</mml:mo>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2202;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula id="inf9">
<mml:math id="m13">
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mo>&#x2202;</mml:mo>
<mml:mn>2</mml:mn>
</mml:msup>
<mml:mi>l</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2202;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2202;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</inline-formula> were the first- and the second-order gradients of the loss function with respect to <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>y</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mi>t</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> respectively. <inline-formula id="inf11">
<mml:math id="m15">
<mml:mrow>
<mml:mtext>&#x3a9;</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> was defined by<disp-formula id="e5">
<mml:math id="m16">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3a9;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;T</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:munderover>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where T was the number of leaf nodes and <inline-formula id="inf12">
<mml:math id="m17">
<mml:mrow>
<mml:msub>
<mml:mi>&#x3c9;</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> was the weight of the j-th leaf node. The objective function was equivalently rewritten as<disp-formula id="e6">
<mml:math id="m18">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#xa0;obj</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:msub>
<mml:mi mathvariant="bold-italic">h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">f</mml:mi>
<mml:mi mathvariant="bold-italic">t</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;T</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mi mathvariant="bold-italic">&#x3bb;</mml:mi>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:munderover>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>
</p>
<p>The set of instances of the leaf node j was defined by<disp-formula id="e7">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mtext>&#x7c;</mml:mtext>
<mml:mi mathvariant="bold-italic">q</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
</p>
<p>The objective function was further represented as<disp-formula id="e8">
<mml:math id="m20">
<mml:mrow>
<mml:mi mathvariant="bold-italic">obj</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munder>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;T</mml:mi>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>Given a fixed tree <inline-formula id="inf13">
<mml:math id="m21">
<mml:mrow>
<mml:mtext>q</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mtext>x</mml:mtext>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, the optimal value of each leaf node was calculated by<disp-formula id="e9">
<mml:math id="m22">
<mml:mrow>
<mml:msubsup>
<mml:mi mathvariant="bold-italic">&#x3c9;</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mi mathvariant="bold-italic">&#x2a;</mml:mi>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>and the optimal value of the whole tree was calculated by<disp-formula id="e10">
<mml:math id="m23">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#xa0;ob</mml:mi>
<mml:msup>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mi mathvariant="bold-italic">&#x2a;</mml:mi>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;T</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>It was expensive and impossible to exhaust all the possible trees for the training data. In practice, the greedy algorithm was used, which started from one node and iteratively split the node. Assume that before the node was split, the objective function of the tree was<disp-formula id="e11">
<mml:math id="m24">
<mml:mrow>
<mml:mi mathvariant="bold-italic">ob</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
<p>After the node k was split into the left tree<inline-formula id="inf14">
<mml:math id="m25">
<mml:mrow>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and the right tree <inline-formula id="inf15">
<mml:math id="m26">
<mml:mrow>
<mml:msub>
<mml:mi>I</mml:mi>
<mml:mi>R</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, the objective function was<disp-formula id="e12">
<mml:math id="m27">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#xa0;ob</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">T</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi mathvariant="bold-italic">&#xa0;</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2014;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
<p>The gain of node splitting was calculated by<disp-formula id="e13">
<mml:math id="m28">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#xa0;gain</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">ob</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">ob</mml:mi>
<mml:msub>
<mml:mi mathvariant="bold-italic">j</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">L</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2b;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">R</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="bold-italic">g</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:msub>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi mathvariant="bold-italic">I</mml:mi>
<mml:mi mathvariant="bold-italic">k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi mathvariant="bold-italic">i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mstyle>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#xa0;&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x3b3;</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>The gain was used to assess the split candidates.</p>
</sec>
</sec>
<sec id="s4">
<title>Evaluation Metrics</title>
<p>In the area of machine learning, the frequently used evaluation metrics include accuracy (ACC), Recall, Precision, F1-score (F1), and Matthews correlation coefficient (MCC) which are respectively calculated by the following formulas:<disp-formula id="e14">
<mml:math id="m29">
<mml:mrow>
<mml:mi mathvariant="bold-italic">ACC</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">TN</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">TN</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
<disp-formula id="e15">
<mml:math id="m30">
<mml:mrow>
<mml:mi mathvariant="bold-italic">Recall</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FN</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
<disp-formula id="e16">
<mml:math id="m31">
<mml:mrow>
<mml:mi mathvariant="bold-italic">Precision</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FP</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(16)</label>
</disp-formula>
<disp-formula id="e17">
<mml:math id="m32">
<mml:mrow>
<mml:mi mathvariant="bold-italic">F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Sensitivity</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="bold-italic">Precision</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="bold-italic">Sensitivity</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">Precision</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(17)</label>
</disp-formula>
<disp-formula id="e18">
<mml:math id="m33">
<mml:mrow>
<mml:mi mathvariant="bold-italic">MCC</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="bold-italic">TN</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="bold-italic">FP</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi mathvariant="bold-italic">FN</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FP</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">TP</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FN</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">TN</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FP</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xd7;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi mathvariant="bold-italic">TN</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">FN</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(18)</label>
</disp-formula>where TP and TN denote respectively the numbers of the true positive and the true negative samples, and FP and FN denote the numbers of the false positive and false negative samples. The F1-score ranges from 0 to 1. F1-score values close to 1 indicated the best prediction. The MCC represents the correlation coefficient between the actual classification and the predicted classification. The range of MCC values is &#x2212;1 to 1, where 1 meant perfect prediction, and &#x2212;1 indicated the worst prediction. The area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC) were also used to evaluate the performances.</p>
</sec>
<sec id="s5">
<title>Experiments</title>
<sec id="s5-1">
<title>Visualization of Preprocessing Features</title>
<p>To investigate the ability of the features to discriminate protein-protein interaction sites from non-interaction sites, we used the Uniform Manifold Approximation and Projection (UMAP) (<xref ref-type="bibr" rid="B64">McInnes et&#x20;al., 2020</xref>) to depict the first two principal components. The UMAP is a powerful tool for dimension reduction and visualization. As shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>, the features processed by the DeepPPISP demonstrated a tighter cluster than the raw features, indicating that features generated by the DeepPPISP were more discriminative. To further evaluate the performance of the preprocessed features, we performed 5-fold cross-validation and independent tests. <xref ref-type="fig" rid="F3">Figure&#x20;3A</xref> showed the ROC curves of the 5-fold cross-validation over both the preprocessing features and raw features, while <xref ref-type="fig" rid="F3">Figure&#x20;3B</xref> depicted the ROC curves of the independent tests. The performance of preprocessed features is equivalent to or better than those of raw features. It must be pointed out that the user-defined parameters were identical in the XGBoost classifiers. Comparison with other methods</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>UMAP diagrams of <bold>(A)</bold> raw features of the training set, <bold>(B)</bold> preprocessing features of the training set, <bold>(C)</bold> raw features of the testing set, and <bold>(D)</bold> preprocessing feature of the testing&#x20;set.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The ROC curves of <bold>(A)</bold> 5-fold cross validation and <bold>(B)</bold> independent test. The red dotted line is a control line on which AUROC &#x3d; 0.5.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g003.tif"/>
</fig>
<p>Due to its versatile roles in the cellular process, the identification of protein-protein interaction sites is increasingly becoming a hot topic and is also a challenging task. Over the past decades, more than 10 methods have been proposed to predict protein-protein interaction sites (<xref ref-type="bibr" rid="B70">Patel et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B30">Du et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B65">Murakami and Mizuguchi, 2010</xref>; <xref ref-type="bibr" rid="B89">Wang et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B101">Zhang and Kurgan, 2019</xref>; <xref ref-type="bibr" rid="B67">Northey et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B13">Chen et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B81">&#x160;iki&#x107; et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B34">Fiorucci and Zacharias, 2010</xref>; <xref ref-type="bibr" rid="B29">Doszt&#xe1;nyi et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B55">La and Kihara, 2012</xref>; <xref ref-type="bibr" rid="B7">Bradford and Westhead, 2005</xref>; <xref ref-type="bibr" rid="B16">Chen and Jeong, 2009</xref>; <xref ref-type="bibr" rid="B18">Chung et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B33">Fernandez-Recio et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B80">Shoemaker et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B68">Ofran and Rost, 2007</xref>; <xref ref-type="bibr" rid="B73">Qin and Zhou, 2007</xref>; <xref ref-type="bibr" rid="B60">Liang et&#x20;al., 2006</xref>; <xref ref-type="bibr" rid="B57">Li et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B104">Zhou and Shan, 2001</xref>; <xref ref-type="bibr" rid="B66">Neuvirth et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B72">Porollo and Meller, 2007</xref>; <xref ref-type="bibr" rid="B78">Segura et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B74">Qiu and Wang, 2012</xref>; <xref ref-type="bibr" rid="B95">Wei et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B105">Zhu et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B37">Guo et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B53">Kuo and Li, 2016</xref>; <xref ref-type="bibr" rid="B92">Wang et&#x20;al., 2021</xref>; <xref ref-type="bibr" rid="B63">Maheshwari and Brylinski, 2015</xref>; <xref ref-type="bibr" rid="B59">Li, 2020</xref>; <xref ref-type="bibr" rid="B28">Dick and Green, 2016</xref>; <xref ref-type="bibr" rid="B91">Wang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B103">Zhao et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B43">Jia et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B25">Deng et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B82">Singh et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B40">Hou et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B56">Li et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B91">Wang et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B3">Bagchi, 2015</xref> &#x23;412; <xref ref-type="bibr" rid="B101">Zhang and Kurgan, 2019</xref>). We compared the proposed method with six other state-of-the-art methods. These six competing methods were DeepPPISP (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>), SCRIBER (<xref ref-type="bibr" rid="B100">Zhang et&#x20;al., 2019</xref>), IntPred (<xref ref-type="bibr" rid="B67">Northey et&#x20;al., 2018</xref>), RF_PPI (<xref ref-type="bibr" rid="B40">Hou et&#x20;al., 2017</xref>), SPRINGS (<xref ref-type="bibr" rid="B82">Singh et&#x20;al., 2014</xref>), PSIVER (<xref ref-type="bibr" rid="B65">Murakami and Mizuguchi, 2010</xref>), ISIS (<xref ref-type="bibr" rid="B68">Ofran and Rost, 2007</xref>), and SPPIDER (<xref ref-type="bibr" rid="B72">Porollo and Meller, 2007</xref>). PSIVER was a Na&#xef;ve Bayes-based classifier that used features from PSSM and accessibility, while SPPIDER combined fingerprints with information from the sequences and structures for PPIS predictio. Both SPRINGS and ISIS were neural network-based methods. The former used evolutionary information, averaged cumulative hydropathy, and predicted relative solvent accessibility, while the latter used structural features and evolutionary information. RF_PPI was a random forest-based classifier for PPIS prediction, while the DeepPPISP was a deep learning-based classifier. The performances of these seven methods over the independent test were listed in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison with other state-of-the-art methods.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">ACC</th>
<th align="center">Precision</th>
<th align="center">Recall</th>
<th align="center">F1</th>
<th align="center">AUROC</th>
<th align="center">AUPRC</th>
<th align="center">MCC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">SPPIDER<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.622</td>
<td align="char" char=".">0.209</td>
<td align="char" char=".">0.459</td>
<td align="char" char=".">0.287</td>
<td align="center">&#x2014;</td>
<td align="center">0.23</td>
<td align="char" char=".">0.089</td>
</tr>
<tr>
<td align="left">ISIS<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">
<bold>0.694</bold>
</td>
<td align="char" char=".">0.211</td>
<td align="char" char=".">0.362</td>
<td align="char" char=".">0.267</td>
<td align="center">&#x2014;</td>
<td align="center">0.24</td>
<td align="char" char=".">0.097</td>
</tr>
<tr>
<td align="left">PSIVER<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.653</td>
<td align="char" char=".">0.253</td>
<td align="char" char=".">0.468</td>
<td align="char" char=".">0.328</td>
<td align="center">&#x2014;</td>
<td align="center">0.25</td>
<td align="char" char=".">0.138</td>
</tr>
<tr>
<td align="left">SPRINGS<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.631</td>
<td align="char" char=".">0.248</td>
<td align="char" char=".">
<italic>0.598</italic>
</td>
<td align="char" char=".">0.35</td>
<td align="center">&#x2014;</td>
<td align="center">0.28</td>
<td align="char" char=".">0.181</td>
</tr>
<tr>
<td align="left">RF.PPI<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.598</td>
<td align="char" char=".">0.173</td>
<td align="char" char=".">0.512</td>
<td align="char" char=".">0.258</td>
<td align="center">&#x2014;</td>
<td align="center">0.21</td>
<td align="char" char=".">0.118</td>
</tr>
<tr>
<td align="left">IntPred<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">
<italic>0.672</italic>
</td>
<td align="char" char=".">0.247</td>
<td align="char" char=".">0.508</td>
<td align="char" char=".">0.332</td>
<td align="center">&#x2014;</td>
<td align="center">&#x2014;</td>
<td align="char" char=".">0.165</td>
</tr>
<tr>
<td align="left">SCRIBER<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.616</td>
<td align="char" char=".">0.274</td>
<td align="char" char=".">0.569</td>
<td align="char" char=".">0.37</td>
<td align="center">0.635</td>
<td align="center">0.307</td>
<td align="char" char=".">0.159</td>
</tr>
<tr>
<td align="left">DeepPPISP<xref ref-type="table-fn" rid="Tfn1">
<sup>a</sup>
</xref>
</td>
<td align="char" char=".">0.655</td>
<td align="char" char=".">
<bold>0.303</bold>
</td>
<td align="char" char=".">0.577</td>
<td align="char" char=".">
<italic>0.397</italic>
</td>
<td align="center">
<italic>0.671</italic>
</td>
<td align="center">
<italic>0.32</italic>
</td>
<td align="char" char=".">
<italic>0.206</italic>
</td>
</tr>
<tr>
<td align="left">DeepPPISP-XGB</td>
<td align="char" char=".">0.633</td>
<td align="char" char=".">
<italic>0.296</italic>
</td>
<td align="char" char=".">
<bold>0.624</bold>
</td>
<td align="char" char=".">
<bold>0.402</bold>
</td>
<td align="center">
<bold>0.681</bold>
</td>
<td align="center">
<bold>0.339</bold>
</td>
<td align="char" char=".">
<bold>0.209</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="Tfn1">
<label>a</label>
<p>Results reported by DeepPPISP (<xref ref-type="bibr" rid="B99">Zeng et&#x20;al., 2020</xref>).</p>
</fn>
<fn>
<p>The highest results are highlighted in bold and the second-highest results are marked in italics. Values that were not reported by the corresponding source are indicated by &#x201c;&#x2014;&#x201d;.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The DeepPPISP-XGB method achieved the highest value in terms of Recall, F1-score, AUROC, AUPRC, and MCC, and it reached the second-highest performance in terms of Precision. Although ISIS got the best ACC, its performance in other respects was lower than those of DeepPPISP-XGB. The DeepPPISP-XGB method improved the Recall by 4.7%, 5.5%, 11.6%, 11.2%, 2.6%, 15.6%, 26.2%, and 16.5%, in comparison with DeepPPISP, SCRIBER, IntPred, RF.PPI, SPRINGS, PSIVER, ISIS, and SPPIDER, respectively. The DeepPPISP-XGB method increased F1-score and MCC by 0.5% and 0.3%, and the AUROC by 1%, in comparison with DeepPPISP.</p>
<p>K-fold cross-validation is a common method in regression or classification questions. In the k-fold cross-validation, the training set was split into k parts. One part was tested and other k&#x2212;1 parts were trained. The procedure was performed k times. We carried out 10-fold cross-validations, and the principle was shown (<xref ref-type="sec" rid="s12">Supplementary Figure S1</xref>). <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> showed ROC curves for the 10-fold cross-validations. The mean and the standard deviation of the AUROCs were 0.741 and 0.006, respectively. <xref ref-type="sec" rid="s12">Supplementary Table S1</xref> lists the ACC, Precision, Recall, F1-score, AUROC, AUPRC, and MCC for each cross-validation.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The ROC curves of 10-fold cross validation on the train set. The minimum AUROC value cross validation is 0.730 at the first fold. The maximum value of the cross validation is 0.752 at the ten-th fold. The green line represents the ROC curve of the cross validation mean. The mean value of AUROC is 0.741. The red dotted line is a control line on which AUROC &#x3d; 0.5.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g004.tif"/>
</fig>
<p>To further evaluate the predictive performance of the DeepPPISP-XGB method, four machine learning algorithms were used for PPIS prediction. Decision tree (<xref ref-type="bibr" rid="B76">Safavian and Landgrebe, 1991</xref>) is a widely utilized classification algorithm, which is made up of the root node, internal nodes, and leaf node. Random forest (RF) (<xref ref-type="bibr" rid="B9">Breiman, 2001</xref>) is an ensemble learning algorithm. It consists of many weak classifiers which determine the sample category. Extremely randomized tree (ERT) (<xref ref-type="bibr" rid="B35">Geurts et&#x20;al., 2006</xref>) is similar to RF but the decision tree of ERT is randomly divided. Support vector machine (SVM) is a statistical algorithm proposed by Boser et&#x20;al. (<xref ref-type="bibr" rid="B6">Boser et&#x20;al., 1992</xref>). These classifiers were implemented in the Scikit-Learn package (v0.24.2) which has been widely utilized in computational biology. The ROC curves and the precision-recall curves are shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. The XGBoost classifier obtained an AUROC value of 0.681 and an AUPRC value of 0.339 on the independent test, significantly better than four classifiers.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The ROC curves <bold>(A)</bold> and precision-recall curves <bold>(B)</bold> for 5 algorithms on the independent&#x20;test.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g005.tif"/>
</fig>
</sec>
<sec id="s5-2">
<title>The Effects of the Global Features</title>
<p>After removing global features, we trained DeepPPISP-XGB. The user-defined parameters of the DeepPPISP-XGB were the same as the previous. <xref ref-type="table" rid="T2">Table&#x20;2</xref> shows the performance of predicting PPIS by using local features alone. The ROC and the precision-recall curves were displayed in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>. The experimental results showed that the inclusion of the global features was beneficial to improve PPIS prediction, which was in agreement with the findings of Zeng et&#x20;al. (2020).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Predictive performance when using local features and using combined local and global features with the DeepPPISP-XGB&#x20;model.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Features</th>
<th align="center">ACC</th>
<th align="center">Precision</th>
<th align="center">Recall</th>
<th align="center">F1</th>
<th align="center">MCC</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Local features</td>
<td align="char" char=".">0.654</td>
<td align="char" char=".">0.276</td>
<td align="char" char=".">0.461</td>
<td align="char" char=".">0.345</td>
<td align="char" char=".">0.138</td>
</tr>
<tr>
<td align="left">Global &#x26; local features</td>
<td align="char" char=".">0.633</td>
<td align="char" char=".">0.296</td>
<td align="char" char=".">0.624</td>
<td align="char" char=".">0.402</td>
<td align="char" char=".">0.209</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The ROC curves <bold>(A)</bold> and the precision-recall curves <bold>(B)</bold> for both local and global &#x26; local features on the independent&#x20;test.</p>
</caption>
<graphic xlink:href="fgene-12-752732-g006.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s6">
<title>Conclusion</title>
<p>We presented a PPIS prediction algorithm based on the DeepPPISP and the XGBoost. The DeepPPISP served as a feature extractor to remove redundant information of the protein sequences. The XGBoost was used to construct a classifier for predicting PPIS. The DeepPPISP-XGB achieved competitive performances with other state-of-the-art methods.</p>
</sec>
<sec id="s7">
<title>Source Code</title>
<p>Source code is available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/fatancy2580/DeepPPISPXGB-master">https://github.com/fatancy2580/DeepPPISPXGB-master</ext-link>.</p>
</sec>
</body>
<back>
<sec id="s8">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s12">Supplementary Material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s9">
<title>Author Contributions</title>
<p>GH and Z-GY conceived a concept and methodology. PW collected data, conducted the experiments, analyzed the results, and wrote the manuscript. GZ analyzed results. GH revised the manuscript.</p>
</sec>
<sec id="s10">
<title>Funding</title>
<p>This work is supported by the National Natural Science Foundation of China (11871061), by the Natural Science Foundation of Hunan Province (2020JJ4034), by the Scientific Research Fund of Hunan Provincial Education Department (19A215), by the open project of Hunan Key Laboratory for Computation and Simulation in Science and Engineering (2019LCESE03), and by Shaoyang University Innovation Foundation for Postgraduate (CX2021SY041, CX2021SY001).</p>
</sec>
<sec sec-type="COI-statement" id="s11">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s12" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s13">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.752732/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2021.752732/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Figure5.TIF" id="SM1" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Figure6.TIF" id="SM2" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Figure1.TIF" id="SM3" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Figure4.TIF" id="SM4" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Figure3.TIF" id="SM5" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM6" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Figure2.TIF" id="SM7" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Madden</surname>
<given-names>T. L.</given-names>
</name>
<name>
<surname>Sch&#xe4;ffer</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>1997</year>). <article-title>Gapped BLAST and PSI-BLAST: a new generation of protein database search programs</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume> (<issue>17</issue>), <fpage>3389</fpage>&#x2013;<lpage>3402</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.17.3389</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aumentado-Armstrong</surname>
<given-names>T. T.</given-names>
</name>
<name>
<surname>Istrate</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Murgita</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Algorithmic approaches to protein-protein interaction site prediction</article-title>. <source>Algorithms Mol. Biol.</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1186/s13015-015-0033-9</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bagchi</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Use of Machine Learning Features to Detect Protein-Protein Interaction Sites at the Molecular Level</article-title>. <source>Inf. Syst. Des. Intell. Appl.</source>, <fpage>49</fpage>&#x2013;<lpage>54</lpage>. <comment>Springer</comment>. <pub-id pub-id-type="doi">10.1007/978-81-322-2247-7_6</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bendell</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Aumentado-Armstrong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Istrate</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Cernek</surname>
<given-names>P. T.</given-names>
</name>
<name>
<surname>Khan</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Transient protein-protein interface prediction: datasets, features, algorithms, and the RAD-T predictor</article-title>. <source>BMC bioinformatics</source> <volume>15</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-15-82</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Berman</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Westbrook</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gilliland</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bhat</surname>
<given-names>T. N.</given-names>
</name>
<name>
<surname>Weissig</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2000</year>). <article-title>The protein data bank</article-title>. <source>Nucleic Acids Res.</source> <volume>28</volume>, <fpage>235</fpage>&#x2013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.235</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boser</surname>
<given-names>B. E.</given-names>
</name>
<name>
<surname>Guyon</surname>
<given-names>I. M.</given-names>
</name>
<name>
<surname>Vapnik</surname>
<given-names>V. N.</given-names>
</name>
</person-group> (<year>1992</year>). <article-title>A training algorithm for optimal margin classifiers</article-title>. <source>Proc. fifth Annu. Workshop Comput. Learn. Theor.</source>, <fpage>144</fpage>&#x2013;<lpage>152</lpage>. <pub-id pub-id-type="doi">10.1145/130385.130401</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bradford</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Westhead</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Improved prediction of protein-protein binding sites using a support vector machines approach</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>1487</fpage>&#x2013;<lpage>1494</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti242</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bradshaw</surname>
<given-names>R. T.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>B. H.</given-names>
</name>
<name>
<surname>Tate</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Leatherbarrow</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Gould</surname>
<given-names>I. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Comparing experimental and computational alanine scanning techniques for probing a prototypical protein-protein interaction</article-title>. <source>Protein Eng. Des. Selection</source> <volume>24</volume>, <fpage>197</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="doi">10.1093/protein/gzq047</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breiman</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Random forests</article-title>. <source>Machine Learn.</source> <volume>45</volume>, <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/A:1010933404324</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Caffrey</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Somaroo</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hughes</surname>
<given-names>J.&#x20;D.</given-names>
</name>
<name>
<surname>Mintseris</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>E. S.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Are protein-protein interfaces more conserved in sequence than the rest of the protein surface</article-title>. <source>Protein Sci.</source> <volume>13</volume>, <fpage>190</fpage>&#x2013;<lpage>202</lpage>. <pub-id pub-id-type="doi">10.1110/ps.03323604</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Callaway</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>&#x27;It will change everything&#x27;: DeepMind&#x27;s AI makes gigantic leap in solving protein structures</article-title>. <source>Nature</source> <volume>588</volume>, <fpage>203</fpage>&#x2013;<lpage>204</lpage>. <pub-id pub-id-type="doi">10.1038/d41586-020-03348-4</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carl</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Konc</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jane&#x17e;i&#x10d;</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Protein surface conservation in binding sites</article-title>. <source>J.&#x20;Chem. Inf. Model.</source> <volume>48</volume>, <fpage>1279</fpage>&#x2013;<lpage>1286</lpage>. <pub-id pub-id-type="doi">10.1021/ci8000315</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>C.-T.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>H.-P.</given-names>
</name>
<name>
<surname>Jian</surname>
<given-names>J.-W.</given-names>
</name>
<name>
<surname>Tsai</surname>
<given-names>K.-C.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>E.-W.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Protein-protein interaction site predictions with three-dimensional probability distributions of interacting atoms on protein surfaces</article-title>. <source>PloS one</source> <volume>7</volume>, <fpage>e37706</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0037706</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>H.-X.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Prediction of interface residues in protein-protein complexes by a consensus neural network method: Test against NMR data</article-title>. <source>Proteins</source> <volume>61</volume>, <fpage>21</fpage>&#x2013;<lpage>35</lpage>. <pub-id pub-id-type="doi">10.1002/prot.20514</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Guestrin</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Xgboost: A scalable tree boosting system</article-title>,&#x201d; in <source>Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>ACM</publisher-name>, <fpage>785</fpage>&#x2013;<lpage>794</lpage>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>X.-w.</given-names>
</name>
<name>
<surname>Jeong</surname>
<given-names>J.&#x20;C.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Sequence-based prediction of protein interaction sites with an integrative method</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>585</fpage>&#x2013;<lpage>591</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp039</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Choi</surname>
<given-names>Y. S.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-S.</given-names>
</name>
<name>
<surname>Choi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ryu</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Evolutionary conservation in multiple faces of protein interaction</article-title>. <source>Proteins</source> <volume>77</volume>, <fpage>14</fpage>&#x2013;<lpage>25</lpage>. <pub-id pub-id-type="doi">10.1002/prot.22410</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chung</surname>
<given-names>J.-L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Bourne</surname>
<given-names>P. E.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Exploiting sequence and structure homologs to identify protein-protein binding sites</article-title>. <source>Proteins</source> <volume>62</volume>, <fpage>630</fpage>&#x2013;<lpage>640</lpage>. <pub-id pub-id-type="doi">10.1002/prot.20741</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cohen</surname>
<given-names>F. E.</given-names>
</name>
<name>
<surname>Prusiner</surname>
<given-names>S. B.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>Pathologic conformations of prion proteins</article-title>. <source>Annu. Rev. Biochem.</source> <volume>67</volume>, <fpage>793</fpage>&#x2013;<lpage>819</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.biochem.67.1.793</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Das</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chakrabarti</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Classification and prediction of protein-protein interaction interface using machine learning algorithm</article-title>. <source>Sci. Rep.</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-020-80900-2</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dayal</surname>
<given-names>P. V.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Busenlehner</surname>
<given-names>L. S.</given-names>
</name>
<name>
<surname>Ellis</surname>
<given-names>H. R.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Exposing the Alkanesulfonate Monooxygenase Protein-Protein Interaction Sites</article-title>. <source>Biochemistry</source> <volume>54</volume>, <fpage>7531</fpage>&#x2013;<lpage>7538</lpage>. <pub-id pub-id-type="doi">10.1021/acs.biochem.5b00935</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Moraes</surname>
<given-names>F. R.</given-names>
</name>
<name>
<surname>Neshich</surname>
<given-names>I. A. P.</given-names>
</name>
<name>
<surname>Mazoni</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Yano</surname>
<given-names>I. H.</given-names>
</name>
<name>
<surname>Pereira</surname>
<given-names>J.&#x20;G. C.</given-names>
</name>
<name>
<surname>Salim</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Improving predictions of protein-protein interfaces by combining amino acid-specific classifiers based on structural and physicochemical descriptors with their weighted neighbor averages</article-title>. <source>Plos one</source> <volume>9</volume>, <fpage>e87107</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0087107</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Vries</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bonvin</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>How proteins get in touch: interface prediction in the study of biomolecular complexes</article-title>. <source>Cpps</source> <volume>9</volume>, <fpage>394</fpage>&#x2013;<lpage>406</lpage>. <pub-id pub-id-type="doi">10.2174/138920308785132712</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dehzangi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>L&#xf3;pez</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lal</surname>
<given-names>S. P.</given-names>
</name>
<name>
<surname>Taherzadeh</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Michaelson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sattar</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>PSSM-suc: Accurately predicting succinylation using position specific scoring matrix into bigram for feature extraction</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>425</volume>, <fpage>97</fpage>&#x2013;<lpage>102</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2017.05.005</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fan</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Developing computational model to predict protein-protein interaction sites based on the XGBoost algorithm</article-title>. <source>Ijms</source> <volume>21</volume>, <fpage>2274</fpage>. <pub-id pub-id-type="doi">10.3390/ijms21072274</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Prediction of protein-protein interaction sites using an ensemble method</article-title>. <source>BMC bioinformatics</source> <volume>10</volume>, <fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-10-426</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dias</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kolaczkowski</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Improving the accuracy of high-throughput protein-protein affinity prediction may require better training data</article-title>. <source>BMC bioinformatics</source> <volume>18</volume>, <fpage>7</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1533-z</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dick</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Comparison of sequence-and structure-based protein-protein interaction sites</article-title>. <source>IEEE EMBS Int. Student Conf. (Isc)</source>, <fpage>1</fpage>&#x2013;<lpage>4</lpage>. <comment>IEEE</comment>. <pub-id pub-id-type="doi">10.1109/embsisc.2016.7508605</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doszt&#xe1;nyi</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>M&#xe9;sz&#xe1;ros</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Simon</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>ANCHOR: web server for predicting protein binding regions in disordered proteins</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2745</fpage>&#x2013;<lpage>2746</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp518</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Improved prediction of protein binding sites from sequences using genetic algorithm</article-title>. <source>Protein J.</source> <volume>28</volume>, <fpage>273</fpage>&#x2013;<lpage>280</lpage>. <pub-id pub-id-type="doi">10.1007/s10930-009-9192-1</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Accelerated profile HMM searches</article-title>. <source>Plos Comput. Biol.</source> <volume>7</volume>, <fpage>e1002195</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002195</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Engelen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Trojan</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Sacquin-Mora</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lavery</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Carbone</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Joint evolutionary trees: a large-scale method to predict protein interfaces based on sequence sampling</article-title>. <source>Plos Comput. Biol.</source> <volume>5</volume>, <fpage>e1000267</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000267</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fern&#xe1;ndez-Recio</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Totrov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Abagyan</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Identification of Protein-Protein Interaction Sites from Docking Energy Landscapes</article-title>. <source>J.&#x20;Mol. Biol.</source> <volume>335</volume>, <fpage>843</fpage>&#x2013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmb.2003.10.069</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fiorucci</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zacharias</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Prediction of protein-protein interaction sites using electrostatic desolvation profiles</article-title>. <source>Biophysical J.</source> <volume>98</volume>, <fpage>1921</fpage>&#x2013;<lpage>1930</lpage>. <pub-id pub-id-type="doi">10.1016/j.bpj.2009.12.4332</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Geurts</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Ernst</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wehenkel</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Extremely randomized trees</article-title>. <source>Mach. Learn.</source> <volume>63</volume>, <fpage>3</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-006-6226-1</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guharoy</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Chakrabarti</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Secondary structure based analysis and classification of biological interfaces: identification of binding motifs in protein-protein interactions</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>1909</fpage>&#x2013;<lpage>1918</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm274</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Predicting protein-protein interaction sites using modified support vector machine</article-title>. <source>Int. J.&#x20;Mach. Learn. Cyber.</source> <volume>9</volume>, <fpage>393</fpage>&#x2013;<lpage>398</lpage>. <pub-id pub-id-type="doi">10.1007/s13042-015-0450-6</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>EPTool: A New Enhancing PSSM Tool for Protein Secondary Structure Prediction</article-title>. <source>J.&#x20;Comput. Biol.</source> <volume>28</volume>, <fpage>362</fpage>&#x2013;<lpage>364</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2020.0417</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ren</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>Deep residual learning for image recognition</article-title>,&#x201d; in <source>Proceedings of the IEEE conference on computer vision and pattern recognition</source>, <fpage>770</fpage>&#x2013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1109/cvpr.2016.90</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hou</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>De Geest</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Vranken</surname>
<given-names>W. F.</given-names>
</name>
<name>
<surname>Heringa</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Feenstra</surname>
<given-names>K. A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Seeing the Trees through the Forest: Sequence-based Homo- and Heteromeric Protein-protein Interaction sites prediction using Random Forest</article-title>. <source>Bioinformatics</source> <volume>33</volume>, <fpage>btx005</fpage>&#x2013;<lpage>1487</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx005</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Prediction of S-nitrosylation modification sites based on kernel sparse representation classification and mRMR algorithm</article-title>. <source>Biomed. Research International</source> <volume>2014</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1155/2014/438341</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.-Q.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>Y.-D.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Prediction of carbamylated lysine sites based on the one-class k-nearest neighbor method</article-title>. <source>Mol. Biosyst.</source> <volume>9</volume>, <fpage>2729</fpage>&#x2013;<lpage>2740</lpage>. <pub-id pub-id-type="doi">10.1039/c3mb70195f</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jia</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xiao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Chou</surname>
<given-names>K.-C.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>iPPBS-Opt: a sequence-based ensemble classifier for identifying protein-protein binding sites by optimizing imbalanced training datasets</article-title>. <source>Molecules</source> <volume>21</volume>, <fpage>95</fpage>. <pub-id pub-id-type="doi">10.3390/molecules21010095</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>L. S.</given-names>
</name>
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Portugaly</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Hidden Markov model speed heuristic and iterative HMM search procedure</article-title>. <source>BMC bioinformatics</source> <volume>11</volume>, <fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-11-431</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Thornton</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Analysis of protein-protein interaction sites using surface patches 1&#x20;1Edited by G.Von Heijne</article-title>. <source>J.&#x20;Mol. Biol.</source> <volume>272</volume>, <fpage>121</fpage>&#x2013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1997.1234</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Thornton</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>Prediction of protein-protein interaction sites using patch analysis 1&#x20;1Edited by G. von Heijne</article-title>. <source>J.&#x20;Mol. Biol.</source> <volume>272</volume>, <fpage>133</fpage>&#x2013;<lpage>143</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1997.1233</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jordan</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>el-Manzalawy</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Dobbs</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Honavar</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Predicting protein-protein interface residues using local surface structural similarity</article-title>. <source>BMC bioinformatics</source> <volume>13</volume>, <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-13-41</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ke</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Meng</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Finley</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Lightgbm: A highly efficient gradient boosting decision tree</article-title>. <source>Adv. Neural Inf. Process. Syst.</source> <volume>30</volume>, <fpage>3146</fpage>&#x2013;<lpage>3154</lpage>. </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kerrien</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Alam-Faruque</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Aranda</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Bancarz</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Bridge</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Derow</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>IntAct--open source resource for molecular interaction data</article-title>. <source>Nucleic Acids Res.</source> <volume>35</volume>, <fpage>D561</fpage>&#x2013;<lpage>D565</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkl958</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Keshava Prasad</surname>
<given-names>T. S.</given-names>
</name>
<name>
<surname>Goel</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kandasamy</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Keerthikumar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kumar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mathivanan</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Human Protein Reference Database--2009 update</article-title>. <source>Nucleic Acids Res.</source> <volume>37</volume>, <fpage>D767</fpage>&#x2013;<lpage>D772</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn892</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Krizhevsky</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sutskever</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hinton</surname>
<given-names>G. E.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Imagenet classification with deep convolutional neural networks</article-title>. <source>Commun. ACM</source> <volume>60</volume>, <fpage>84</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1145/3065386</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kr&#xef;&#xbf;&#xbd;ger</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Gohlke</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>DrugScorePPI webserver: fast and accurate in&#x20;silico alanine scanning for scoring protein-protein interactions</article-title>. <source>Nucleic Acids Res.</source> <volume>38</volume>, <fpage>W480</fpage>&#x2013;<lpage>W486</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkq471</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuo</surname>
<given-names>T.-H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>K.-B.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Predicting Protein-Protein Interaction Sites Using Sequence Descriptors and Site Propensity of Neighboring Amino Acids</article-title>. <source>Ijms</source> <volume>17</volume>, <fpage>1788</fpage>. <pub-id pub-id-type="doi">10.3390/ijms17111788</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuzmanov</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Emili</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Protein-protein interaction networks: probing disease mechanisms using model systems</article-title>. <source>Genome Med.</source> <volume>5</volume>, <fpage>37</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1186/gm441</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>La</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>A novel method for protein-protein interaction site prediction using phylogenetic substitution models</article-title>. <source>Proteins</source> <volume>80</volume>, <fpage>126</fpage>&#x2013;<lpage>141</lpage>. <pub-id pub-id-type="doi">10.1002/prot.23169</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>B.-Q.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>K.-Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>Y.-D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Prediction of Protein-Protein Interaction Sites by Random Forest Algorithm with mRMR and IFS</article-title>. <source>PLoS ONE</source> <volume>7</volume>, <fpage>e43927</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0043927</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.-H.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.-L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Protein protein interaction site prediction based on conditional random fields</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>597</fpage>&#x2013;<lpage>604</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl660</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>F.-X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Control principles for complex biological networks</article-title>. <source>Brief. Bioinformatics</source> <volume>20</volume>, <fpage>2253</fpage>&#x2013;<lpage>2266</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bby088</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Computational Methods for Predicting Protein-protein Interactions and Binding Sites</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Western University</publisher-name>. </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Protein binding site prediction using an empirical scoring function</article-title>. <source>Nucleic Acids Res.</source> <volume>34</volume>, <fpage>3698</fpage>&#x2013;<lpage>3707</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkl454</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>SNB&#x2010;PSSM : A spatial neighbor&#x2010;based PSSM used for protein-RNA binding site prediction</article-title>. <source>J.&#x20;Mol. Recognit</source> <volume>34</volume>, <fpage>e2887</fpage>. <pub-id pub-id-type="doi">10.1002/jmr.2887</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Loregian</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Marsden</surname>
<given-names>H. S.</given-names>
</name>
<name>
<surname>Pal&#xf9;</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Protein-protein interactions as targets for antiviral chemotherapy</article-title>. <source>Rev. Med. Virol.</source> <volume>12</volume>, <fpage>239</fpage>&#x2013;<lpage>262</lpage>. <pub-id pub-id-type="doi">10.1002/rmv.356</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maheshwari</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Brylinski</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Prediction of protein-protein interaction sites from weakly homologous template structures using meta-threading and machine learning</article-title>. <source>J.&#x20;Mol. Recognit.</source> <volume>28</volume>, <fpage>35</fpage>&#x2013;<lpage>48</lpage>. <pub-id pub-id-type="doi">10.1002/jmr.2410</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>McInnes</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Healy</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Melville</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <source>UMAP: uniform manifold approximation and projection for dimension reduction</source>. <comment>ArXiv [Preprint]. arXiv:1802.03426</comment>. </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Murakami</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mizuguchi</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Applying the Na&#xef;ve Bayes classifier with kernel density estimation to the prediction of protein-protein interaction sites</article-title>. <source>Bioinformatics</source> <volume>26</volume>, <fpage>1841</fpage>&#x2013;<lpage>1848</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq302</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Neuvirth</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Raz</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Schreiber</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>ProMate: A Structure Based Prediction Program to Identify the Location of Protein-Protein Binding Sites</article-title>. <source>J.&#x20;Mol. Biol.</source> <volume>338</volume>, <fpage>181</fpage>&#x2013;<lpage>199</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmb.2004.02.040</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Northey</surname>
<given-names>T. C.</given-names>
</name>
<name>
<surname>Bare&#x161;i&#x107;</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>A. C. R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>IntPred: a structure-based predictor of protein-protein interaction sites</article-title>. <source>Bioinformatics</source> <volume>34</volume>, <fpage>223</fpage>&#x2013;<lpage>229</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx585</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ofran</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Rost</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>ISIS: interaction sites identified from sequence</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>e13</fpage>&#x2013;<lpage>e16</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl303</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Orii</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Ganapathiraju</surname>
<given-names>M. K.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Wiki-pi: a web-server of annotated human protein-protein interactions to aid in discovery of protein function</article-title>. <source>PloS one</source> <volume>7</volume>, <fpage>e49029</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0049029</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Patel</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Pillay</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jawa</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2006</year>) <article-title>Information of binding sites improves prediction of protein-protein interaction</article-title>. In <comment>2006 5th International Conference on Machine Learning and Applications (</comment>
<source>ICMLA</source>06) pp. <fpage>205</fpage>&#x2013;<lpage>212</lpage>. <comment>IEEE</comment>. <pub-id pub-id-type="doi">10.1109/icmla.2006.29</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Petta</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Lievens</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Libert</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Tavernier</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>De Bosscher</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Modulation of Protein-Protein Interactions for the Development of Novel Therapeutics</article-title>. <source>Mol. Ther.</source> <volume>24</volume>, <fpage>707</fpage>&#x2013;<lpage>718</lpage>. <pub-id pub-id-type="doi">10.1038/mt.2015.214</pub-id> </citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Porollo</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Meller</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Prediction-based fingerprints of protein-protein interactions</article-title>. <source>Proteins</source> <volume>66</volume>, <fpage>630</fpage>&#x2013;<lpage>645</lpage>. <pub-id pub-id-type="doi">10.1002/prot.21248</pub-id> </citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>H.-X.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>meta-PPISP: a meta web server for protein-protein interaction site prediction</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>3386</fpage>&#x2013;<lpage>3387</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm434</pub-id> </citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Qiu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Prediction of protein-protein interaction sites using patch-based residue characterization</article-title>. <source>J.&#x20;Theor. Biol.</source> <volume>293</volume>, <fpage>143</fpage>&#x2013;<lpage>150</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2011.10.021W</pub-id> </citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Remmert</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Biegert</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hauser</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>S&#xf6;ding</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>HHblits: lightning-fast iterative protein sequence searching by HMM-HMM alignment</article-title>. <source>Nat. Methods</source> <volume>9</volume>, <fpage>173</fpage>&#x2013;<lpage>175</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.1818</pub-id> </citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Safavian</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Landgrebe</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>1991</year>). <article-title>A survey of decision tree classifier methodology</article-title>. <source>IEEE Trans. Syst. Man. Cybern.</source> <volume>21</volume>, <fpage>660</fpage>&#x2013;<lpage>674</lpage>. <pub-id pub-id-type="doi">10.1109/21.97458</pub-id> </citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Salwinski</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Pettit</surname>
<given-names>F. K.</given-names>
</name>
<name>
<surname>Bowie</surname>
<given-names>J.&#x20;U.</given-names>
</name>
<name>
<surname>Eisenberg</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>The database of interacting proteins: 2004 update</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>449D</fpage>&#x2013;<lpage>451D</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh086</pub-id> </citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Segura</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>P. F.</given-names>
</name>
<name>
<surname>Fernandez-Fuentes</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Improving the prediction of protein binding sites by combining heterogeneous data and Voronoi diagrams</article-title>. <source>BMC bioinformatics</source> <volume>12</volume>, <fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-12-352</pub-id> </citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Selkoe</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>1998</year>). <article-title>The cell biology of &#x3b2;-amyloid precursor protein and presenilin in Alzheimer&#x27;s disease</article-title>. <source>Trends Cell Biology</source> <volume>8</volume>, <fpage>447</fpage>&#x2013;<lpage>453</lpage>. <pub-id pub-id-type="doi">10.1016/s0962-8924(98)01363-4</pub-id> </citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shoemaker</surname>
<given-names>B. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Thangudu</surname>
<given-names>R. R.</given-names>
</name>
<name>
<surname>Tyagi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Fong</surname>
<given-names>J.&#x20;H.</given-names>
</name>
<name>
<surname>Marchler-Bauer</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Inferred Biomolecular Interaction Server-a web server to analyze and predict protein interacting partners and binding sites</article-title>. <source>Nucleic Acids Res.</source> <volume>38</volume>, <fpage>D518</fpage>&#x2013;<lpage>D524</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkp842</pub-id> </citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>&#x160;iki&#x107;</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tomi&#x107;</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vlahovi&#x10d;ek</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Prediction of Protein-Protein Interaction Sites in Sequences and 3D Structures by Random Forests</article-title>. <source>Plos Comput. Biol.</source> <volume>5</volume>, <fpage>e1000278</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1000278</pub-id> </citation>
</ref>
<ref id="B82">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Dhole</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Pai</surname>
<given-names>P. P.</given-names>
</name>
<name>
<surname>Mondal</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>SPRINGS: prediction of protein-protein interaction sites using artificial neural networks</article-title>. <source>PeerJ&#x20;PrePrints</source>. <pub-id pub-id-type="doi">10.13188/2572-8679.1000001</pub-id> </citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sperandio</surname>
<given-names>O.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Editorial: [Hot Topics: Toward the Design of Drugs on Protein-Protein Interactions]</article-title>. <source>Cpd</source> <volume>18</volume>, <fpage>4585</fpage>. <pub-id pub-id-type="doi">10.2174/138161212802651661</pub-id> </citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Taechalertpaisarn</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lyu</surname>
<given-names>R.-L.</given-names>
</name>
<name>
<surname>Arancillo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>C.-M.</given-names>
</name>
<name>
<surname>Perez</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Ioerger</surname>
<given-names>T. R.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Correlations between secondary structure- and protein-protein interface-mimicry: the interface mimicry hypothesis</article-title>. <source>Org. Biomol. Chem.</source> <volume>17</volume>, <fpage>3267</fpage>&#x2013;<lpage>3274</lpage>. <pub-id pub-id-type="doi">10.1039/c9ob00204a</pub-id> </citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tjong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Qin</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>H.-X.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>PI2PE: protein interface/interior prediction engine</article-title>. <source>Nucleic Acids Res.</source> <volume>35</volume>, <fpage>W357</fpage>&#x2013;<lpage>W362</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkm231</pub-id> </citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Touw</surname>
<given-names>W. G.</given-names>
</name>
<name>
<surname>Baakman</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Black</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>te&#xa0;Beek</surname>
<given-names>T. A. H.</given-names>
</name>
<name>
<surname>Krieger</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Joosten</surname>
<given-names>R. P.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>A series of PDB-related databanks for everyday needs</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>D364</fpage>&#x2013;<lpage>D368</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku1028</pub-id> </citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Von Mering</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Jensen</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Snel</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hooper</surname>
<given-names>S. D.</given-names>
</name>
<name>
<surname>Krupp</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Foglierini</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2004</year>). <article-title>STRING: known and predicted protein-protein associations, integrated and transferred across organisms</article-title>. <source>Nucleic Acids Res.</source> <volume>33</volume>, <fpage>D433</fpage>&#x2013;<lpage>D437</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki005</pub-id> </citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Mei</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>M.-T.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>C.-H.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Imbalance data processing strategy for protein interaction sites prediction</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>18</volume>, <fpage>985</fpage>&#x2013;<lpage>994</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2019.2953908</pub-id> </citation>
</ref>
<ref id="B89">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yan</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Fast prediction of protein-protein interaction sites based on Extreme Learning Machines</article-title>. <source>Neurocomputing</source> <volume>128</volume>, <fpage>258</fpage>&#x2013;<lpage>266</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2012.12.062</pub-id> </citation>
</ref>
<ref id="B90">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Fei</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Efficient utilization on PSSM combining with recurrent neural network for membrane protein types prediction</article-title>. <source>Comput. Biol. Chem.</source> <volume>81</volume>, <fpage>9</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2019.107094</pub-id> </citation>
</ref>
<ref id="B91">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Protein-protein interaction sites prediction by ensemble random forests with synthetic minority oversampling technique</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>2395</fpage>&#x2013;<lpage>2402</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty995</pub-id> </citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Salhi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Prediction of protein-protein interaction sites through eXtreme gradient boosting with kernel principal component analysis</article-title>. <source>Comput. Biol. Med.</source> <volume>134</volume>, <fpage>104516</fpage>. <pub-id pub-id-type="doi">10.1016/j.compbiomed.2021.104516</pub-id> </citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Mei</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhen</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Semi-supervised prediction of protein interaction sites from unlabeled sample information</article-title>. <source>BMC bioinformatics</source> <volume>20</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1186/s12859-019-3274-7</pub-id> </citation>
</ref>
<ref id="B94">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Dai</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Using Recursive Feature Selection with Random Forest to Improve Protein Structural Class Prediction for Low-Similarity Sequences</article-title>. <source>Comput. Math. Methods Med.</source>, <fpage>2021</fpage>. <pub-id pub-id-type="doi">10.1155/2021/5529389</pub-id> </citation>
</ref>
<ref id="B95">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wei</surname>
<given-names>Z.-S.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>J.-Y.</given-names>
</name>
<name>
<surname>Shen</surname>
<given-names>H.-B.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>D.-J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Protein-protein interaction sites prediction by ensembling SVM and sample-weighted random forests</article-title>. <source>Neurocomputing</source> <volume>193</volume>, <fpage>201</fpage>&#x2013;<lpage>212</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2016.02.022</pub-id> </citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wheeler</surname>
<given-names>T. J.</given-names>
</name>
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>nhmmer: DNA homology search with profile HMMs</article-title>. <source>Bioinformatics</source> <volume>29</volume>, <fpage>2487</fpage>&#x2013;<lpage>2489</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt403</pub-id> </citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xue</surname>
<given-names>L. C.</given-names>
</name>
<name>
<surname>Dobbs</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Honavar</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>HomPPI: a class of sequence homology based protein-protein interface prediction methods</article-title>. <source>BMC bioinformatics</source> <volume>12</volume>, <fpage>1</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-12-244</pub-id> </citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zellner</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Staudigel</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Trenner</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Bittkowski</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wolowski</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Icking</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Prescont: Predicting protein-protein interfaces utilizing four residue properties</article-title>. <source>Proteins</source> <volume>80</volume>, <fpage>154</fpage>&#x2013;<lpage>168</lpage>. <pub-id pub-id-type="doi">10.1002/prot.23172</pub-id> </citation>
</ref>
<ref id="B99">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeng</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>F.-X.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Protein-protein interaction site prediction through combining local and global features with deep neural networks</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>1114</fpage>&#x2013;<lpage>1120</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz699</pub-id> </citation>
</ref>
<ref id="B100">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Quan</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>L&#xfc;</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Sequence-based prediction of protein-protein interaction sites by simplified long short-term memory network</article-title>. <source>Neurocomputing</source> <volume>357</volume>, <fpage>86</fpage>&#x2013;<lpage>100</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2019.05.013</pub-id> </citation>
</ref>
<ref id="B101">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kurgan</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>SCRIBER: accurate and partner type-specific prediction of protein-binding residues from proteins sequences</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>i343</fpage>&#x2013;<lpage>i353</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz324</pub-id> </citation>
</ref>
<ref id="B102">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>Q. C.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Fisher</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Honig</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Petrey</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>PredUs: a web server for predicting protein interfaces using structural neighbors</article-title>. <source>Nucleic Acids Res.</source> <volume>39</volume>, <fpage>W283</fpage>&#x2013;<lpage>W287</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr311</pub-id> </citation>
</ref>
<ref id="B103">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Bao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yin</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>PPIs Meta: A Meta-predictor of Protein-Protein Interaction Sites with Weighted Voting Strategy</article-title>. <source>Cp</source> <volume>14</volume>, <fpage>186</fpage>&#x2013;<lpage>193</lpage>. <pub-id pub-id-type="doi">10.2174/1570164614666170306164127</pub-id> </citation>
</ref>
<ref id="B104">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>H.-X.</given-names>
</name>
<name>
<surname>Shan</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Prediction of protein interaction sites from sequence profile and residue neighbor list</article-title>. <source>Proteins</source> <volume>44</volume>, <fpage>336</fpage>&#x2013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1002/prot.1099</pub-id> </citation>
</ref>
<ref id="B105">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>ConvsPPIS: identifying protein-protein interaction sites by an ensemble convolutional neural network with feature graph</article-title>. <source>Cbio</source> <volume>15</volume>, <fpage>368</fpage>&#x2013;<lpage>378</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666191105155713</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>