<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2021.735329</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Novel Network-Based Algorithm for Predicting Protein-Protein Interactions Using Gene Ontology</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Hu</surname> <given-names>Lun</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/945401/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Wang</surname> <given-names>Xiaojuan</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Huang</surname> <given-names>Yu-An</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1001316/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Hu</surname> <given-names>Pengwei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/712365/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>You</surname> <given-names>Zhu-Hong</given-names></name>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x0002A;</sup></xref>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Sciences</institution>, <addr-line>&#x000DC;r&#x000FC;mqi</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>School of Computer Science and Technology, Wuhan University of Technology</institution>, <addr-line>Wuhan</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>College of Computer Science and Software Engineering, Shenzhen University</institution>, <addr-line>Shenzhen</addr-line>, <country>China</country></aff>
<aff id="aff4"><sup>4</sup><institution>School of Computer Science, Northwestern Polytechnical University</institution>, <addr-line>Xi&#x00027;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Qi Zhao, University of Science and Technology Liaoning, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Yi Xiong, Shanghai Jiao Tong University, China; Xiaofei Zhang, Central China Normal University, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Pengwei Hu <email>hupengwei&#x00040;hotmail.com</email></corresp>
<corresp id="c002">Zhu-Hong You <email>zhuhongyou&#x00040;nwpu.edu.cn</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Systems Microbiology, a section of the journal Frontiers in Microbiology</p></fn></author-notes>
<pub-date pub-type="epub">
<day>25</day>
<month>08</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>735329</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2021 Hu, Wang, Huang, Hu and You.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Hu, Wang, Huang, Hu and You</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract><p>Proteins are one of most significant components in living organism, and their main role in cells is to undertake various physiological functions by interacting with each other. Thus, the prediction of protein-protein interactions (PPIs) is crucial for understanding the molecular basis of biological processes, such as chronic infections. Given the fact that laboratory-based experiments are normally time-consuming and labor-intensive, computational prediction algorithms have become popular at present. However, few of them could simultaneously consider both the structural information of PPI networks and the biological information of proteins for an improved accuracy. To do so, we assume that the prior information of functional modules is known in advance and then simulate the generative process of a PPI network associated with the biological information of proteins, i.e., Gene Ontology, by using an established Bayesian model. In order to indicate to what extent two proteins are likely to interact with each other, we propose a novel scoring function by combining the membership distributions of proteins with network paths. Experimental results show that our algorithm has a promising performance in terms of several independent metrics when compared with state-of-the-art prediction algorithms, and also reveal that the consideration of modularity in PPI networks provides us an alternative, yet much more flexible, way to accurately predict PPIs.</p></abstract>
<kwd-group>
<kwd>protein-protein interaction</kwd>
<kwd>prediction</kwd>
<kwd>network topology</kwd>
<kwd>gene ontology</kwd>
<kwd>modularity</kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="8"/>
<ref-count count="26"/>
<page-count count="11"/>
<word-count count="6154"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>As one of the most common and indispensable molecules in cells, proteins are critical in regulating various biological processes observed in living organisms by interacting with other different proteins through protein-protein interactions (PPIs) (Hu et al., <xref ref-type="bibr" rid="B9">2021a</xref>). Since PPIs are of great significance to undertake many physiological functions, there is a necessity for us to identify PPIs from cells in order to fully explore the cellular mechanism behind biological processes.</p>
<p>In the last decades, a large number of prediction methods have been developed to verify the interacting relationship between pairwise proteins, and they are divided into two categories, one is laboratory-based and the other is computational-based. The technologies in the former category include, but not limited to, yeast two-hybrid (Fields and Sternglanz, <xref ref-type="bibr" rid="B2">1994</xref>), TAP-tagging (Ho et al., <xref ref-type="bibr" rid="B4">2002</xref>), and protein chips (Zhu et al., <xref ref-type="bibr" rid="B26">2001</xref>). They normally suffer the disadvantage of being time-consuming and labor-intensive, thus resulting in an inefficient identification of PPIs. To overcome these problems, attempts have been made to develop different computational algorithms for PPI prediction. In particular, computational algorithms mainly put their efforts on extracting useful features from the biological information of proteins, such as protein sequences (Zahiri et al., <xref ref-type="bibr" rid="B23">2013</xref>; Hu and Chan, <xref ref-type="bibr" rid="B6">2015</xref>), protein structures (Zhang et al., <xref ref-type="bibr" rid="B25">2012</xref>; Mirabello and Wallner, <xref ref-type="bibr" rid="B17">2017</xref>), and co-evolutionary profiles (Hsin Liu et al., <xref ref-type="bibr" rid="B5">2013</xref>; Hu and Chan, <xref ref-type="bibr" rid="B7">2017</xref>), that are able to explicitly represent the characteristics of proteins, and then solve the problem of PPI prediction as a binary classification problem. Though efficient, most of them are unable to handle the structural information of PPI networks for better performing the prediction task. Moreover, regarding the fact that the amount of PPI data have also increased significantly with the development of high-throughput technologies, studies have been conducted to develop various prediction algorithms that are able to complete the task of PPI prediction in a distributed manner (You et al., <xref ref-type="bibr" rid="B22">2014</xref>; Hu et al., <xref ref-type="bibr" rid="B10">2017</xref>).</p>
<p>As a recent attempt in network-based PPI prediction, L3 (Kov&#x000E1;cs et al., <xref ref-type="bibr" rid="B12">2019</xref>) reckons that the traditional triadic closure principle is inappropriate for predicting PPIs from a given PPI network, as two proteins are more likely to interact if one of them is similar to the other&#x00027;s partners rather than sharing many common interacting partners. Experimental results demonstrate that L3 significantly outperforms existing link prediction methods when applied to solve the PPI prediction problem. Given two proteins, since L3 only considers their common interacting partners, the network paths involved are with the same length, i.e., 3. In this regard, L3 is incapable of determining the interaction between proteins that are far away from each other without any common neighbors. To address this problem, Wang et al. (<xref ref-type="bibr" rid="B21">2020</xref>) design a novel stochastic block model, namely PPISB, for predicting PPIs without specifying the length of network paths in advance. PPISB can capture the latent structural features of proteins in a PPI network, thus verifying whether two proteins interact with each other or not. However, a major concern for network-based algorithms is the quality of PPI networks. In particular, when composing a PPI network, the PPI data generated by high-throughput technology is characterized by high false-positive and false-negative rates, and accordingly the accuracy performance of network-based prediction algorithms is severely affected. Similar to L3 and PPISB, network-based distance Analysis can also be applied to predict lncRNA-miRNA Interactions (Zhang et al., <xref ref-type="bibr" rid="B24">2021</xref>).</p>
<p>As has been pointed out by Hu et al. (<xref ref-type="bibr" rid="B11">2021b</xref>), proteins in the functional modules are densely connected with each other. In other words, for two proteins in the same functional module, their probability of being interacting should be considerably larger than those across different functional modules. Moreover, the neighboring relationship between molecules has also been verified to be useful for predicting their interactions (Liu et al., <xref ref-type="bibr" rid="B15">2020</xref>). Hence, we believe that the performance of PPI prediction can be further improved by taking into account this motivation. In this work, we target to integrate the biological information of proteins, specifically Gene Ontology (GO), into a given PPI network, thus alleviating the negative influence of noise data. Motivated by the aforementioned intuition that proteins in the same functional module are more likely to interact with each other, we adopt an established Bayesian model proposed by Hu et al. (<xref ref-type="bibr" rid="B8">2020</xref>) to simulate the generative process of PPI networks together with associated GO information by assuming that the prior information of functional modules are known in advance. After that, a novel scoring function is designed to compute the interaction probability of two proteins according to their membership distributions and network paths. Following this pipeline, we develop a new algorithm, namely NGPM, to complete the task of PPI prediction. To evaluate the performance of NGPM, a series of extensive experiments have been conducted by comparing it with several state-of-the-art PPI prediction algorithms on five practical PPI networks collected from different species, and an in-depth discussion about experimental results is provided to demonstrate the superiority of NGPM in predicting PPIs.</p>
<p>The rest of this paper is organized as follows. In section 2, the details of NGPM are described. Experimental results are presented in section 3, following which we end with an in-depth discussion in section 4.</p></sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<p>Given the fact that proteins interact with each other in cells to form functional modules, a single protein is possible to be involved in multiple protein complexes and thereby undertake different physiological functions. For a PPI network associated with GO information of proteins, we first assume that a total of <italic>K</italic> functional modules are existed and the details of generating such a PPI network is first presented by adopting the Bayesian model proposed by Hu et al. (<xref ref-type="bibr" rid="B8">2020</xref>). After that, we describe the complete procedure of NGPM.</p>
<sec>
<title>2.1. Mathematical Preliminaries</title>
<p>A PPI network of interest is formally denoted as a four-element tuple <italic>G</italic> &#x0003D; {<italic>V, A, X</italic>, &#x0039B;}, where <italic>V</italic> &#x0003D; {<italic>v</italic><sub><italic>i</italic></sub>}(1 &#x02264; <italic>i</italic> &#x02264; <italic>n</italic><sub><italic>V</italic></sub>) is a set of all <italic>n</italic><sub><italic>V</italic></sub> proteins, <italic>A</italic> &#x0003D; [<italic>A</italic><sub><italic>ij</italic></sub>] is a <italic>n</italic><sub><italic>V</italic></sub>&#x000D7;<italic>n</italic><sub><italic>V</italic></sub> adjacency matrix where <italic>A</italic><sub><italic>i</italic></sub><italic>j</italic> &#x0003D; 1 if two proteins, i.e., <italic>v</italic><sub><italic>i</italic></sub> and <italic>v</italic><sub><italic>j</italic></sub>, interact with each other and 0 otherwise, <italic>X</italic> &#x0003D; {<italic>X</italic><sub><italic>i</italic></sub>}(1 &#x02264; <italic>i</italic> &#x02264; <italic>n</italic><sub><italic>V</italic></sub>) consists of the GO information of proteins in <italic>V</italic>, and &#x0039B; &#x0003D; {&#x0039B;<sub><italic>m</italic></sub>}(1 &#x02264; <italic>m</italic> &#x02264; <italic>n</italic><sub><italic>V</italic></sub>) denotes a set of total <italic>n</italic><sub>&#x0039B;</sub> GO categories that are available to be associated with proteins. Obviously, <italic>A</italic> and <italic>X</italic> describe <italic>G</italic> from the perspectives of network topology and GO, respectively. In this regard, an instance of <italic>G</italic> can thus be obtained if <italic>A</italic> and <italic>X</italic> are determined.</p>
<p>Regarding <italic>X</italic>, each element, i.e., <italic>X</italic><sub><italic>i</italic></sub> &#x0003D; {<italic>x</italic><sub><italic>ip</italic></sub>}, denotes the set of GO annotations taken by <italic>v</italic><sub><italic>i</italic></sub> without considering GO categories, and the size of <italic>X</italic><sub><italic>i</italic></sub> is |<italic>X</italic><sub><italic>i</italic></sub>|. The combination of <italic>X</italic><sub><italic>i</italic></sub> and &#x0039B; preserves the necessary details to sample the GO information for each protein. Assuming that &#x0039B;<sub><italic>ip</italic></sub>&#x02208;&#x0039B; is the GO category of <italic>x</italic><sub><italic>ip</italic></sub> and <italic>dom</italic>(&#x0039B;<sub><italic>m</italic></sub>) is a set of possible GO annotations in &#x0039B;<sub><italic>m</italic></sub>, we have <italic>x</italic><sub><italic>ip</italic></sub>&#x02208;<italic>dom</italic>(&#x0039B;<sub><italic>m</italic></sub>) if &#x0039B;<sub><italic>ip</italic></sub> &#x0003D; &#x0039B;<sub><italic>m</italic></sub>. The size of <italic>dom</italic>(&#x0039B;<sub><italic>m</italic></sub>) is denoted as |<italic>dom</italic>(&#x0039B;<sub><italic>m</italic></sub>)|.</p>
<p>To indicate the functional modules of proteins, we adopt a <italic>n</italic><sub><italic>V</italic></sub>&#x000D7;1 vector, i.e., <italic>C</italic> &#x0003D; [<italic>C</italic><sub><italic>i</italic></sub>](1 &#x02264; <italic>i</italic> &#x02264; <italic>n</italic><sub><italic>V</italic></sub>, 1 &#x02264; <italic>C</italic><sub><italic>i</italic></sub> &#x02264; <italic>K</italic>), where <italic>C</italic><sub><italic>i</italic></sub> represents the functional module label of <italic>v</italic><sub><italic>i</italic></sub>. Therefore, for an arbitrary protein, i.e., <italic>v</italic><sub><italic>i</italic></sub>, its <italic>C</italic><sub><italic>i</italic></sub> is equal to <italic>k</italic> if it is in the <italic>k</italic>-th functional module.</p></sec>
<sec>
<title>2.2. Generating Functional Module Labels</title>
<p>For an arbitrary protein denoted as <italic>v</italic><sub><italic>i</italic></sub>, its functional module label <italic>C</italic><sub><italic>i</italic></sub> is chosen from a Multinomial distribution, which is defined as (1).</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M1"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>|</mml:mo><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mi>K</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003B1;<sub><italic>k</italic></sub> is the probability of a protein that is assigned to the <italic>k</italic>-th functional module and <inline-formula><mml:math id="M2"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. Instead of predetermining the value of each element in &#x003B1;, we consider &#x003B1; as a random variable and sample it by using a Dirichlet distribution with a parameter &#x003B6;.</p></sec>
<sec>
<title>2.3. Generating GO Information of Proteins</title>
<p>In order to completely retain the relationship between GO categories and their corresponding annotations, we sample the GO annotations of <italic>v</italic><sub><italic>i</italic></sub> with two steps. Specifically, to obtain <italic>x</italic><sub><italic>ip</italic></sub>, we first choose its GO category, i.e., &#x0039B;<sub><italic>ip</italic></sub>, from a Multinomial distribution that is specific to the functional module of <italic>v</italic><sub><italic>i</italic></sub>. Hence, we have</p>
<disp-formula id="E2"><label>(2)</label><mml:math id="M3"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x0039B;</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x0039B;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x02264;</mml:mo><mml:msub><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x0039B;</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the above equation, &#x003B8;<sub><italic>C</italic><sub><italic>i</italic></sub></sub> is a <italic>n</italic><sub>&#x0039B;</sub>-dimensional variable randomly selected from the Dirichlet distribution with a parameter &#x003BB;<sub><italic>C</italic><sub><italic>i</italic></sub></sub>. As the subscript of &#x003BB;<sub><italic>C</italic><sub><italic>i</italic></sub></sub>, <italic>C</italic><sub><italic>i</italic></sub> indicates that the probability distribution of &#x003BB;<sub><italic>C</italic><sub><italic>i</italic></sub></sub> is conditioned on the functional module label of <italic>v</italic><sub><italic>i</italic></sub>.</p>
<p>Once the GO category of <italic>x</italic><sub><italic>ip</italic></sub> is determined, the next step is to select the annotation of <italic>x</italic><sub><italic>ip</italic></sub> from the domain of &#x0039B;<sub><italic>ip</italic></sub>. Assuming that &#x0039B;<sub><italic>ip</italic></sub> is actually the <italic>m</italic>&#x02212;th category in &#x0039B;, i.e., &#x0039B;<sub><italic>m</italic></sub>, the value of <italic>x</italic><sub><italic>ip</italic></sub> is then sampled from a Multinomial distribution defined as:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M4"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003B2;</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x02264;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mi>m</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x0039B;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>val</italic><sub><italic>mt</italic></sub> is the <italic>t</italic>-th annotation in <italic>dom</italic>(&#x0039B;<sub><italic>m</italic></sub>). Regarding the subscripts <italic>C</italic><sub><italic>i</italic></sub> and <italic>m</italic>, their combination indicates that the Multinomial distribution of <italic>x</italic><sub><italic>ip</italic></sub> is specific to the functional module of <italic>v</italic><sub><italic>i</italic></sub> and the GO category &#x0039B;<sub><italic>m</italic></sub>. In other words, proteins in the same functional module share similar Multinomial distributions of GO annotations, which can differ across different GO categories or functional modules. To generate &#x003B2;<sub><italic>C</italic><sub><italic>i</italic></sub><italic>m</italic></sub>, we also place a Dirichlet distribution over it with a prior parameter &#x003BC;<sub><italic>C</italic><sub><italic>i</italic></sub><italic>m</italic></sub>. The graphical presentation of generating the GO information of proteins is presented in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Graphical model representation of generating the GO information of proteins.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0001.tif"/>
</fig></sec>
<sec>
<title>2.4. Generating PPIs</title>
<p>As mentioned before, we introduce <italic>A</italic> to represent the interaction relationships for all pairwise proteins in <italic>G</italic>. Hence, generating PPIs in a PPI network is identical to generate <italic>A</italic>. Following the observation that proteins in the same functional module are densely connected, the value of <italic>A</italic><sub><italic>ij</italic></sub> is dependent on a finite mixture of functional modules labels according to Stochastic BlockModel (Nowicki and Snijders, <xref ref-type="bibr" rid="B18">2001</xref>).</p>
<p>Given two proteins, i.e., <italic>v</italic><sub><italic>i</italic></sub> and <italic>v</italic><sub><italic>j</italic></sub>, the probability of <italic>v</italic><sub><italic>i</italic></sub> interacting with <italic>v</italic><sub><italic>j</italic></sub> follows a Multinomial distribution described below.</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M5"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mtext class="textrm" mathvariant="normal">A</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B5;</mml:mi></mml:mstyle></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B5;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the above equation, the parameter &#x003B5;<sub><italic>kl</italic></sub> is conditioned on the functional module labels of <italic>v</italic><sub><italic>i</italic></sub> and <italic>v</italic><sub><italic>j</italic></sub>. The interaction probabilities between all pairs of functional modules are therefore parameterized by <bold>&#x003B5;</bold>, which is a <italic>K</italic>&#x000D7;<italic>K</italic> matrix. With (4), proteins in the same functional modules present similar regularities when interacting with other proteins. Similarly, we also place a Dirichlet distribution with a prior parameter <bold>&#x003C4;</bold><sub><italic>k</italic></sub> to determine <bold>&#x003B5;</bold><sub><italic>k</italic></sub>. The graphical presentation of generating PPIs is presented in <xref ref-type="fig" rid="F2">Figure 2</xref>.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Graphical model representation of generating PPIs.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0002.tif"/>
</fig>
<p>So far, the generative process of <italic>G</italic> is completed by the above generative process that involves several latent variables <bold>&#x003B1;</bold>, <bold>&#x003B8;</bold>, <bold>&#x003B2;</bold>, and <bold>&#x003B5;</bold>. Regarding the values of these variables, we also define corresponding prior parameters, i.e., <bold>&#x003B6;</bold>, <bold>&#x003BB;</bold>, <bold>&#x003BC;</bold>, and <bold>&#x003C4;</bold>, to sample them in a Bayesian manner.</p></sec>
<sec>
<title>2.5. Bayesian Decision</title>
<p>According to the above generative process, a PPI network, i.e., <italic>G</italic>, is represented as a collection of proteins, PPIs and GO annotations. To indicate the functional module label of each protein, we need to compute the probability of each possible <italic>C</italic> conditioning on both <italic>A</italic> and <italic>X</italic>, and select the one with the maximum posterior probability as the optimal result. Hence, we can formulate an optimization problem as below.</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M6"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>&#x00108;</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo class="qopname">arg</mml:mo><mml:mo class="qopname">max</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>A</mml:mi><mml:mo>,</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003B6;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003BB;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003BC;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x003C4;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>To address this problem, we apply the solution developed in Hu et al. (<xref ref-type="bibr" rid="B8">2020</xref>). Instead of explicitly determining <italic>C</italic>, this solution yields the optimal membership matrix, i.e., <inline-formula><mml:math id="M7"><mml:mover accent="true"><mml:mrow><mml:mstyle mathvariant="bold"><mml:mi>&#x003B1;</mml:mi></mml:mstyle></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> to derive &#x00108;. Specifically, for <italic>v</italic><sub><italic>i</italic></sub>, its functional module label <italic>C</italic><sub><italic>i</italic></sub> is more likely to be equal to <italic>k</italic> if <inline-formula><mml:math id="M8"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is larger.</p></sec>
<sec>
<title>2.6. Computing Interaction Probability</title>
<p>To indicate to what extent two proteins are likely to interact, a scoring function is designed by taking into account their membership distributions and network paths simultaneously. The motivation of designing such a function is twofold. First of all, for two proteins, the probability of being grouped in the same functional module is larger if their membership distributions are more similar, and accordingly they are more likely to interact with each other. On the other hand, two proteins are less likely to interact if the network path connecting them is longer. Assuming that <italic>L</italic><sub><italic>v</italic><sub><italic>i</italic></sub><italic>v</italic><sub><italic>j</italic></sub></sub> is a set of all network paths connecting <italic>v</italic><sub><italic>i</italic></sub> and <italic>v</italic><sub><italic>j</italic></sub> in <italic>G</italic> and its size is |<italic>L</italic><sub><italic>v</italic><sub><italic>i</italic></sub><italic>v</italic><sub><italic>j</italic></sub></sub>|, the scoring function is defined as below.</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M9"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>s</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:munderover></mml:mstyle><mml:mi>w</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>In the above scoring function, <italic>weight</italic>(<italic>L</italic><sub><italic>w</italic></sub>) evaluates the strength of <italic>L</italic><sub><italic>w</italic></sub> in terms of providing evidence to support the interaction between <italic>v</italic><sub><italic>i</italic></sub> and <italic>v</italic><sub><italic>j</italic></sub> and its definition is given as:</p>
<disp-formula id="E7"><label>(7)</label><mml:math id="M10"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>w</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munderover accentunder="false" accent="false"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mrow></mml:munderover></mml:mstyle><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>z</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <italic>k</italic> is the value of <italic>C</italic><sub><italic>i</italic></sub>, |<italic>L</italic><sub><italic>w</italic></sub>| is the number of proteins in <italic>L</italic><sub><italic>w</italic></sub> and <inline-formula><mml:math id="M11"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003B1;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>z</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the membership over the <italic>k</italic>-th function module for the <italic>z</italic>-th protein along the path <italic>L</italic><sub><italic>w</italic></sub>. Obviously, the value of <italic>weight</italic>(<italic>L</italic><sub><italic>w</italic></sub>) is determined by the likelihood of being group in the function module of <italic>v</italic><sub><italic>i</italic></sub> for the remaining proteins in <italic>L</italic><sub><italic>w</italic></sub>.</p>
<p>Regarding <italic>decay</italic>(<italic>L</italic><sub><italic>w</italic></sub>), the motivation of introducing this term is that it is less likely to interact with each other if two proteins are located far away from each other in a given PPI network. Hence, the definition of <italic>decay</italic>(<italic>L</italic><sub><italic>w</italic></sub>) is given by (8) where &#x003C6; is the decay coefficient and usually set to be greater than or equal to 1. Since the value of <italic>weight</italic>(<italic>L</italic><sub><italic>w</italic></sub>) ranges from 0 to 1, <italic>decay</italic>(<italic>L</italic><sub><italic>w</italic></sub>) has a decay effect as an exponentiation. The longer the length of <italic>L</italic><sub><italic>w</italic></sub> is, the more obvious the decay effect of <italic>decay</italic>(<italic>L</italic><sub><italic>w</italic></sub>) has. To achieve a balance between accuracy and time, the value of |<italic>L</italic><sub><italic>w</italic></sub>| is set to be 3 in our experiments.</p>
<disp-formula id="E8"><label>(8)</label><mml:math id="M12"><mml:mtable class="eqnarray" columnalign="right center left"><mml:mtr><mml:mtd><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x003C6;</mml:mi><mml:mo>&#x000D7;</mml:mo><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>For each pair of testing proteins, we propose a novel prediction algorithm, namely NGPM, to calculate their interacting probability. To begin with the prediction, NGPM ranks the scores of all pairs of proteins including known PPIs and newly predicted PPIs. Since a predicted PPI is more likely to be real if it is surrounded by more already known PPIs, a sliding window is set by NGPM by selecting the upper and lower 50 pairs of proteins as a reference for the given pair of proteins. NGPM calculates the percentage of known PPIs to all pairs of proteins in this window, and then regards this percentage as the interacting probability for the given pair of testing proteins.</p></sec></sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<p>In this section, the performance of NGPM has been compared with several state-of-the-art prediction algorithms on five practical PPI networks and the evaluation metrics include Precision, Recall, f-measure, AUC, and PR-AUC.</p>
<sec>
<title>3.1. Experimental Setup</title>
<p>In the experiments, five independent PPI networks collected from different species are used, and they are denoted as Yeast-Tong (Tong et al., <xref ref-type="bibr" rid="B20">2004</xref>), Yeast-Krogan (Krogan et al., <xref ref-type="bibr" rid="B13">2006</xref>), Human (Rolland et al., <xref ref-type="bibr" rid="B19">2014</xref>; Kov&#x000E1;cs et al., <xref ref-type="bibr" rid="B12">2019</xref>), <italic>Escherichia coli</italic> (<italic>E. coli</italic>) (Gagarinova et al., <xref ref-type="bibr" rid="B3">2016</xref>), and Mouse (Malty et al., <xref ref-type="bibr" rid="B16">2017</xref>) respectively. The first two datasets are obtained from the species of yeast, and the Human dataset is composed of three human PPI networks, i.e., HI-II-14 (Rolland et al., <xref ref-type="bibr" rid="B19">2014</xref>), HI-III (Rolland et al., <xref ref-type="bibr" rid="B19">2014</xref>), and HI-tested (Kov&#x000E1;cs et al., <xref ref-type="bibr" rid="B12">2019</xref>). The rest datasets are generated from other species as indicated by their names. The statistics of all these PPI networks are presented in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Statistics of PPI networks used in the experiments.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="center"><bold><italic>N</italic></bold></th>
<th valign="top" align="center"><bold>E</bold></th>
<th valign="top" align="center"><bold>k<sub><italic>av</italic></sub></bold></th>
<th valign="top" align="center"><bold>CC</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Yeast-Tong</td>
<td valign="top" align="center">964</td>
<td valign="top" align="center">3,846</td>
<td valign="top" align="center">7.98</td>
<td valign="top" align="center">0.15</td>
</tr>
<tr>
<td valign="top" align="left">Yeast-Krogan</td>
<td valign="top" align="center">2,708</td>
<td valign="top" align="center">7,123</td>
<td valign="top" align="center">5.26</td>
<td valign="top" align="center">0.19</td>
</tr>
<tr>
<td valign="top" align="left">Human</td>
<td valign="top" align="center">6,657</td>
<td valign="top" align="center">32,307</td>
<td valign="top" align="center">9.52</td>
<td valign="top" align="center">0.07</td>
</tr>
<tr>
<td valign="top" align="left"><italic>E. coli</italic></td>
<td valign="top" align="center">312</td>
<td valign="top" align="center">5,108</td>
<td valign="top" align="center">32.74</td>
<td valign="top" align="center">0.19</td>
</tr>
<tr>
<td valign="top" align="left">Mouse</td>
<td valign="top" align="center">786</td>
<td valign="top" align="center">1,975</td>
<td valign="top" align="center">5.03</td>
<td valign="top" align="center">0.15</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>N, number of proteins; E, number of PPIs; k<sub>av</sub>, average graph distance; CC, clustering coefficient</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>In the experiments, a five-fold cross-validation has been conducted to yield convincing results and the performance of NGPM is compared with that of ASNE (Liao et al., <xref ref-type="bibr" rid="B14">2018</xref>) and L3 (Kov&#x000E1;cs et al., <xref ref-type="bibr" rid="B12">2019</xref>) to demonstrate its superiority in PPI prediction. When generating the negative samples, i.e., non-interacting proteins, we adopt the same strategy as L3 for conducting a fair comparison. In particular, for each PPI network, a total of 244 pairs of non-adjacent proteins are randomly selected as negative samples and 100 pairs of them should contain at least one of proteins listed in the top 500 PPIs predicted by L3.</p></sec>
<sec>
<title>3.2. Parameter Sensitivity Analysis</title>
<p>As the most important parameter involved in NGPM, <italic>K</italic> determines the number of functional modules observed from a given PPI network. To investigate the sensitivity of NGPM to the change of <italic>K</italic>, we present the performance of NGPM by varying the value of <italic>K</italic> from 2 to 20 at a step size of 1. In doing so, we are able to determine the best value of <italic>K</italic> for each dataset.</p>
<p>Given different values of <italic>K</italic>, the performance of NGPM is presented in <xref ref-type="fig" rid="F3">Figure 3</xref>. Fluctuations are observed for the AUC and PR curves, while the f-measure cures are more stable for all datasets except Mouse. A possible reason for that phenomenon is that f-measure is a harmonic mean of Precision and Recall. Since the increase in the score of <italic>K</italic> results in opposite changes of Precision and Recall, the fluctuation in the curve of f-measure is alleviated.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The performance of NGPM given different values of <italic>K</italic>.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0003.tif"/>
</fig>
<p>Among all kinds of curves in <xref ref-type="fig" rid="F3">Figure 3</xref>, we also note that the robustness of NGPM in terms of AUC is the worst, as the AUC curves are more extensively fluctuated when compared with other curves. After investigating the experimental results, we find that the false-positive rates obtained by NGPM with different values of <italic>K</italic> are different, thus having a significant impact to the change of AUC curves. Another point worth noting is that the AUC curves are below the PR and f-measure curves for all datasets except Mouse. The reason for the unsatisfactory performance of AUC is due to the imbalance between positive and negative samples in the testing datasets.</p>
<p>According to <xref ref-type="fig" rid="F3">Figure 3</xref>, the best values of <italic>K</italic> for Yeast-Tong, Yeast-Krogan, Human, <italic>E. coli</italic>, and Mouse are 14, 6, 10, 18, and 2, respectively. Hence, in the following experiments, we use the best performance of NGPM obtained by using these values for comparison.</p></sec>
<sec>
<title>3.3. Performance Comparison</title>
<p>During the comparison, since ASNE can use different measurements to calculate the similarity between two proteins and determine their interacting probability accordingly, two most commonly used measurements including Euclidean similarity and cosine similarity are chosen in our experiments, and they are denoted as eASNE and cASNE, respectively. The results of performance comparison are shown in <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> and <xref ref-type="table" rid="T2">Table 2</xref> where <xref ref-type="fig" rid="F4">Figures 4</xref>, <xref ref-type="fig" rid="F5">5</xref> show the ROC and PR curves of L3, NGPM, and ASNE obtained in each dataset, and <xref ref-type="table" rid="T2">Table 2</xref> records the exact scores yielded by each prediction algorithm.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>The ROC curves of L3, eASNE, cASNE, and NGPM.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0004.tif"/>
</fig>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>The PR curves of L3, eASNE, cASNE, and NGPM.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0005.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>The performance of models.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Dataset</bold></th>
<th valign="top" align="center"><bold>Model</bold></th>
<th valign="top" align="center" colspan="3" style="border-bottom: thin solid #000000;"><bold>f-measure</bold></th>
<th valign="top" align="center"><bold>AUC</bold></th>
<th valign="top" align="center"><bold>PR-AUC</bold></th>
</tr>
<tr>
<th/>
<th/>
<th valign="top" align="center"><bold>Precision</bold></th>
<th valign="top" align="center"><bold>Recall</bold></th>
<th valign="top" align="center"><bold>f-measure</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Yeast-Tong</td>
<td valign="top" align="center">L3</td>
<td valign="top" align="center"><bold>0.84</bold></td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.54</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center"><bold>0.87</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">eASNE</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.40</td>
<td valign="top" align="center">0.72</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">cASNE</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.46</td>
<td valign="top" align="center">0.73</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td valign="top" align="center">NGPM</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center"><bold>0.96</bold></td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center"><bold>0.68</bold></td>
<td valign="top" align="center"><bold>0.87</bold></td>
</tr>
 <tr>
<td valign="top" align="left">Yeast-Krogan</td>
<td valign="top" align="center">L3</td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center">0.19</td>
<td valign="top" align="center">0.31</td>
<td valign="top" align="center"><bold>0.57</bold></td>
<td valign="top" align="center"><bold>0.91</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">eASNE</td>
<td valign="top" align="center">0.78</td>
<td valign="top" align="center">0.61</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center">0.30</td>
<td valign="top" align="center">0.79</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">cASNE</td>
<td valign="top" align="center">0.69</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.01</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.80</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td valign="top" align="center">NGPM</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center"><bold>0.88</bold></td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.89</td>
</tr>
 <tr>
<td valign="top" align="left">Human</td>
<td valign="top" align="center">L3</td>
<td valign="top" align="center">0.98</td>
<td valign="top" align="center">0.39</td>
<td valign="top" align="center">0.56</td>
<td valign="top" align="center"><bold>0.61</bold></td>
<td valign="top" align="center"><bold>0.98</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">eASNE</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.90</td>
<td valign="top" align="center">0.43</td>
<td valign="top" align="center">0.96</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">cASNE</td>
<td valign="top" align="center"><bold>1.00</bold></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.03</td>
<td valign="top" align="center">0.47</td>
<td valign="top" align="center">0.96</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td valign="top" align="center">NGPM</td>
<td valign="top" align="center">0.96</td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center"><bold>0.95</bold></td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.95</td>
</tr>
 <tr>
<td valign="top" align="left"><italic>E. coli</italic></td>
<td valign="top" align="center">L3</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.53</td>
<td valign="top" align="center">0.64</td>
<td valign="top" align="center"><bold>0.56</bold></td>
<td valign="top" align="center"><bold>0.85</bold></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">eASNE</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">1.00</td>
<td valign="top" align="center">0.89</td>
<td valign="top" align="center">0.52</td>
<td valign="top" align="center">0.82</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">cASNE</td>
<td valign="top" align="center"><bold>0.85</bold></td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.50</td>
<td valign="top" align="center">0.81</td>
</tr>
<tr style="border-bottom: thin solid #000000;">
<td/>
<td valign="top" align="center">NGPM</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center"><bold>1.00</bold></td>
<td valign="top" align="center"><bold>0.89</bold></td>
<td valign="top" align="center"><bold>0.56</bold></td>
<td valign="top" align="center">0.82</td>
</tr>
 <tr>
<td valign="top" align="left">Mouse</td>
<td valign="top" align="center">L3</td>
<td valign="top" align="center"><bold>0.91</bold></td>
<td valign="top" align="center">0.23</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.59</td>
<td valign="top" align="center">0.75</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">eASNE</td>
<td valign="top" align="center">0.55</td>
<td valign="top" align="center">0.77</td>
<td valign="top" align="center">0.65</td>
<td valign="top" align="center">0.37</td>
<td valign="top" align="center">0.56</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">cASNE</td>
<td valign="top" align="center">0.60</td>
<td valign="top" align="center">0.02</td>
<td valign="top" align="center">0.04</td>
<td valign="top" align="center">0.44</td>
<td valign="top" align="center">0.57</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">NGPM</td>
<td valign="top" align="center">0.68</td>
<td valign="top" align="center"><bold>0.94</bold></td>
<td valign="top" align="center"><bold>0.79</bold></td>
<td valign="top" align="center"><bold>0.86</bold></td>
<td valign="top" align="center"><bold>0.91</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>Best scores are bolded</italic>.</p>
</table-wrap-foot>
</table-wrap>
<p>When compared with ASNE, NGPM obtains a better performance on each metric across all the datasets except for Human and <italic>E. coli</italic>. On average, NGPM performs better by 6.28, 17.28, 12.08, 49.50, and 15.32% in terms of Precision, Recall, f-measure, AUC, and PR-AUC, respectively than eASNE while cASNE yields the worst performance among them. However, both NGPM and ASNE do not perform well on the Human dataset in terms of AUC. A main reason for that phenomenon is due to the serious imbalance between interacting and non-interacting proteins in the Human dataset. As mentioned before, the strategy of selecting negative samples in NGPM is as same as in L3, but it leads to the imbalance of interacting samples and non-interacting samples. Since the Human dataset is the largest one, it has more than 30,000 positive samples, while the negative sample is only 244. Thus it suffers the disadvantage of imbalance seriously and smaller AUC scores are obtained by all algorithms when compared with the other datasets.</p>
<p>In order to more specifically illustrate the advantage of NGPM compared to ASNE in PPI prediction, we take the prediction results of NGPM on the Human dataset as an example. In particular, the nodes in <xref ref-type="fig" rid="F6">Figure 6</xref> represent proteins, and an edge connecting two nodes represents the interaction between them. Regarding the two proteins UBE2D3 and CLNS1A, they are classified as the negative sample in the testing dataset and thus there is no edge between them. However, ASNE predicts that they can interact with a probability as high as 0.76, thus leading to a wrong conclusion. NGPM accurately predicts the true relationship between UBE2D3 and CLNS1A. In the prediction result of NGPM, the interacting probability between these two proteins is &#x0003C; 0.4. Hence, NGPM is believed to be more reliable than ASNE when predicting PPIs. In addition, PPIs indicated by red lines are all successfully predicted by NGPM but incorrectly predicted by ASNE. These interactions have been verified by the BioGRID database (Chatr-Aryamontri et al., <xref ref-type="bibr" rid="B1">2017</xref>) and can provide help for understanding the biological processes in the cell. Among them, CLNS1A, SNRPD1, EPB41, SNRPG, SNRPD3, and LSM6 are all important components of the cytoplasm, they can form protein complexes together. It is for this reason that NGPM is able to provide a precise prediction result for these proteins. Besides, all the proteins except EPB41 can participate in the process of RNA molecular interaction. Proteins UBE2D3 and RNF115 can also add ubiquitin groups to the proteins in cells to help them form ubiquitin chains, so that they can complete the catalysis of the ubiquitin reaction due to the interaction between them.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>An illustration of PPIs correctly identified by NGPM in the Human dataset.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0006.tif"/>
</fig>
<p>In addition to correctly predict known PPIs, NGPM is also capable of predicting novel PPIs that are not found in the testing dataset. Since NGPM allows each protein to be associated with a membership distribution and also finds the path between two proteins, the interacting probability can be determined by NGPM for any pair of proteins in a PPI network given such information. As indicated by <xref ref-type="fig" rid="F7">Figure 7</xref>, several pairs of proteins extracted from the Yeast-Tong dataset are presented. PPIs represented by the edges are novel PPIs predicted by NGPM and these interactions have been confirmed by BioGRID database (Chatr-Aryamontri et al., <xref ref-type="bibr" rid="B1">2017</xref>). In this regard, the ability of NGPM in predicting novel PPIs could thus be verified.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>An illustration of novel PPIs identified by NGPM in Yeast-Tong dataset.</p></caption>
<graphic xlink:href="fmicb-12-735329-g0007.tif"/>
</fig>
<p>In order to verify whether NGPM can effectively eliminate the negative impact imposed by noise data such as false positives and false negatives after combining gene ontology and network topology, we compare the performance of NGPM on five PPI network with L3. From <xref ref-type="table" rid="T2">Table 2</xref>, NGPM obtains the best Recall and f-measure scores on all datasets. Specifically, when compared with L3, the performance of NGPM is better by 174.57, 79.42, 1.68, and 1.83% in terms of Recall, f-measure, AUC, and PR-AUC, respectively, and hence NGPM can reduce the negative impact caused by the noise data for PPI prediction. However, NGPM does not achieve the best performance on Precision, there are several reasons for this phenomenon. First of all The performance of NGPM is constrained by the existence of network paths. If there is no path between two proteins, NGPM can not predict the interaction between them and hence it will consider their interacting probability as 0. In doing so, a part of PPIs in the testing dataset are able to be predicted as non-interacting protein pairs, thus increasing the false negatives in the prediction result. Secondly, when predicting the interacting probability for proteins pairs, the longest path is set to be 3 in experiments, which is constrained by the computational efficiency of NGPM. A longer path will consume more time and we may be unable to obtain the prediction result after an acceptable period. Although the longer a path is, the less impact it has on determining the interacting probability between two proteins and consequently some PPIs are falsely predicted by NGPM. In this regard, the number of false positive samples obtained by NGPM is larger than the other algorithms, thus reducing the prediction accuracy of NGPM.</p></sec></sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<p>In this paper, an efficient network-based prediction algorithm, namely NGPM, is proposed to predict PPIs by additionally considering the GO information of protein. The motivation behind NGPM is to make use of the property of functional modularity observed in PPI networks and also to combine the GO knowledge to alleviate the negative impact imposed by the noise data. Hence, by simulating the generative process of a PPI network, NGPM is able to incorporate these two kinds of information and optimize the membership distributions of proteins over functional modules. After that, a new scoring function is then designed to compute the interacting probability between two proteins. Experimental results have demonstrated that NGPM could better solve the prediction problem of PPIs as it yields a superior performance in terms of several independent metrics when compared with state-of-the-art prediction algorithms. In this regard, the novel PPIs predicted by NGPM may probably missed due to the constraints of laboratory experiments.</p>
<p>Several reasons can be summarized to explain the promising accuracy of NGPM. First of all, for a given protein, the modularity property of PPI networks allows NGPM to search potential interacting partners in a more accurate range, as proteins in the same functional module are more likely to interact with each other. However, there is no such a prior knowledge about the existence of functional modules in a PPI network before PPI prediction. By assuming the existence of total <italic>K</italic> functional modules embedded in a given PPI network, NGPM combines both network structure and GO to simulate the generative process of this network and then adopts an efficient solution to infer the membership distributions of proteins over functional modules. In doing so, the accuracy of PPI prediction can be improved. Secondly, to indicate how likely two proteins interact with each other, a novel scoring function is specifically designed by taking into account both network paths and membership distributions of proteins. It is also meaningful from a biological view. In particular, two proteins are more likely to interact with each other if they share many common interacting partners and are grouped into the same functional module together with these partners. Lastly, unlike conventional PPI prediction algorithms, NGPM does not rely on the selection of classifiers nor the generation of negative samples, thus making its performance more robust. One should note that the strategy of generating negative samples we describe in section 3.1 is only used for testing rather than training.</p>
<p>In addition to GO, there are also other kinds of biological information that can be used to characterize proteins. It is possible for NGPM to incorporate these biological information. Specifically, when generating the GO information of proteins, NGPM adopts different Multinomial distributions to sample the GO category and corresponding annotations. Hence, given a particular kind of biological information, we are able to incorporate it into NGPM if it can be represented as a set of attribute values taken by proteins.</p>
<p>Regarding future work, we would like to unfold it from three aspects. Firstly, since the longest length of paths used in (6) affects the performance of NGPM in some ways and we currently set it as 3 in the experiments, we intend to release this constraint by allowing NPGM to consider more path information. However, the increase in the longest length of paths could result in a consequence that more time will be taken by NPGM. Furthermore, there are many variational parameters that have to be optimized. The increase in the scale of PPI networks will obviously take more time to optimize these variational parameters. Hence, the current version of NGPM is not applicable for large-scale PPI prediction. To overcome this limitation, we would like to develop a distributed version of NGPM by following the MapReduce framework. Furthermore, regarding <italic>K</italic>, we have performed several trials to find its best value and thus we are also interested in providing a simpler, yet effective, strategy to determine its value. Lastly, since self-supervised pre-training has proven beneficial for many computer vision tasks, we would like to explore the possibility of pre-training NGPM on a different dataset when predicting PPIs.</p></sec>
<sec sec-type="data-availability-statement" id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: <ext-link ext-link-type="uri" xlink:href="https://gitee.com/allenv5/NGPM">https://gitee.com/allenv5/NGPM</ext-link>.</p></sec>
<sec id="s6">
<title>Author Contributions</title>
<p>LH conceived of the study and drafted the manuscript. XW implemented the algorithms and carried out the experiments. LH, PH, and Z-HY conceived of the study, participated in its design and coordination, and helped to draft the manuscript. XW and Y-AH performed the statistical analysis. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p></sec>
<sec sec-type="disclaimer" id="s7">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p></sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chatr-Aryamontri</surname> <given-names>A.</given-names></name> <name><surname>Oughtred</surname> <given-names>R.</given-names></name> <name><surname>Boucher</surname> <given-names>L.</given-names></name> <name><surname>Rust</surname> <given-names>J.</given-names></name> <name><surname>Chang</surname> <given-names>C.</given-names></name> <name><surname>Kolas</surname> <given-names>N. K.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>The biogrid interaction database: 2017 update</article-title>. <source>Nucleic Acids Res</source>. <volume>45</volume>, <fpage>D369</fpage>&#x02013;<lpage>D379</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1102</pub-id><pub-id pub-id-type="pmid">27980099</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fields</surname> <given-names>S.</given-names></name> <name><surname>Sternglanz</surname> <given-names>R.</given-names></name></person-group> (<year>1994</year>). <article-title>The two-hybrid system: an assay for protein-protein interactions</article-title>. <source>Trends Genet</source>. <volume>10</volume>, <fpage>286</fpage>&#x02013;<lpage>292</lpage>. <pub-id pub-id-type="doi">10.1016/0168-9525(90)90012-U</pub-id><pub-id pub-id-type="pmid">11700327</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gagarinova</surname> <given-names>A.</given-names></name> <name><surname>Stewart</surname> <given-names>G.</given-names></name> <name><surname>Samanfar</surname> <given-names>B.</given-names></name> <name><surname>Phanse</surname> <given-names>S.</given-names></name> <name><surname>White</surname> <given-names>C. A.</given-names></name> <name><surname>Aoki</surname> <given-names>H.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Systematic genetic screens reveal the dynamic global functional organization of the bacterial translation machinery</article-title>. <source>Cell Rep</source>. <volume>17</volume>, <fpage>904</fpage>&#x02013;<lpage>916</lpage>. <pub-id pub-id-type="doi">10.1016/j.celrep.2016.09.040</pub-id><pub-id pub-id-type="pmid">27732863</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ho</surname> <given-names>Y.</given-names></name> <name><surname>Gruhler</surname> <given-names>A.</given-names></name> <name><surname>Heilbut</surname> <given-names>A.</given-names></name> <name><surname>Bader</surname> <given-names>G. D.</given-names></name> <name><surname>Moore</surname> <given-names>L.</given-names></name> <name><surname>Adams</surname> <given-names>S.-L.</given-names></name> <etal/></person-group>. (<year>2002</year>). <article-title>Systematic identification of protein complexes in saccharomyces cerevisiae by mass spectrometry</article-title>. <source>Nature</source> <volume>415</volume>, <fpage>180</fpage>&#x02013;<lpage>183</lpage>. <pub-id pub-id-type="doi">10.1038/415180a</pub-id><pub-id pub-id-type="pmid">11805837</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hsin Liu</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>K.-C.</given-names></name> <name><surname>Yuan</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>Human protein-protein interaction prediction by a novel sequence-based co-evolution method: co-evolutionary divergence</article-title>. <source>Bioinformatics</source> <volume>29</volume>, <fpage>92</fpage>&#x02013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts620</pub-id><pub-id pub-id-type="pmid">23080115</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Chan</surname> <given-names>K. C.</given-names></name></person-group> (<year>2015</year>). <article-title>Discovering variable-length patterns in protein sequences for protein-protein interaction prediction</article-title>. <source>IEEE Trans. Nanobiosci</source>. <volume>14</volume>, <fpage>409</fpage>&#x02013;<lpage>416</lpage>. <pub-id pub-id-type="doi">10.1109/TNB.2015.2429672</pub-id><pub-id pub-id-type="pmid">26011889</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Chan</surname> <given-names>K. C.</given-names></name></person-group> (<year>2017</year>). <article-title>Extracting coevolutionary features from protein sequences for predicting protein-protein interactions</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinform</source>. <volume>14</volume>, <fpage>155</fpage>&#x02013;<lpage>166</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2016.2520923</pub-id><pub-id pub-id-type="pmid">26812730</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Chan</surname> <given-names>K. C.</given-names></name> <name><surname>Yuan</surname> <given-names>X.</given-names></name> <name><surname>Xiong</surname> <given-names>S.</given-names></name></person-group> (<year>2020</year>). <article-title>A variational bayesian framework for cluster analysis in a complex network</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>32</volume>, <fpage>2115</fpage>&#x02013;<lpage>2128</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2019.2914200</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>Y.-A.</given-names></name> <name><surname>Hu</surname> <given-names>P.</given-names></name> <name><surname>You</surname> <given-names>Z.-H.</given-names></name></person-group> (<year>2021a</year>). <article-title>A survey on computational models for predicting protein-protein interactions</article-title>. <source>Brief. Bioinform</source>. <volume>2021</volume>:<fpage>bbab036</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab036</pub-id><pub-id pub-id-type="pmid">33693513</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Yuan</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>P.</given-names></name> <name><surname>Chan</surname> <given-names>K. C.</given-names></name></person-group> (<year>2017</year>). <article-title>Efficiently predicting large-scale protein-protein interactions using mapreduce</article-title>. <source>Comput. Biol. Chem</source>. <volume>69</volume>, <fpage>202</fpage>&#x02013;<lpage>206</lpage>. <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2017.03.009</pub-id><pub-id pub-id-type="pmid">28396055</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hu</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Pan</surname> <given-names>X.</given-names></name> <name><surname>Yan</surname> <given-names>H.</given-names></name> <name><surname>You</surname> <given-names>Z.-H.</given-names></name></person-group> (<year>2021b</year>). <article-title>HISCF: leveraging higher-order structures for clustering analysis in biological networks</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>542</fpage>&#x02013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa775</pub-id><pub-id pub-id-type="pmid">32931549</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kov&#x000E1;cs</surname> <given-names>I. A.</given-names></name> <name><surname>Luck</surname> <given-names>K.</given-names></name> <name><surname>Spirohn</surname> <given-names>K.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>Pollis</surname> <given-names>C.</given-names></name> <name><surname>Schlabach</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Network-based prediction of protein interactions</article-title>. <source>Nat. Commun</source>. <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-09177-y</pub-id><pub-id pub-id-type="pmid">30886144</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krogan</surname> <given-names>N. J.</given-names></name> <name><surname>Cagney</surname> <given-names>G.</given-names></name> <name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Zhong</surname> <given-names>G.</given-names></name> <name><surname>Guo</surname> <given-names>X.</given-names></name> <name><surname>Ignatchenko</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2006</year>). <article-title>Global landscape of protein complexes in the yeast saccharomyces cerevisiae</article-title>. <source>Nature</source> <volume>440</volume>, <fpage>637</fpage>&#x02013;<lpage>643</lpage>. <pub-id pub-id-type="doi">10.1038/nature04670</pub-id><pub-id pub-id-type="pmid">16554755</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>L.</given-names></name> <name><surname>He</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>H.</given-names></name> <name><surname>Chua</surname> <given-names>T.-S.</given-names></name></person-group> (<year>2018</year>). <article-title>Attributed social network embedding</article-title>. <source>IEEE Trans. Knowl. Data Eng</source>. <volume>30</volume>, <fpage>2257</fpage>&#x02013;<lpage>2270</lpage>. <pub-id pub-id-type="doi">10.1109/TKDE.2018.2819980</pub-id></citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Ren</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>Q.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name></person-group> (<year>2020</year>). <article-title>Predicting lncRNA-miRNA interactions based on logistic matrix factorization with neighborhood regularized</article-title>. <source>Knowl. Based Syst</source>. <volume>191</volume>:<fpage>105261</fpage>. <pub-id pub-id-type="doi">10.1016/j.knosys.2019.105261</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Malty</surname> <given-names>R. H.</given-names></name> <name><surname>Aoki</surname> <given-names>H.</given-names></name> <name><surname>Kumar</surname> <given-names>A.</given-names></name> <name><surname>Phanse</surname> <given-names>S.</given-names></name> <name><surname>Amin</surname> <given-names>S.</given-names></name> <name><surname>Zhang</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>A map of human mitochondrial protein interactions linked to neurodegeneration reveals new mechanisms of redox homeostasis and nf-&#x003BA;b signaling</article-title>. <source>Cell Syst</source>. <volume>5</volume>, <fpage>564</fpage>&#x02013;<lpage>577</lpage>. <pub-id pub-id-type="doi">10.1016/j.cels.2017.10.010</pub-id><pub-id pub-id-type="pmid">29128334</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mirabello</surname> <given-names>C.</given-names></name> <name><surname>Wallner</surname> <given-names>B.</given-names></name></person-group> (<year>2017</year>). <article-title>Interpred: a pipeline to identify and model protein-protein interactions</article-title>. <source>Proteins</source> <volume>85</volume>, <fpage>1159</fpage>&#x02013;<lpage>1170</lpage>. <pub-id pub-id-type="doi">10.1002/prot.25280</pub-id><pub-id pub-id-type="pmid">28263438</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nowicki</surname> <given-names>K.</given-names></name> <name><surname>Snijders</surname> <given-names>T. A. B.</given-names></name></person-group> (<year>2001</year>). <article-title>Estimation and prediction for stochastic blockstructures</article-title>. <source>J. Am. Stat. Assoc</source>. <volume>96</volume>, <fpage>1077</fpage>&#x02013;<lpage>1087</lpage>. <pub-id pub-id-type="doi">10.1198/016214501753208735</pub-id></citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rolland</surname> <given-names>T.</given-names></name> <name><surname>Ta&#x000E7;san</surname> <given-names>M.</given-names></name> <name><surname>Charloteaux</surname> <given-names>B.</given-names></name> <name><surname>Pevzner</surname> <given-names>S. J.</given-names></name> <name><surname>Zhong</surname> <given-names>Q.</given-names></name> <name><surname>Sahni</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>A proteome-scale map of the human interactome network</article-title>. <source>Cell</source> <volume>159</volume>, <fpage>1212</fpage>&#x02013;<lpage>1226</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2014.10.050</pub-id><pub-id pub-id-type="pmid">25416956</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tong</surname> <given-names>A. H. Y.</given-names></name> <name><surname>Lesage</surname> <given-names>G.</given-names></name> <name><surname>Bader</surname> <given-names>G. D.</given-names></name> <name><surname>Ding</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Xin</surname> <given-names>X.</given-names></name> <etal/></person-group>. (<year>2004</year>). <article-title>Global mapping of the yeast genetic interaction network</article-title>. <source>Science</source> <volume>303</volume>, <fpage>808</fpage>&#x02013;<lpage>813</lpage>. <pub-id pub-id-type="doi">10.1126/science.1091317</pub-id><pub-id pub-id-type="pmid">14764870</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Hu</surname> <given-names>P.</given-names></name> <name><surname>Hu</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>A novel stochastic block model for network-based prediction of protein-protein interactions,</article-title> in <source>International Conference on Intelligent Computing</source> (<publisher-loc>Bari</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>621</fpage>&#x02013;<lpage>632</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-60802-6_54</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>You</surname> <given-names>Z.-H.</given-names></name> <name><surname>Yu</surname> <given-names>J.-Z.</given-names></name> <name><surname>Zhu</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Wen</surname> <given-names>Z.-K.</given-names></name></person-group> (<year>2014</year>). <article-title>A mapreduce based parallel SVM for large-scale predicting protein-protein interactions</article-title>. <source>Neurocomputing</source> <volume>145</volume>, <fpage>37</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2014.05.072</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zahiri</surname> <given-names>J.</given-names></name> <name><surname>Yaghoubi</surname> <given-names>O.</given-names></name> <name><surname>Mohammad-Noori</surname> <given-names>M.</given-names></name> <name><surname>Ebrahimpour</surname> <given-names>R.</given-names></name> <name><surname>Masoudi-Nejad</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>PPIevo: Protein-protein interaction prediction from PSSM based evolutionary information</article-title>. <source>Genomics</source> <volume>102</volume>, <fpage>237</fpage>&#x02013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2013.05.006</pub-id><pub-id pub-id-type="pmid">23747746</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Yang</surname> <given-names>P.</given-names></name> <name><surname>Feng</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Using network distance analysis to predict lncrna-mirna interactions</article-title>. <source>Interdiscip. Sci</source>. <volume>13</volume>, <fpage>535</fpage>&#x02013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1007/s12539-021-00458-z</pub-id><pub-id pub-id-type="pmid">34232474</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Q. C.</given-names></name> <name><surname>Petrey</surname> <given-names>D.</given-names></name> <name><surname>Deng</surname> <given-names>L.</given-names></name> <name><surname>Qiang</surname> <given-names>L.</given-names></name> <name><surname>Shi</surname> <given-names>Y.</given-names></name> <name><surname>Thu</surname> <given-names>C. A.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Structure-based prediction of protein-protein interactions on a genome-wide scale</article-title>. <source>Nature</source> <volume>490</volume>, <fpage>556</fpage>&#x02013;<lpage>560</lpage>. <pub-id pub-id-type="doi">10.1038/nature11503</pub-id><pub-id pub-id-type="pmid">23023127</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>H.</given-names></name> <name><surname>Bilgin</surname> <given-names>M.</given-names></name> <name><surname>Bangham</surname> <given-names>R.</given-names></name> <name><surname>Hall</surname> <given-names>D.</given-names></name> <name><surname>Casamayor</surname> <given-names>A.</given-names></name> <name><surname>Bertone</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2001</year>). <article-title>Global analysis of protein activities using proteome chips</article-title>. <source>Science</source> <volume>293</volume>, <fpage>2101</fpage>&#x02013;<lpage>2105</lpage>. <pub-id pub-id-type="doi">10.1126/science.1062191</pub-id><pub-id pub-id-type="pmid">11474067</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This work has been supported in part by the Natural Science Foundation of Xinjiang Uygur Autonomous Region under grant 2021D01D05 and in part by the Pioneer Hundred Talents Program of Chinese Academy of Sciences.</p>
</fn>
</fn-group>
</back>
</article> 