<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">758131</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.758131</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Protein Function Prediction Based on PPI Networks: Network Reconstruction vs Edge Enrichment</article-title>
<alt-title alt-title-type="left-running-head">Zhou et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Network Reconstruction and Edge Enrichment</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhou</surname>
<given-names>Jiaogen</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1272427/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Xiong</surname>
<given-names>Wei</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Yang</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Guan</surname>
<given-names>Jihong</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Jiangsu Provincial Engineering Research Center for Intelligent Monitoring and Ecological Management of Pond and Reservoir Water Environment, Huaiyin Normal University, <addr-line>Huian</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Shanghai Key Lab of Intelligent Information Processing, and School of Computer Science, Fudan University, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Department of Computer Science and Technology, Tongji University, <addr-line>Shanghai</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/562375/overview">Liang Cheng</ext-link>, Harbin Medical University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/775059/overview">Cheng Liang</ext-link>, Shandong Normal University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1466937/overview">Yongjun Tang</ext-link>, Central South University, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Jihong Guan, <email>jhguan@tongji.edu.cn</email>
</corresp>
<fn fn-type="equal" id="fn1">
<label>
<sup>&#x2020;</sup>
</label>
<p>These authors have contributed equally to this&#x20;work</p>
</fn>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>14</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>758131</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>11</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Zhou, Xiong, Wang and Guan.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Zhou, Xiong, Wang and Guan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Over the past decades, massive amounts of protein-protein interaction (PPI) data have been accumulated due to the advancement of high-throughput technologies, and but data quality issues (noise or incompleteness) of PPI have been still affecting protein function prediction accuracy based on PPI networks. Although two main strategies of <italic>network reconstruction</italic> and <italic>edge enrichment</italic> have been reported on the effectiveness of boosting the prediction performance in numerous literature studies, there still lack comparative studies of the performance differences between <italic>network reconstruction</italic> and <italic>edge enrichment</italic>. Inspired by the question, this study first uses three protein similarity metrics (local, global and sequence) for network reconstruction and edge enrichment in PPI networks, and then evaluates the performance differences of network reconstruction, edge enrichment and the original networks on two real PPI datasets. The experimental results demonstrate that edge enrichment work better than both network reconstruction and original networks. Moreover, for the edge enrichment of PPI networks, the sequence similarity outperformes both local and global similarity. In summary, our study can help biologists select suitable pre-processing schemes and achieve better protein function prediction for PPI networks.</p>
</abstract>
<kwd-group>
<kwd>edge enrichment</kwd>
<kwd>network reconstruction</kwd>
<kwd>protein-protein interaction networks</kwd>
<kwd>protein function prediction</kwd>
<kwd>protein sequence annotation</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Over the past decades, massive amounts of un-annotated protein sequence data have been accumulated with the advancement of high-throughput biological technologies. Due to high costs and time-consummation of experimental determining protein function annotation, the proportion of annotated proteins has been still relatively low (<xref ref-type="bibr" rid="B21">Sharan et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B5">Barrell et&#x20;al., 2009</xref>). The increasing efforts have been made to predict protein functions.</p>
<p>As the best-known and early method of protein function prediction, homology-based prediction method indeed gave rise to a series of protein function prediction methods based on protein sequence or structural similarity (<xref ref-type="bibr" rid="B23">Sleator and Walsh, 2010</xref>). At the same time, the emerging of available protein databases, such as FATCAT (<xref ref-type="bibr" rid="B36">Ye and Godzik, 2004</xref>), PAST (<xref ref-type="bibr" rid="B27">T&#xe4;ubig et&#x20;al., 2006</xref>) and PROCAT (<xref ref-type="bibr" rid="B30">Wallace et&#x20;al., 1996</xref>), has further helped to improve the effectiveness of protein prediction. However, the low sequence similarity scores often occur when comparing target protein sequences with source protein sequences (<xref ref-type="bibr" rid="B17">Ofran et&#x20;al., 2005</xref>), and thus this significantly reduces the effective application of homology-based prediction methods.</p>
<p>With the increasing amounts of the measured protein-protein interaction (PPI) data, more and more protein function prediction methods based on PPI networks are proposed and generally outperform the above homology-based prediction methods. In PPI networks, proteins and protein-protein interactions are represented by nodes and edges, respectively (<xref ref-type="bibr" rid="B21">Sharan et&#x20;al., 2007</xref>; <xref ref-type="bibr" rid="B8">Chen et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B32">Wu et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B29">Waiho et&#x20;al., 2021</xref>). Up to now, numerous algorithms have been used in protein function prediction based on PPI networks, such as edge-betweenness clustering (<xref ref-type="bibr" rid="B11">Dunn et&#x20;al., 2005</xref>), Graphlet-based edge clustering (<xref ref-type="bibr" rid="B24">Solava et&#x20;al., 2012</xref>), clique percolation (<xref ref-type="bibr" rid="B2">Adamcsek et&#x20;al., 2006</xref>), GRAAL (<xref ref-type="bibr" rid="B13">Kuchaiev et&#x20;al., 2010</xref>), hybrid-property based method (<xref ref-type="bibr" rid="B12">Hu et&#x20;al., 2011</xref>), and IsoRank (<xref ref-type="bibr" rid="B22">Singh et&#x20;al., 2008</xref>). Moreover, advanced machine learning and deep learning techniques have also been used for protein function prediction, including collective classification (<xref ref-type="bibr" rid="B33">Xiong et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B31">Wu et&#x20;al., 2014</xref>), active learning (<xref ref-type="bibr" rid="B34">Xiong et&#x20;al., 2014</xref>), DeepInteract (<xref ref-type="bibr" rid="B18">Sunil et&#x20;al., 2017</xref>), ConvsPPIS (<xref ref-type="bibr" rid="B37">Zhu et&#x20;al., 2020</xref>), PhosIDN (<xref ref-type="bibr" rid="B35">Yang et&#x20;al., 2021</xref>) and WinBinVec (<xref ref-type="bibr" rid="B1">Abdollahi et&#x20;al., 2021</xref>),&#x20;etc.</p>
<p>The above methods mainly use existing PPI data. However, current PPI data mainly generated by high-throughput or TAP-MS techniques (<xref ref-type="bibr" rid="B6">Berggard et&#x20;al., 2007</xref>), are often in presence of noise and incompleteness, and this unavoidably causes adverse effects on the prediction performance. Two main methods of <italic>network reconstruction</italic> and <italic>edge enrichment</italic> are proposed to effectively boost the prediction performance. Different strategies are used for network reconstruction or edge enrichment. For example, <xref ref-type="bibr" rid="B7">Bogdanov and Singh (2010)</xref> presented a network reconstruction approach by extracting functional neighborhood features using random walk with restart. <xref ref-type="bibr" rid="B9">Chua et&#x20;al. (2007)</xref> used weighting strategies to enrich PPI networks, and adopted a local prediction method to predict the functions of un-annotated proteins. <xref ref-type="bibr" rid="B33">Xiong et&#x20;al. (2013)</xref> applied collective classification to PPI networks with enriched edges to predict protein functions.</p>
<p>Although the above two types of approaches achieve promising performance improvements, there still lack comparative studies of the performance differences between network reconstruction and edge enrichment. We do not still know which one is better in performance, or specifically, which one should be applied for different situations. Inspired by the question, we conducte a comprehensive comparison of two network transformation of network reconstruction and edge enrichment for boosting the performance of PPI network-based protein functional annotation. Concretely, we first use three different protein similarity metrics for network reconstruction and edge enrichment of PPI networks, and then evaluate the performance differences between the two transformed networks (network reconstruction and edge enrichment) and original networks on two real PPI datasets. The results of experiments demonstrate that edge enrichment work better than both network reconstruction and original networks. Moreover, for the edge enrichment of PPI networks, the sequence similarity outperformes both local and global similarity. More detailed work will be presented in later sections.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Similarity Metrics</title>
<p>As we point out above, the noise and incompleteness of PPI network data adversely affects the performance of protein functional annotation. Network reconstruction and edge enrichment are major approaches to improve PPI data quality. In this work, we carry out comparison study on these two approaches by reconstructing and enriching original networks using various protein similarity metrics, including sequence similarity, local similarity and global similarity. In what follows, we describe and discuss these similarity measures in detail.</p>
<sec id="s2-1-1">
<title>2.1.1 Protein Sequence Similarity</title>
<p>BLAST method (<xref ref-type="bibr" rid="B3">Altschul et&#x20;al., 1997</xref>) is used to measure the similarity between any two proteins in this study. The similarity of a given protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub> with other proteins is defined as<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>V</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <italic>S</italic>
<sub>
<italic>x</italic>,<italic>i</italic>
</sub> is the similarity score between the pair of proteins <italic>V</italic>
<sub>
<italic>x</italic>
</sub> and <italic>V</italic>
<sub>
<italic>i</italic>
</sub>. Due to ignoring self-similarity, <italic>S</italic>
<sub>
<italic>x</italic>,<italic>i</italic>
</sub> &#x3d; 0 is set when <italic>x</italic>&#x20;&#x3d;&#x20;<italic>i</italic>.</p>
</sec>
<sec id="s2-1-2">
<title>2.1.2 Local Similarity Indices</title>
<p>We consider three kinds of local similarity indices, including <italic>Common Neighbors</italic> (CN), <italic>Jaccard Index</italic> and Functional Similarity&#x20;(FS).</p>
<p>
<italic>Common Neighbors</italic>. Given nodes <italic>u</italic> and <italic>v</italic>, their neighboring sets are <italic>N</italic>
<sub>
<italic>u</italic>
</sub> and <italic>N</italic>
<sub>
<italic>v</italic>
</sub>, respectively. The CN is defined as the neighborhood overlap of the nodes (<xref ref-type="bibr" rid="B16">Newman, 2001</xref>). The more identical neighbors two nodes have, the higher the CN value is. The measure of CN is as follows:<disp-formula id="e2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>
<italic>Jaccard Index</italic>. Given nodes <italic>u</italic> and <italic>v</italic> and their corresponding neighboring sets of <italic>N</italic>
<sub>
<italic>u</italic>
</sub> and <italic>N</italic>
<sub>
<italic>v</italic>
</sub>, Jaccard index is used to measure the similarity between the <italic>N</italic>
<sub>
<italic>u</italic>
</sub> and <italic>N</italic>
<sub>
<italic>v</italic>
</sub> sets, and it is calculated as:<disp-formula id="e3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>J</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>d</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x222a;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>
<italic>Functional Similarity</italic> (FS). For a PPI network, FS index was first used to measure the similarity of any pair of proteins (<xref ref-type="bibr" rid="B10">Chua et&#x20;al., 2006</xref>), and it is defined as follows:<disp-formula id="e4">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>S</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>max</mml:mi>
<mml:mfenced open="" close=")">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>v</mml:mi>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>, and by using the <italic>&#x3bb;</italic>
<sub>
<italic>u</italic>,<italic>v</italic>
</sub> factor, similarity weights between protein pairs are penalized when their common neighbors are too few. <italic>n</italic>
<sub>
<italic>avg</italic>
</sub> is the average number of close neighbors that each node has in the network. In a weighted PPI network, the labeled weights of edges mean interaction confidences between pairs of proteins. Thus, we can modify the FS index to take into account the confidence of each interaction. The extended FS index for weighted PPI networks, named FS.R, is defined as follows:<disp-formula id="e5">
<mml:math id="m6">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mi>S</mml:mi>
<mml:mo>.</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn mathvariant="normal">2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mfrac>
<mml:mrow>
<mml:mn mathvariant="normal">2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2229;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>v</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>u</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(5)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-3">
<title>2.1.3 Global Similarity Indices</title>
<p>Two global similarity indices are considered in this paper, they are Katz index and random walk with restart.</p>
<p>
<italic>Katz Index</italic>. This index is proposed by <xref ref-type="bibr" rid="B15">L&#xfc; and Zhou (2011)</xref>. It sums the set of paths directly and deals with the paths by length so that the shorter paths get more weights. Formally,<disp-formula id="e6">
<mml:math id="m7">
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>K</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>z</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>&#x22c5;</mml:mo>
<mml:mfenced open="|" close="|">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x3e;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2026;</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf2">
<mml:math id="m8">
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>v</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x3e;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the set of the paths, which connect the nodes of <italic>u</italic> and <italic>v</italic> with a path length of L. The parameter of <italic>&#x3b2;</italic> controls the path weights.</p>
<p>
<italic>Random Walk with Restart (RWR)</italic>. <xref ref-type="bibr" rid="B28">Tong et&#x20;al. (2008)</xref> used RWR index to measure the relevance score between node <italic>j</italic> and node <italic>i</italic> in a PPI network. Given the adjacency matrix <italic>W</italic>
<sub>
<italic>n</italic>,<italic>n</italic>
</sub> of a PPI network, a random walker transmits from the starting node <italic>i</italic> to one of its neighbors at random with probability c, and returns to the node <italic>i</italic> with the probability 1 &#x2212; <italic>c</italic>. Finally, the walker will stay stably at node <italic>j</italic> with probability <italic>R</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub>. The steady-state probability <italic>R</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> is defined as RWR index. We have<disp-formula id="e7">
<mml:math id="m9">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>c</mml:mi>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf3">
<mml:math id="m10">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is the starting vector, the <italic>i</italic>th element is 1 and the other elements are 0. <inline-formula id="inf4">
<mml:math id="m11">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is a weighted matrix. For an unweighted network, <inline-formula id="inf5">
<mml:math id="m12">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>/</mml:mo>
<mml:mi>m</mml:mi>
</mml:math>
</inline-formula> (where <italic>m</italic> is the number of neighbors that node <italic>i</italic> has) if <italic>i</italic> and <italic>j</italic> are connected, and <inline-formula id="inf6">
<mml:math id="m13">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:math>
</inline-formula> otherwise. For a weighted network,<disp-formula id="e8">
<mml:math id="m14">
<mml:mfenced open="{" close="">
<mml:mrow>
<mml:mtable class="cases">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>i</mml:mi>
<mml:mi>f</mml:mi>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>i</mml:mi>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>d</mml:mi>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>j</mml:mi>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>.</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>.</mml:mo>
<mml:mspace width="1em"/>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-1-4">
<title>2.2 Network Reconstruction and Edge Enrichment</title>
<p>
<italic>Network reconstruction</italic> is carried out as follows: First, the similarity scores between protein pairs in the original PPI network are calculated according to the above similarity indexes. Next, some interactions are selected to reconstruct the PPI network based on the similarity scores. As in <xref ref-type="bibr" rid="B14">Liben-Nowell and Kleinberg (2007)</xref>, an appropriate score threshold is used such that the number of protein pairs with higher scores than the threshold is as same as possible to the interaction number of the original network. Then, a new network is formed by using the protein pairs with higher scores over the threshold. However, this approach may lead to absence of some proteins in the new network. Alternatively, for any node <italic>N</italic>
<sub>
<italic>i</italic>
</sub> in the original network, we first remove all its interactions. We find the top <italic>k</italic> neighbors most similar to the node <italic>N</italic>
<sub>
<italic>i</italic>
</sub>. Then, the <italic>k</italic> edges from the node <italic>N</italic>
<sub>
<italic>i</italic>
</sub> to its top <italic>k</italic> neighbors are created, and their similarity scores are used as edge weights in the new network. Thus, we have<disp-formula id="e9">
<mml:math id="m15">
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>
<italic>Edge enrichment</italic> is also performed in two steps as in network reconstruction, the only difference is that all interactions in the original network are preserved. An enriched network has two types of edges: <italic>explicit edges</italic> (old edges) and similarity-inferred edges (new edges). Here, there are two questions to be addressed: One is how to combine the edge weights with different semantics, and another is how many edges are added for each protein, that is, how to optimize the parameter <italic>k</italic> (see <xref ref-type="disp-formula" rid="e9">Eq. 9</xref>). The questions will be discussed in the following sections.</p>
</sec>
</sec>
<sec id="s2-2">
<title>2.3 Protein Function Prediction Approaches</title>
<p>In this study, protein function predictions on two real PPI datasets are performed using two different approaches.The first one is majority method, which is a local neighbor counting approach (<xref ref-type="bibr" rid="B19">Schwikowski et&#x20;al., 2000</xref>). The second is a global protein function prediction approach, which is common called <italic>collective classification</italic> (<xref ref-type="bibr" rid="B33">Xiong et&#x20;al., 2013</xref>). Details of this approach are presented in the following subsections.</p>
<table-wrap id="alg1" position="float">
<label>ALGORITHM 1</label>
<caption>
<p>Gibbs sampling</p>
</caption>
<table>
<tbody>
<tr>
<td>
<inline-graphic xlink:href="fgene-12-758131-fx1.tif"/>
</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2-3">
<title>2.4 Gibbs Sampling Based Collective Classification</title>
<p>Gibbs sampling (GS) includes two main processes of <italic>bootstrapping</italic> and <italic>iterative classification</italic> (<xref ref-type="bibr" rid="B20">Sen et&#x20;al., 2008</xref>). The pseudo-code is illustrated&#x20;below.</p>
<sec id="s2-3-1">
<title>2.4.1 Bootstrapping</title>
<p>The closer the proteins to each other, the more similar their functions become in a PPI network. For an unannotated protein, its probability distribution is estimated using a weighted voting method. In the original or reconstructed network, there is only one kind of annotated neighbors to vote. An unannotated protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub> has the corresponding explicit neighbors of <italic>N</italic>
<sub>
<italic>x</italic>
</sub> or <italic>k</italic> similarity-inferred neighbors. For the above neighbor sets, we have their edge weights as follows:<disp-formula id="e10">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
<label>(10)</label>
</disp-formula>
</p>
<p>The probability of <italic>V</italic>
<sub>
<italic>x</italic>
</sub> having the <italic>j</italic>th function <italic>F</italic>
<sub>
<italic>j</italic>
</sub> (<italic>V</italic>
<sub>
<italic>x</italic>
</sub>
<italic>F</italic>
<sub>
<italic>j</italic>
</sub>) is calculated as follows:<disp-formula id="e11">
<mml:math id="m17">
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(11)</label>
</disp-formula>where <inline-formula id="inf7">
<mml:math id="m18">
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf8">
<mml:math id="m19">
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are the normalizers:<disp-formula id="e12">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:mspace width="0.3333em" class="nbsp"/>
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(12)</label>
</disp-formula>
</p>
<p>However, in the enriched network, there are both old (explicit) and new (similarity-inferred) neighbors which need to be voted. So, the parameter <italic>&#x3bb;</italic> &#x2208; (0, 1) is used to combine the two types of different neighbors. Given a query protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub>, the <italic>V</italic>
<sub>
<italic>x</italic>
</sub>
<italic>F</italic>
<sub>
<italic>j</italic>
</sub> probability is calculated as follows:<disp-formula id="e13">
<mml:math id="m21">
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>Z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(13)</label>
</disp-formula>
</p>
<p>A higher <inline-formula id="inf9">
<mml:math id="m22">
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> value indicates a higher probability that protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub> is more likely to have <italic>j</italic>th function <italic>F</italic>
<sub>
<italic>j</italic>
</sub>. The <italic>V</italic>
<sub>
<italic>x</italic>
</sub>
<italic>F</italic>
<sub>
<italic>j</italic>
</sub> probability distribution is represented as:<disp-formula id="e14">
<mml:math id="m23">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4.2 Iterative Classification</title>
<p>Iterative classification has two main steps of burn-in and sampling. In burn-in period, iteration number is fixed, and <inline-formula id="inf10">
<mml:math id="m24">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> is updated in each iteration. In sampling period, we update <inline-formula id="inf11">
<mml:math id="m25">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> in each iteration, and also count how many times the <italic>j</italic>th function <italic>F</italic>
<sub>
<italic>j</italic>
</sub> for protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub> are sampled. Considering each protein with one or more functions, therefore, we define the most likely function of the protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub> as follow:<disp-formula id="e15">
<mml:math id="m26">
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
<label>(15)</label>
</disp-formula>where <inline-formula id="inf12">
<mml:math id="m27">
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> represents the <italic>j</italic>th most likely function of the protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub>, that is the jth-rank result. We further use <inline-formula id="inf13">
<mml:math id="m28">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> vector to record all ranking results in the <italic>i</italic>th iteration.<disp-formula id="e16">
<mml:math id="m29">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(16)</label>
</disp-formula>
</p>
<p>The matrix <italic>M</italic>
<sub>
<italic>x</italic>
</sub> with <italic>s</italic> rows and <italic>m</italic> columns is produced after running the predetermined <italic>s</italic> number of iterations.<disp-formula id="e17">
<mml:math id="m30">
<mml:msub>
<mml:mrow>
<mml:mi>M</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(17)</label>
</disp-formula>
</p>
<p>Finally, we obtain the required m-dimensional vector <inline-formula id="inf14">
<mml:math id="m31">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> for query protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub>:<disp-formula id="e18">
<mml:math id="m32">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>&#x20d7;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
<mml:mo>.</mml:mo>
</mml:math>
<label>(18)</label>
</disp-formula>where <inline-formula id="inf15">
<mml:math id="m33">
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the first ranked prediction in the <italic>i</italic>th column of&#x20;<italic>M</italic>
<sub>
<italic>x</italic>
</sub>.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and Discussion</title>
<sec id="s3-1">
<title>3.1 Data Preprocessing and Experimental Workflow</title>
<p>The two PPI datasets of A and B are used in our study. The datasets A and B are downloaded from the databases of BioGRID (<xref ref-type="bibr" rid="B25">Stark et&#x20;al., 2011</xref>) and STRING (<xref ref-type="bibr" rid="B26">Szklarczyk et&#x20;al., 2011</xref>), respectively. The datasets A and B are annotated as in <xref ref-type="bibr" rid="B4">Ashburner et&#x20;al. (2000)</xref>. The datasets in this study are based on Gene Ontology (GO) annotation. GO annotations consist of three basic namespaces: molecular function, biological process and cellular component. We construct one protein interaction network for each GO namespace using only physical interactions.Therefore, there are totally six PPI networks (three for <italic>S.cerevisiae</italic> and the other three for <italic>M.musculus</italic>) in Dataset A. For Dataset B, we construct two PPI networks (one for <italic>S.cerevisiae</italic> and another for <italic>M.musculus</italic>).More detailed information was listed in the supplementary material (<xref ref-type="sec" rid="s10">Supplementary Table&#x20;S1</xref>).</p>
<p>The comparison of the function prediction performance on the reconstructed and enriched networks with that on the original networks is first performed using the cross validation of leave-one-out method (LOOM). LOOM takes each protein in turn as a query protein, and carries out function prediction with the remaining proteins in the network. As the bootstrapping in <italic>Gibbs sampling</italic> based collective classification does not result in updating of the query protein, therefore we use the <italic>majority</italic> method to predict protein functions in LOOM cross validation. Then, the annotated protein proportion is changed from 10% to 90%, and the average performance of 10 experiments is reported for each of all proportions. The <italic>majority</italic> method is not suitable in this setting because it is a local neighbor counting approach and does not work well in sparsely-labeled network. Thus, the <italic>Gibbs sampling</italic> based collective classification is used to predict protein functions. The main hardware configuration of an Inter dual-core processor (3&#xa0;GHz) and 16GB RAM, with a Linux operating system, and Python 3.0 is as the programming environment for running the algorithms.</p>
<p>Finally, as in <xref ref-type="bibr" rid="B7">Bogdanov and Singh (2010)</xref>, the ratio of the number of <italic>true positive</italic> (TP) predictions to the number of <italic>false positive</italic>predictions (FP) is produced in the cross validation, i.e. TP/FP is used to assess prediction accuracy of PPI networks. We define the overall <italic>i</italic>th rank <italic>true positive</italic> (TP) as the number of proteins whose <italic>i</italic>th rank predicted function <inline-formula id="inf16">
<mml:math id="m34">
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is one of the true functions of protein <italic>V</italic>
<sub>
<italic>x</italic>
</sub>, and the overall <italic>i</italic>th rank <italic>false positive</italic> (FP) as the number of proteins whose <italic>i</italic>th rank predicted function <inline-formula id="inf17">
<mml:math id="m35">
<mml:msubsup>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is not one of the true functions of protein&#x20;<italic>V</italic>
<sub>
<italic>x</italic>
</sub>.</p>
</sec>
<sec id="s3-2">
<title>3.2 Similarity Index Selection and the Effect of the Parameters <italic>k</italic> and <italic>&#x3bb;</italic>
</title>
<p>In this study, in addition to sequence similarity, the PPI networks are reconstructed and enriched by using three local similarity indices (CN, Jaccard and FS)and two global similarity indices (Katz and RWR). In order to choose the best ones for the following experiments, the performance differences between the five similarity indices are evaluated over the two datasets of A and B. The experimental results over the dataset A are presented in <xref ref-type="sec" rid="s10">Supplementary Table S3</xref> and <xref ref-type="table" rid="T1">Table&#x20;1</xref>, and ones over the Dataset B listed in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. Using FS as the local similarity index and RWR as the global similarity index generally achieve the best performance. Hence, FS and RWR are selected as the local similarity index and global similarity index, respectively in the following experiments.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison of performance differences between similarity indices (Dataset A: <italic>M.musculus</italic>).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">Molecular function</th>
<th colspan="4" align="center">Biological process</th>
<th colspan="3" align="center">Cellular component</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">4th rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">&#xa0;&#xa0;Origin</td>
<td align="center">0.28</td>
<td align="center">0.12</td>
<td align="center">0.10</td>
<td align="center">0.39</td>
<td align="center">0.23</td>
<td align="center">0.13</td>
<td align="center">0.09</td>
<td align="center">1.63</td>
<td align="center">0.45</td>
<td align="center">0.24</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;CN</td>
<td align="center">0.21</td>
<td align="center">0.09</td>
<td align="center">0.07</td>
<td align="center">0.27</td>
<td align="center">0.22</td>
<td align="center">0.14</td>
<td align="center">0.09</td>
<td align="center">1.44</td>
<td align="center">0.47</td>
<td align="center">0.16</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;Jaccard</td>
<td align="center">0.30</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.11</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">
<bold>0.30</bold>
</td>
<td align="center">0.129</td>
<td align="center">0.11</td>
<td align="center">1.94</td>
<td align="center">0.56</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;FS</td>
<td align="center">
<bold>0.33</bold>
</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">0.47</td>
<td align="center">0.28</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.12</bold>
</td>
<td align="center">
<bold>2.13</bold>
</td>
<td align="center">
<bold>0.61</bold>
</td>
<td align="center">
<bold>0.27</bold>
</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;CN&#x2b;</td>
<td align="center">0.27</td>
<td align="center">0.14</td>
<td align="center">0.10</td>
<td align="center">0.37</td>
<td align="center">0.26</td>
<td align="center">0.14</td>
<td align="center">0.11</td>
<td align="center">1.70</td>
<td align="center">0.54</td>
<td align="center">0.21</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;Jaccard&#x2b;</td>
<td align="center">0.35</td>
<td align="center">0.16</td>
<td align="center">0.12</td>
<td align="center">
<bold>0.54</bold>
</td>
<td align="center">
<bold>0.34</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.12</td>
<td align="center">2.03</td>
<td align="center">0.62</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;FS&#x2b;</td>
<td align="center">
<bold>0.38</bold>
</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">0.52</td>
<td align="center">0.30</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.69</bold>
</td>
<td align="center">
<bold>0.29</bold>
</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;Katz</td>
<td align="center">0.29</td>
<td align="center">0.13</td>
<td align="center">0.12</td>
<td align="center">0.45</td>
<td align="center">0.23</td>
<td align="center">
<bold>0.17</bold>
</td>
<td align="center">0.11</td>
<td align="center">1.70</td>
<td align="center">0.54</td>
<td align="center">0.28</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;RWR</td>
<td align="center">
<bold>0.32</bold>
</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">
<bold>0.13</bold>
</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">
<bold>0.26</bold>
</td>
<td align="center">0.16</td>
<td align="center">
<bold>0.12</bold>
</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.61</bold>
</td>
<td align="center">
<bold>0.30</bold>
</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;Katz&#x2b;</td>
<td align="center">0.31</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.14</td>
<td align="center">0.47</td>
<td align="center">0.26</td>
<td align="center">
<bold>0.19</bold>
</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">2.13</td>
<td align="center">0.59</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">&#xa0;&#xa0;RWR&#x2b;</td>
<td align="center">
<bold>0.35</bold>
</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.52</bold>
</td>
<td align="center">
<bold>0.28</bold>
</td>
<td align="center">0.17</td>
<td align="center">0.12</td>
<td align="center">
<bold>2.45</bold>
</td>
<td align="center">
<bold>0.64</bold>
</td>
<td align="center">
<bold>0.33</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Comparison of performance differences between similarity indices (Dataset B).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">
<italic>S.cerevisiae</italic>
</th>
<th colspan="3" align="center">
<italic>M.musculus</italic>
</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Origin</td>
<td align="center">2.23</td>
<td align="center">0.75</td>
<td align="center">0.49</td>
<td align="center">1.94</td>
<td align="center">1.49</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">CN</td>
<td align="center">1.50</td>
<td align="center">0.54</td>
<td align="center">0.29</td>
<td align="center">1.28</td>
<td align="center">0.69</td>
<td align="center">0.43</td>
</tr>
<tr>
<td align="left">Jaccard</td>
<td align="center">1.55</td>
<td align="center">0.62</td>
<td align="center">0.39</td>
<td align="center">1.51</td>
<td align="center">1.13</td>
<td align="center">
<bold>0.79</bold>
</td>
</tr>
<tr>
<td align="left">FS</td>
<td align="center">
<bold>1.70</bold>
</td>
<td align="center">
<bold>0.64</bold>
</td>
<td align="center">
<bold>0.41</bold>
</td>
<td align="center">
<bold>1.56</bold>
</td>
<td align="center">
<bold>1.22</bold>
</td>
<td align="center">0.75</td>
</tr>
<tr>
<td align="left">CN&#x2b;</td>
<td align="center">1.85</td>
<td align="center">0.67</td>
<td align="center">0.43</td>
<td align="center">1.63</td>
<td align="center">1.27</td>
<td align="center">0.72</td>
</tr>
<tr>
<td align="left">Jaccard&#x2b;</td>
<td align="center">1.95</td>
<td align="center">0.65</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">1.78</td>
<td align="center">1.33</td>
<td align="center">0.78</td>
</tr>
<tr>
<td align="left">FS&#x2b;</td>
<td align="center">
<bold>2.13</bold>
</td>
<td align="center">
<bold>0.72</bold>
</td>
<td align="center">0.47</td>
<td align="center">
<bold>1.92</bold>
</td>
<td align="center">
<bold>1.51</bold>
</td>
<td align="center">
<bold>0.81</bold>
</td>
</tr>
<tr>
<td align="left">Katz</td>
<td align="center">1.63</td>
<td align="center">0.62</td>
<td align="center">0.41</td>
<td align="center">1.70</td>
<td align="center">
<bold>1.33</bold>
</td>
<td align="center">0.75</td>
</tr>
<tr>
<td align="left">RWR</td>
<td align="center">
<bold>1.78</bold>
</td>
<td align="center">
<bold>0.67</bold>
</td>
<td align="center">
<bold>0.43</bold>
</td>
<td align="center">
<bold>1.78</bold>
</td>
<td align="center">1.27</td>
<td align="center">
<bold>0.79</bold>
</td>
</tr>
<tr>
<td align="left">Katz&#x2b;</td>
<td align="center">1.86</td>
<td align="center">0.64</td>
<td align="center">0.47</td>
<td align="center">1.86</td>
<td align="center">
<bold>1.51</bold>
</td>
<td align="center">0.79</td>
</tr>
<tr>
<td align="left">RWR&#x2b;</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.75</bold>
</td>
<td align="center">
<bold>0.52</bold>
</td>
<td align="center">
<bold>2.03</bold>
</td>
<td align="center">1.49</td>
<td align="center">
<bold>0.85</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The effect of two parameters on the performance of network reconstruction and edge enrichment are also examined in our study. The first one is the number of similarity-inferred edges <italic>k</italic>. The prediction performance on the Datasets of A and B is listed in <xref ref-type="sec" rid="s10">Supplementary Table S4</xref>, <xref ref-type="table" rid="T3">Table&#x20;3</xref>, and <xref ref-type="table" rid="T4">Table&#x20;4</xref>, with the varying values of <italic>k</italic>. For both the datasets A and B, experimental results show that BLAST roughly achieves the best performance by setting <italic>k</italic> &#x3d; 5. When the values of <italic>k</italic>&#x20;&#x3d; {10, 30, 50, 100} are used for FS and RWR, using <italic>k</italic>&#x20;&#x3d; 30 or <italic>k</italic>&#x20;&#x3d; 50 generally works best in most cases, and the overall performance is relatively robust for the reconstructed or enriched networks. Hence, in the following experiments, the parameter value of <italic>k</italic> is used as 5, 30, 30 for BLAST, FS and RWR, respectively.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The influence of the parameter of <italic>k</italic> (<italic>M.musculus</italic> in Dataset A).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">Molecular function</th>
<th colspan="4" align="center">Biological process</th>
<th colspan="3" align="center">Cellular component</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">4th rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Origin</td>
<td align="center">0.28</td>
<td align="center">0.12</td>
<td align="center">0.10</td>
<td align="center">0.39</td>
<td align="center">0.23</td>
<td align="center">0.13</td>
<td align="center">0.09</td>
<td align="center">1.63</td>
<td align="center">0.45</td>
<td align="center">0.24</td>
</tr>
<tr>
<td align="left">BLAST 1</td>
<td align="center">0.34</td>
<td align="center">0.19</td>
<td align="center">0.09</td>
<td align="center">0.30</td>
<td align="center">0.16</td>
<td align="center">0.08</td>
<td align="center">0.04</td>
<td align="center">0.79</td>
<td align="center">0.34</td>
<td align="center">0.13</td>
</tr>
<tr>
<td align="left">BLAST 5</td>
<td align="center">0.43</td>
<td align="center">0.26</td>
<td align="center">0.13</td>
<td align="center">
<bold>0.45</bold>
</td>
<td align="center">
<bold>0.22</bold>
</td>
<td align="center">0.12</td>
<td align="center">0.08</td>
<td align="center">
<bold>0.98</bold>
</td>
<td align="center">
<bold>0.38</bold>
</td>
<td align="center">0.18</td>
</tr>
<tr>
<td align="left">BLAST 10</td>
<td align="center">
<bold>0.45</bold>
</td>
<td align="center">
<bold>0.27</bold>
</td>
<td align="center">0.11</td>
<td align="center">0.43</td>
<td align="center">0.18</td>
<td align="center">
<bold>0.13</bold>
</td>
<td align="center">0.09</td>
<td align="center">0.96</td>
<td align="center">0.35</td>
<td align="center">0.17</td>
</tr>
<tr>
<td align="left">BLAST 15</td>
<td align="center">0.41</td>
<td align="center">0.21</td>
<td align="center">
<bold>0.13</bold>
</td>
<td align="center">0.42</td>
<td align="center">0.20</td>
<td align="center">0.11</td>
<td align="center">
<bold>0.09</bold>
</td>
<td align="center">0.92</td>
<td align="center">0.33</td>
<td align="center">
<bold>0.19</bold>
</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;1</td>
<td align="center">0.39</td>
<td align="center">0.24</td>
<td align="center">0.15</td>
<td align="center">0.47</td>
<td align="center">0.26</td>
<td align="center">0.15</td>
<td align="center">0.12</td>
<td align="center">1.71</td>
<td align="center">0.49</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;5</td>
<td align="center">0.47</td>
<td align="center">
<bold>0.29</bold>
</td>
<td align="center">
<bold>0.23</bold>
</td>
<td align="center">
<bold>0.56</bold>
</td>
<td align="center">0.30</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">
<bold>2.02</bold>
</td>
<td align="center">
<bold>0.67</bold>
</td>
<td align="center">0.32</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;10</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">0.24</td>
<td align="center">0.21</td>
<td align="center">0.54</td>
<td align="center">0.32</td>
<td align="center">0.14</td>
<td align="center">0.11</td>
<td align="center">1.94</td>
<td align="center">0.58</td>
<td align="center">0.29</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;15</td>
<td align="center">0.46</td>
<td align="center">0.27</td>
<td align="center">0.20</td>
<td align="center">0.49</td>
<td align="center">
<bold>0.34</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.12</td>
<td align="center">1.86</td>
<td align="center">0.62</td>
<td align="center">
<bold>0.33</bold>
</td>
</tr>
<tr>
<td align="left">FS 10</td>
<td align="center">0.30</td>
<td align="center">0.14</td>
<td align="center">0.12</td>
<td align="center">0.42</td>
<td align="center">0.24</td>
<td align="center">0.14</td>
<td align="center">0.09</td>
<td align="center">1.71</td>
<td align="center">0.54</td>
<td align="center">0.23</td>
</tr>
<tr>
<td align="left">FS 30</td>
<td align="center">0.33</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">
<bold>0.47</bold>
</td>
<td align="center">0.28</td>
<td align="center">0.16</td>
<td align="center">0.12</td>
<td align="center">
<bold>2.13</bold>
</td>
<td align="center">0.61</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">FS 50</td>
<td align="center">
<bold>0.35</bold>
</td>
<td align="center">0.16</td>
<td align="center">0.17</td>
<td align="center">0.46</td>
<td align="center">
<bold>0.30</bold>
</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">0.14</td>
<td align="center">2.04</td>
<td align="center">
<bold>0.64</bold>
</td>
<td align="center">
<bold>0.28</bold>
</td>
</tr>
<tr>
<td align="left">FS 100</td>
<td align="center">0.32</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">0.12</td>
<td align="center">0.45</td>
<td align="center">0.27</td>
<td align="center">0.17</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">1.95</td>
<td align="center">0.57</td>
<td align="center">0.26</td>
</tr>
<tr>
<td align="left">FS&#x2b;10</td>
<td align="center">0.32</td>
<td align="center">0.14</td>
<td align="center">0.12</td>
<td align="center">0.26</td>
<td align="center">0.14</td>
<td align="center">0.14</td>
<td align="center">0.10</td>
<td align="center">1.95</td>
<td align="center">0.58</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="left">FS&#x2b;30</td>
<td align="center">0.39</td>
<td align="center">0.16</td>
<td align="center">0.15</td>
<td align="center">0.52</td>
<td align="center">
<bold>0.30</bold>
</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.14</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.70</bold>
</td>
<td align="center">
<bold>0.30</bold>
</td>
</tr>
<tr>
<td align="left">FS&#x2b;50</td>
<td align="center">
<bold>0.41</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.11</td>
<td align="center">
<bold>0.54</bold>
</td>
<td align="center">0.24</td>
<td align="center">0.14</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">2.21</td>
<td align="center">0.67</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="left">FS&#x2b;100</td>
<td align="center">0.38</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.50</td>
<td align="center">0.27</td>
<td align="center">0.16</td>
<td align="center">0.13</td>
<td align="center">2.07</td>
<td align="center">0.64</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">RWR 10</td>
<td align="center">0.25</td>
<td align="center">0.13</td>
<td align="center">0.11</td>
<td align="center">0.41</td>
<td align="center">0.21</td>
<td align="center">0.14</td>
<td align="center">0.09</td>
<td align="center">1.86</td>
<td align="center">0.54</td>
<td align="center">0.24</td>
</tr>
<tr>
<td align="left">RWR 30</td>
<td align="center">
<bold>0.32</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.13</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">
<bold>0.26</bold>
</td>
<td align="center">0.16</td>
<td align="center">0.11</td>
<td align="center">2.23</td>
<td align="center">
<bold>0.61</bold>
</td>
<td align="center">
<bold>0.30</bold>
</td>
</tr>
<tr>
<td align="left">RWR 50</td>
<td align="center">0.31</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.11</td>
<td align="center">0.47</td>
<td align="center">0.21</td>
<td align="center">
<bold>0.17</bold>
</td>
<td align="center">0.12</td>
<td align="center">
<bold>2.33</bold>
</td>
<td align="center">0.58</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">RWR 100</td>
<td align="center">0.29</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">0.44</td>
<td align="center">0.22</td>
<td align="center">0.16</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">2.12</td>
<td align="center">0.55</td>
<td align="center">0.30</td>
</tr>
<tr>
<td align="left">RWR&#x2b;10</td>
<td align="center">0.29</td>
<td align="center">0.14</td>
<td align="center">0.13</td>
<td align="center">0.47</td>
<td align="center">0.23</td>
<td align="center">0.15</td>
<td align="center">0.12</td>
<td align="center">2.13</td>
<td align="center">0.57</td>
<td align="center">0.29</td>
</tr>
<tr>
<td align="left">RWR&#x2b;30</td>
<td align="center">
<bold>0.35</bold>
</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.52</bold>
</td>
<td align="center">
<bold>0.28</bold>
</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">0.12</td>
<td align="center">
<bold>2.45</bold>
</td>
<td align="center">
<bold>0.64</bold>
</td>
<td align="center">0.33</td>
</tr>
<tr>
<td align="left">RWR&#x2b;50</td>
<td align="center">0.34</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.49</td>
<td align="center">0.28</td>
<td align="center">0.14</td>
<td align="center">
<bold>0.13</bold>
</td>
<td align="center">2.36</td>
<td align="center">0.62</td>
<td align="center">0.32</td>
</tr>
<tr>
<td align="left">RWR&#x2b;100</td>
<td align="center">0.31</td>
<td align="center">0.15</td>
<td align="center">0.15</td>
<td align="center">0.46</td>
<td align="center">0.25</td>
<td align="center">0.16</td>
<td align="center">0.12</td>
<td align="center">2.23</td>
<td align="center">0.58</td>
<td align="center">
<bold>0.36</bold>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>The effect of the parameter <italic>k</italic> (Dataset B).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">
<italic>S.cerevisiae</italic>
</th>
<th colspan="3" align="center">
<italic>M.musculus</italic>
</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Origin</td>
<td align="center">2.23</td>
<td align="center">0.75</td>
<td align="center">0.49</td>
<td align="center">1.94</td>
<td align="center">1.49</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">BLAST 1</td>
<td align="center">0.96</td>
<td align="center">0.37</td>
<td align="center">0.17</td>
<td align="center">1.28</td>
<td align="center">0.59</td>
<td align="center">0.35</td>
</tr>
<tr>
<td align="left">BLAST 5</td>
<td align="center">1.18</td>
<td align="center">
<bold>0.43</bold>
</td>
<td align="center">
<bold>0.28</bold>
</td>
<td align="center">1.63</td>
<td align="center">
<bold>0.75</bold>
</td>
<td align="center">0.45</td>
</tr>
<tr>
<td align="left">BLAST 10</td>
<td align="center">
<bold>1.21</bold>
</td>
<td align="center">0.39</td>
<td align="center">0.24</td>
<td align="center">
<bold>1.70</bold>
</td>
<td align="center">0.72</td>
<td align="center">0.41</td>
</tr>
<tr>
<td align="left">BLAST 15</td>
<td align="center">1.15</td>
<td align="center">0.42</td>
<td align="center">0.26</td>
<td align="center">1.56</td>
<td align="center">0.70</td>
<td align="center">
<bold>0.50</bold>
</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;1</td>
<td align="center">2.11</td>
<td align="center">0.64</td>
<td align="center">0.45</td>
<td align="center">2.15</td>
<td align="center">1.51</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;5</td>
<td align="center">
<bold>2.83</bold>
</td>
<td align="center">
<bold>0.82</bold>
</td>
<td align="center">0.64</td>
<td align="center">
<bold>2.45</bold>
</td>
<td align="center">
<bold>1.63</bold>
</td>
<td align="center">
<bold>0.87</bold>
</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;10</td>
<td align="center">2.57</td>
<td align="center">0.75</td>
<td align="center">
<bold>0.65</bold>
</td>
<td align="center">2.33</td>
<td align="center">1.57</td>
<td align="center">0.85</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;15</td>
<td align="center">2.40</td>
<td align="center">0.69</td>
<td align="center">0.62</td>
<td align="center">2.28</td>
<td align="center">1.49</td>
<td align="center">0.76</td>
</tr>
<tr>
<td align="left">FS 10</td>
<td align="center">1.53</td>
<td align="center">0.55</td>
<td align="center">0.38</td>
<td align="center">1.33</td>
<td align="center">1.06</td>
<td align="center">0.68</td>
</tr>
<tr>
<td align="left">FS 30</td>
<td align="center">1.72</td>
<td align="center">
<bold>0.64</bold>
</td>
<td align="center">
<bold>0.41</bold>
</td>
<td align="center">1.56</td>
<td align="center">
<bold>1.22</bold>
</td>
<td align="center">0.75</td>
</tr>
<tr>
<td align="left">FS 50</td>
<td align="center">
<bold>1.75</bold>
</td>
<td align="center">0.57</td>
<td align="center">0.38</td>
<td align="center">1.64</td>
<td align="center">1.19</td>
<td align="center">
<bold>0.79</bold>
</td>
</tr>
<tr>
<td align="left">FS 100</td>
<td align="center">1.63</td>
<td align="center">0.61</td>
<td align="center">0.37</td>
<td align="center">
<bold>1.68</bold>
</td>
<td align="center">1.18</td>
<td align="center">0.73</td>
</tr>
<tr>
<td align="left">FS&#x2b;10</td>
<td align="center">1.93</td>
<td align="center">0.65</td>
<td align="center">0.40</td>
<td align="center">1.85</td>
<td align="center">0.42</td>
<td align="center">0.79</td>
</tr>
<tr>
<td align="left">FS&#x2b;30</td>
<td align="center">
<bold>2.13</bold>
</td>
<td align="center">
<bold>0.72</bold>
</td>
<td align="center">0.47</td>
<td align="center">
<bold>1.92</bold>
</td>
<td align="center">
<bold>1.51</bold>
</td>
<td align="center">
<bold>0.81</bold>
</td>
</tr>
<tr>
<td align="left">FS&#x2b;50</td>
<td align="center">2.05</td>
<td align="center">0.70</td>
<td align="center">0.44</td>
<td align="center">1.83</td>
<td align="center">1.40</td>
<td align="center">0.76</td>
</tr>
<tr>
<td align="left">FS&#x2b;100</td>
<td align="center">1.90</td>
<td align="center">0.67</td>
<td align="center">
<bold>0.49</bold>
</td>
<td align="center">1.92</td>
<td align="center">1.45</td>
<td align="center">0.78</td>
</tr>
<tr>
<td align="left">RWR 10</td>
<td align="center">1.50</td>
<td align="center">0.57</td>
<td align="center">0.36</td>
<td align="center">1.57</td>
<td align="center">1.08</td>
<td align="center">0.69</td>
</tr>
<tr>
<td align="left">RWR 30</td>
<td align="center">
<bold>1.78</bold>
</td>
<td align="center">
<bold>0.67</bold>
</td>
<td align="center">0.43</td>
<td align="center">
<bold>1.78</bold>
</td>
<td align="center">1.27</td>
<td align="center">0.79</td>
</tr>
<tr>
<td align="left">RWR 50</td>
<td align="center">1.72</td>
<td align="center">0.63</td>
<td align="center">0.40</td>
<td align="center">1.69</td>
<td align="center">
<bold>1.31</bold>
</td>
<td align="center">0.74</td>
</tr>
<tr>
<td align="left">RWR 100</td>
<td align="center">1.70</td>
<td align="center">0.61</td>
<td align="center">
<bold>0.45</bold>
</td>
<td align="center">1.64</td>
<td align="center">1.17</td>
<td align="center">
<bold>0.82</bold>
</td>
</tr>
<tr>
<td align="left">RWR&#x2b;10</td>
<td align="center">2.00</td>
<td align="center">0.70</td>
<td align="center">0.46</td>
<td align="center">1.88</td>
<td align="center">1.40</td>
<td align="center">0.69</td>
</tr>
<tr>
<td align="left">RWR&#x2b;30</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">0.75</td>
<td align="center">
<bold>0.52</bold>
</td>
<td align="center">
<bold>2.03</bold>
</td>
<td align="center">
<bold>1.49</bold>
</td>
<td align="center">
<bold>0.85</bold>
</td>
</tr>
<tr>
<td align="left">RWR&#x2b;50</td>
<td align="center">2.11</td>
<td align="center">0.72</td>
<td align="center">0.49</td>
<td align="center">1.94</td>
<td align="center">1.43</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">RWR&#x2b;100</td>
<td align="center">1.94</td>
<td align="center">
<bold>0.81</bold>
</td>
<td align="center">0.48</td>
<td align="center">1.82</td>
<td align="center">1.45</td>
<td align="center">0.75</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The second parameter <italic>&#x3bb;</italic> dominates the tradeoff between explicit edges and similarity-inferred edges. Further, the effect of the parameter <italic>&#x3bb;</italic> is evaluated on the prediction performance when it varies from 0.1 to 0.9. The results on the Dataset A are listed in Supplementary material (see <xref ref-type="sec" rid="s10">Supplementary Table S5</xref>) and <xref ref-type="table" rid="T5">Table&#x20;5</xref>, and ones on the Dataset B in <xref ref-type="table" rid="T6">Table&#x20;6</xref>, respectively. Generally, the <italic>&#x3bb;</italic> value has a relatively small impact on prediction accuracy, unless it is too large or too small. In the following experiments, the <italic>&#x3bb;</italic> value is set uniformly at&#x20;0.7.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>The influence of the parameter <italic>&#x3bb;</italic> (<italic>M.musculus</italic> in Dataset A).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">Molecular function</th>
<th colspan="4" align="center">Biological process</th>
<th colspan="3" align="center">Cellular component</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">4th rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Origin</td>
<td align="center">0.28</td>
<td align="center">0.12</td>
<td align="center">0.10</td>
<td align="center">0.39</td>
<td align="center">0.23</td>
<td align="center">0.13</td>
<td align="center">0.09</td>
<td align="center">1.63</td>
<td align="center">0.45</td>
<td align="center">0.24</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.1</td>
<td align="center">0.38</td>
<td align="center">0.21</td>
<td align="center">0.11</td>
<td align="center">0.43</td>
<td align="center">0.24</td>
<td align="center">0.12</td>
<td align="center">0.13</td>
<td align="center">1.52</td>
<td align="center">0.41</td>
<td align="center">0.20</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.3</td>
<td align="center">0.40</td>
<td align="center">0.25</td>
<td align="center">0.18</td>
<td align="center">0.49</td>
<td align="center">0.29</td>
<td align="center">
<bold>0.17</bold>
</td>
<td align="center">0.11</td>
<td align="center">1.65</td>
<td align="center">0.54</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.5</td>
<td align="center">0.44</td>
<td align="center">
<bold>0.30</bold>
</td>
<td align="center">0.16</td>
<td align="center">0.537</td>
<td align="center">0.27</td>
<td align="center">0.15</td>
<td align="center">
<bold>0.15</bold>
</td>
<td align="center">1.85</td>
<td align="center">0.62</td>
<td align="center">0.28</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.7</td>
<td align="center">
<bold>0.47</bold>
</td>
<td align="center">0.29</td>
<td align="center">
<bold>0.23</bold>
</td>
<td align="center">
<bold>0.56</bold>
</td>
<td align="center">
<bold>0.32</bold>
</td>
<td align="center">0.16</td>
<td align="center">0.14</td>
<td align="center">
<bold>2.02</bold>
</td>
<td align="center">
<bold>0.67</bold>
</td>
<td align="center">
<bold>0.34</bold>
</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.9</td>
<td align="center">0.33</td>
<td align="center">0.16</td>
<td align="center">0.15</td>
<td align="center">0.42</td>
<td align="center">0.23</td>
<td align="center">0.13</td>
<td align="center">0.10</td>
<td align="center">1.76</td>
<td align="center">0.55</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.1</td>
<td align="center">0.31</td>
<td align="center">0.14</td>
<td align="center">0.12</td>
<td align="center">0.49</td>
<td align="center">0.24</td>
<td align="center">0.15</td>
<td align="center">0.13</td>
<td align="center">1.94</td>
<td align="center">0.59</td>
<td align="center">0.27</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.3</td>
<td align="center">0.35</td>
<td align="center">0.16</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">
<bold>0.53</bold>
</td>
<td align="center">
<bold>0.33</bold>
</td>
<td align="center">0.13</td>
<td align="center">0.10</td>
<td align="center">1.86</td>
<td align="center">0.68</td>
<td align="center">0.25</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.5</td>
<td align="center">0.37</td>
<td align="center">0.15</td>
<td align="center">0.13</td>
<td align="center">0.49</td>
<td align="center">0.31</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">0.11</td>
<td align="center">2.04</td>
<td align="center">0.63</td>
<td align="center">0.28</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.7</td>
<td align="center">
<bold>0.39</bold>
</td>
<td align="center">
<bold>0.16</bold>
</td>
<td align="center">0.15</td>
<td align="center">0.52</td>
<td align="center">0.30</td>
<td align="center">0.16</td>
<td align="center">
<bold>0.14</bold>
</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.70</bold>
</td>
<td align="center">
<bold>0.30</bold>
</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.9</td>
<td align="center">0.30</td>
<td align="center">0.13</td>
<td align="center">0.10</td>
<td align="center">0.42</td>
<td align="center">0.222</td>
<td align="center">0.14</td>
<td align="center">0.11</td>
<td align="center">1.86</td>
<td align="center">0.57</td>
<td align="center">0.26</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.1</td>
<td align="center">0.30</td>
<td align="center">0.13</td>
<td align="center">0.12</td>
<td align="center">0.47</td>
<td align="center">0.29</td>
<td align="center">0.16</td>
<td align="center">0.12</td>
<td align="center">2.12</td>
<td align="center">0.59</td>
<td align="center">0.28</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.3</td>
<td align="center">0.33</td>
<td align="center">0.15</td>
<td align="center">0.14</td>
<td align="center">0.50</td>
<td align="center">
<bold>0.31</bold>
</td>
<td align="center">0.11</td>
<td align="center">0.08</td>
<td align="center">2.22</td>
<td align="center">0.64</td>
<td align="center">0.32</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.5</td>
<td align="center">0.35</td>
<td align="center">0.13</td>
<td align="center">
<bold>0.17</bold>
</td>
<td align="center">0.50</td>
<td align="center">0.24</td>
<td align="center">0.16</td>
<td align="center">0.10</td>
<td align="center">2.32</td>
<td align="center">
<bold>0.74</bold>
</td>
<td align="center">0.30</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.7</td>
<td align="center">
<bold>0.37</bold>
</td>
<td align="center">
<bold>0.17</bold>
</td>
<td align="center">0.14</td>
<td align="center">
<bold>0.52</bold>
</td>
<td align="center">0.28</td>
<td align="center">
<bold>0.18</bold>
</td>
<td align="center">
<bold>0.12</bold>
</td>
<td align="center">
<bold>2.45</bold>
</td>
<td align="center">0.64</td>
<td align="center">
<bold>0.33</bold>
</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.9</td>
<td align="center">0.30</td>
<td align="center">0.13</td>
<td align="center">0.10</td>
<td align="center">0.43</td>
<td align="center">0.26</td>
<td align="center">0.14</td>
<td align="center">0.10</td>
<td align="center">1.94</td>
<td align="center">0.57</td>
<td align="center">0.27</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>The influence of the parameter <italic>&#x3bb;</italic> (Dataset B).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Indices</th>
<th colspan="3" align="center">
<italic>S.cerevisiae</italic>
</th>
<th colspan="3" align="center">
<italic>M.musculus</italic>
</th>
</tr>
<tr>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
<th align="center">1st rank</th>
<th align="center">2nd rank</th>
<th align="center">3rd rank</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Origin</td>
<td align="center">2.23</td>
<td align="center">0.75</td>
<td align="center">0.49</td>
<td align="center">1.94</td>
<td align="center">1.49</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.1</td>
<td align="center">1.56</td>
<td align="center">0.63</td>
<td align="center">0.41</td>
<td align="center">1.76</td>
<td align="center">1.28</td>
<td align="center">0.72</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.3</td>
<td align="center">1.89</td>
<td align="center">0.70</td>
<td align="center">0.58</td>
<td align="center">2.01</td>
<td align="center">1.44</td>
<td align="center">0.78</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.5</td>
<td align="center">2.56</td>
<td align="center">0.75</td>
<td align="center">
<bold>0.66</bold>
</td>
<td align="center">2.34</td>
<td align="center">1.37</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.7</td>
<td align="center">
<bold>2.83</bold>
</td>
<td align="center">
<bold>0.82</bold>
</td>
<td align="center">0.64</td>
<td align="center">
<bold>2.45</bold>
</td>
<td align="center">
<bold>1.63</bold>
</td>
<td align="center">
<bold>0.87</bold>
</td>
</tr>
<tr>
<td align="left">BLAST&#x2b;0.9</td>
<td align="center">2.36</td>
<td align="center">0.79</td>
<td align="center">0.56</td>
<td align="center">2.12</td>
<td align="center">1.51</td>
<td align="center">0.85</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.1</td>
<td align="center">1.86</td>
<td align="center">0.66</td>
<td align="center">0.42</td>
<td align="center">1.70</td>
<td align="center">1.33</td>
<td align="center">0.74</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.3</td>
<td align="center">1.93</td>
<td align="center">0.64</td>
<td align="center">0.45</td>
<td align="center">1.86</td>
<td align="center">1.38</td>
<td align="center">0.81</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.5</td>
<td align="center">2.06</td>
<td align="center">0.69</td>
<td align="center">0.43</td>
<td align="center">
<bold>2.02</bold>
</td>
<td align="center">1.44</td>
<td align="center">
<bold>0.87</bold>
</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.7</td>
<td align="center">
<bold>2.13</bold>
</td>
<td align="center">0.72</td>
<td align="center">
<bold>0.47</bold>
</td>
<td align="center">1.92</td>
<td align="center">1.51</td>
<td align="center">0.84</td>
</tr>
<tr>
<td align="left">FS&#x2b;0.9</td>
<td align="center">1.99</td>
<td align="center">
<bold>0.75</bold>
</td>
<td align="center">0.41</td>
<td align="center">1.88</td>
<td align="center">
<bold>1.62</bold>
</td>
<td align="center">0.79</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.1</td>
<td align="center">1.82</td>
<td align="center">0.62</td>
<td align="center">0.42</td>
<td align="center">1.65</td>
<td align="center">1.38</td>
<td align="center">0.77</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.3</td>
<td align="center">1.94</td>
<td align="center">0.65</td>
<td align="center">0.48</td>
<td align="center">1.83</td>
<td align="center">1.44</td>
<td align="center">
<bold>0.83</bold>
</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.5</td>
<td align="center">2.02</td>
<td align="center">0.71</td>
<td align="center">
<bold>0.54</bold>
</td>
<td align="center">1.95</td>
<td align="center">1.46</td>
<td align="center">0.73</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.7</td>
<td align="center">
<bold>2.23</bold>
</td>
<td align="center">
<bold>0.75</bold>
</td>
<td align="center">0.52</td>
<td align="center">
<bold>2.03</bold>
</td>
<td align="center">
<bold>1.49</bold>
</td>
<td align="center">0.82</td>
</tr>
<tr>
<td align="left">RWR&#x2b;0.9</td>
<td align="center">2.12</td>
<td align="center">0.69</td>
<td align="center">0.47</td>
<td align="center">1.92</td>
<td align="center">1.43</td>
<td align="center">0.77</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 Performance Evaluation on Dataset A</title>
<p>The performance comparison of reconstructed and enriched networks with that of the original networks is first carried out by leave-one-out validation. The top protein function prediction is selected according to the average number of useful functions per protein in the PPI networks. Therefore, only the top 2 predictions are performed on the PPI networks of <italic>S.cerevisiae</italic> in the Dataset A, and the top 3 or 4 predictions are examined for <italic>M.musculus</italic> in Dataset&#x20;A.</p>
<p>Obviously, edge enrichment gains more accurate predictions than network reconstruction and original networks, due to the combination of explicit and implicit (similarity-inferred) edges (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref>). The results clearly indicate that edge enrichment indeed gains better prediction performance by adding similarity-inferred edges to PPI networks. BLAST-enriched networks always worke best, while BLAST-reconstructed networks always work worst. This is because BLAST-inferred edges are based on protein sequence information that is short in the original networks. The useful information in the original network greatly increases by adding BLAST-inferred edges, and consequently boosts prediction accuracy. However, in the reconstructed networks, the original PPI edges are put aside first, BLAST-reconstructed networks contain only protein sequence information, and thus performe worst. The experimental results also validate that FS-reconstructed networks and RWR-reconstructed networks work better than the original networks in most cases. This is because the reconstructed networks filter out noisy or spurious interactions in the original PPI networks.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The performance evaluation by leave-one-out validation over the PPI networks (Dataset A: <italic>S.cerevisiae</italic> and <italic>M.musculus</italic>). Here, the sub figures in the horizontal and vertical directions represent the experimental results for the PPI networks of different data sets and function types, respectively. Horizontally, the top three subplots represent ones on <italic>S.cerevisiae</italic>, and the bottom for ones on <italic>M.musculus</italic>. <bold>(A)</bold> and <bold>(D)</bold> Molecular function, <bold>(B)</bold> and <bold>(E)</bold> Biological process, <bold>(C)</bold> and <bold>(F)</bold> Cellular component.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g001.tif"/>
</fig>
<p>We further evaluate prediction accuracy of these three kinds of networks by using Gibbs Sampling in sparse-labeled PPI networks. Concretely, in PPI networks, the annotated protein proportion is changed from 0.1 to 0.9, and the remaining protein functions are predicted. For each proportion of the annotated proteins, the average prediction accuracy of running 10 experiments is presented on the PPI networks of <italic>S.cerevisiae</italic> (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>)and <italic>M.musculus</italic> (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>), respectively. The enrichment gains more accurate predictions than network reconstruction and original networks. The BLAST-enriched networks always work the best, while the BLAST-reconstructed networks always perform the worst. As expected, the experimental results also validate that FS-reconstructed networks and RWR-reconstructed networks generally performe better than the original networks. As the annotated protein proportion in the original networks increases, the prediction performance gets better for most networks, especially for the 1-st rank function. However, the prediction performance of the original network slightly declines as its annotated protein proportion increases (<xref ref-type="fig" rid="F3">Figure&#x20;3G,&#x20;H</xref>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>The performance evaluation over the sparsely-labeled networks (Dataset A: <italic>S.cerevisiae</italic>). Here, the sub figures in the horizontal and vertical directions represent the experimental results for the PPI networks of different function types and rank predicted functions, respectively. Horizontally, the top two subplots represent ones of molecular function, the middle for ones of biological process, and the bottom for ones of cellular component <bold>(A)</bold>, <bold>(C)</bold> and <bold>(E)</bold> first rank predicted function, <bold>(B)</bold>, <bold>(D)</bold> and <bold>(F)</bold> second rank predicted function.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The performance evaluation over the sparsely-labeled networks (Dataset A: <italic>M.musculus</italic>). Here, the sub figures in the horizontal and vertical directions represent the experimental results for the PPI networks of different function types and rank predicted functions, respectively. Horizontally, the top three subplots represent ones on the PPI networks of molecular function, the middle for ones of biological process, and the bottom for ones of cellular component <bold>(A)</bold>, <bold>(D)</bold> and <bold>(G)</bold> first rank predicted function, <bold>(B)</bold>, <bold>(E)</bold> and <bold>(H)</bold> second rank predicted function, <bold>(C)</bold>, <bold>(F)</bold> and <bold>(I)</bold> third rank predicted function.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g003.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Performance Evaluation on Dataset B</title>
<p>As above, the performance of reconstructed and enriched networks is first compared with that of the original networks by leave-one-out validation. Here, the top 3 protein function predictions are considered for both PPI networks of <italic>S. cerevisiae</italic> and <italic>M. musculus</italic>. As expected, edge enrichment gaines higher accurate predictions than network reconstruction and original networks. Moreover, BLAST-enriched networks get best, while the BLAST-reconstructed networks always work worst (<xref ref-type="fig" rid="F4">Figure&#x20;4</xref>). The reasons are the same as for the dataset&#x20;A.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The performance evaluation by leave-one-out validation over the PPI networks (Dataset B: <italic>S.cerevisiae</italic> and <italic>M.musculus</italic>) <bold>(A)</bold> <italic>S.cerevisiae</italic> <bold>(B)</bold> <italic>M.musculus</italic>.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g004.tif"/>
</fig>
<p>Next, we evaluate the prediction performance of these networks in sparse-labeled conditions with the collective classification method. Similarly, the average prediction performance is generated over running 10 experiments, with the annotated-protein proportion varying from 0.1 to 0.9. Generally, the experimental results present a similar trend to the above for the dataset A (<xref ref-type="fig" rid="F5">Figure&#x20;5</xref>). However, FS-reconstructed networks and RWR-reconstructed networks do not outperform the original networks, due to the quality properties of the dataset itself. This is mainly because many informative interactions are deleted and the prediction performance is impaired when reconstructing the networks based on similarity.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The performance evaluation over the sparsely-labeled networks (Dataset B: <italic>S.cerevisiae</italic> and <italic>M.musculus</italic>). Here, the sub figures in the horizontal and vertical directions represent the experimental results for different data types and rank predicted functions, respectively. Horizontally, the top three subplots represent ones over the dataset of <italic>S.cerevisiae</italic>, and the bottom for ones of <italic>M.musculus</italic>. <bold>(A)</bold> and <bold>(D)</bold> first rank predicted function, <bold>(B)</bold> and <bold>(E)</bold> second rank predicted function, <bold>(C)</bold> and <bold>(F)</bold> third rank predicted function.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g005.tif"/>
</fig>
<p>To validate this point, 10% and 50% interactions of the original network of the dataset B are randomly selected to construct two sparse networks. The leave-one-out validation is then performed over the two sparse networks. The selection process have two steps: First, a random weight is assigned to each edge of the original network, and a minimum spanning tree is constructed on the new network. The randomness of the minimum spanning tree (<italic>MST</italic>) is ensured by the random weights, and <italic>MST</italic> ensures the connectivity of the sparse network. Second, the <italic>MST</italic> is expanded by adding a number of edges, which are randomly selected from the original network (but not already on the <italic>MST</italic>). Hence, the number of edges in the sparse network is equal to 10% or 50% of edges in the original network. The sparse network preserves the basic topological properties of the original network.</p>
<p>The final experimental results also confirm the above-mentioned phenomenon. For example, in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref>, the FS-reconstructed networks and the RWR-reconstructed networks work better than the original networks when the networks are very sparse (e.g. 10%). However, as the networks become denser, the FS-reconstructed networks and the RWR-reconstructed networks get worse than the original networks.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The performance evaluation by leave-one-out validation over the PPI networks (Dataset B: <italic>S.cerevisiae</italic> and <italic>M.musculus</italic>). Here, the sub figures in the horizontal and vertical directions represent the experimental results for different data types and rank predicted functions, respectively. Horizontally, the top three subplots represent ones over the dataset of <italic>S.cerevisiae</italic>, and the bottom for ones of <italic>M.musculus</italic> <bold>(A)</bold> and <bold>(D)</bold> first rank predicted function, <bold>(B)</bold> and <bold>(E)</bold> second rank predicted function, <bold>(C)</bold> and <bold>(F)</bold> third rank predicted function.</p>
</caption>
<graphic xlink:href="fgene-12-758131-g006.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="conclusion" id="s4">
<title>4 Conclusion</title>
<p>The systematic comparison of two network transformation approaches (network reconstruction and edge enrichment) is performed using three different protein similarity metrics (sequence similarity, local and global similarity). In summary, edge enrichment performs better than network reconstruction and original networks, while network reconstruction is more effective on relatively small and incomplete PPI networks. The edge enrichment of PPI networks based on sequence similarity outperforms those based on both local and global similarity. As the PPI networks become more and more complete, the effectiveness of both edge enrichment and network reconstruction will decrease or relatively decrease.</p>
<p>Research efforts will be further expanded in future, which include: 1) how the removal of noisy edges and addition of informative edges affect the prediction performance; 2) a combining approach that combines the best properties of all these indices is developed since the similarity indices considered here have different properties and performances.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: (1) Datasets A: BioGRID, <ext-link ext-link-type="uri" xlink:href="https://downloads.thebiogrid.org/BioGRID">https://downloads.thebiogrid.org/BioGRID</ext-link>. (2) Datasets A: STRING, <ext-link ext-link-type="uri" xlink:href="https://string-db.org">https://string-db.org</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>JZ, JG, WX and JG, JH designed and performed the experiments. JZ, JG, WX and YW analyzed the data. The manuscript was written by JZ, JG and WX and approved by all authors.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was supported by National Natural Science Foundation of China (NSFC) under Grants Nos. 41877009, 61772367, 62172300, U1936205 and by the Fundamental Research Funds for the Central Universities.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.758131/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2021.758131/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abdollahi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>P.-C.</given-names>
</name>
<name>
<surname>Chiang</surname>
<given-names>J.-H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Winbinvec: Cancer-Associated Protein-Protein Interaction Extraction and Identification of 20 Various Cancer Types and Metastasis Using Different Deep Learning Models</article-title>. <source>IEEE J.&#x20;Biomed. Health Inform.</source> <volume>25</volume>, <fpage>4052</fpage>&#x2013;<lpage>4063</lpage>. <pub-id pub-id-type="doi">10.1109/JBHI.2021.3093441</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adamcsek</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Palla</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Farkas</surname>
<given-names>I. J.</given-names>
</name>
<name>
<surname>Der&#xe9;nyi</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Vicsek</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>CFinder: Locating Cliques and Overlapping Modules in Biological Networks</article-title>. <source>Bioinformatics</source> <volume>22</volume>, <fpage>1021</fpage>&#x2013;<lpage>1023</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl039</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Madden</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sch&#xe4;ffer</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>1997</year>). <article-title>Gapped BLAST and PSI-BLAST: a New Generation of Protein Database Search Programs</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume>, <fpage>3389</fpage>&#x2013;<lpage>3402</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.17.3389</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ashburner</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ball</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Blake</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>Botstein</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Butler</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Cherry</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<etal/>
</person-group> (<year>2000</year>). <article-title>Gene Ontology: Tool for the Unification of Biology</article-title>. <source>Nat. Genet.</source> <volume>25</volume>, <fpage>25</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1038/75556</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barrell</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Dimmer</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Huntley</surname>
<given-names>R. P.</given-names>
</name>
<name>
<surname>Binns</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>O&#x27;Donovan</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Apweiler</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>The GOA Database in 2009--an Integrated Gene Ontology Annotation Resource</article-title>. <source>Nucleic Acids Res.</source> <volume>37</volume>, <fpage>D396</fpage>&#x2013;<lpage>D403</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn803</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bergg&#xe5;rd</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Linse</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>James</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Methods for the Detection and Analysis of Protein-Protein Interactions</article-title>. <source>Proteomics</source> <volume>7</volume>, <fpage>2833</fpage>&#x2013;<lpage>2842</lpage>. <pub-id pub-id-type="doi">10.1002/pmic.200700131</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bogdanov</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>A. K.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Molecular Function Prediction Using Neighborhood Features</article-title>. <source>Ieee/acm Trans. Comput. Biol. Bioinf.</source> <volume>7</volume>, <fpage>208</fpage>&#x2013;<lpage>217</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2009.81</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Protein Interface Complementarity and Gene Duplication Improve Link Prediction of Protein-Protein Interaction Network</article-title>. <source>Front. Genet.</source> <volume>11</volume>, <fpage>291</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2020.00291</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chua</surname>
<given-names>H. N.</given-names>
</name>
<name>
<surname>Sung</surname>
<given-names>W.-K.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>An Efficient Strategy for Extensive Integration of Diverse Biological Data for Protein Function Prediction</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>3364</fpage>&#x2013;<lpage>3373</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm520</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chua</surname>
<given-names>H. N.</given-names>
</name>
<name>
<surname>Sung</surname>
<given-names>W.-K.</given-names>
</name>
<name>
<surname>Wong</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Exploiting Indirect Neighbours and Topological Weight to Predict Protein Function from Protein-Protein Interactions</article-title>. <source>Bioinformatics</source> <volume>22</volume>, <fpage>1623</fpage>&#x2013;<lpage>1630</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl145</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dunn</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Dudbridge</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sanderson</surname>
<given-names>C. M.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>The Use of Edge-Betweenness Clustering to Investigate Biological Function in Protein Interaction Networks</article-title>. <source>BMC Bioinformatics</source> <volume>6</volume>, <fpage>39</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-6-39</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>W.-C.</given-names>
</name>
<name>
<surname>Cai</surname>
<given-names>Y.-D.</given-names>
</name>
<name>
<surname>Chou</surname>
<given-names>K.-C.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Predicting Functions of Proteins in Mouse Based on Weighted Protein-Protein Interaction Network and Protein Hybrid Properties</article-title>. <source>PLOS ONE</source> <volume>6</volume>, <fpage>e14556</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0014556</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuchaiev</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Milenkovi&#x107;</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Memi&#x161;evi&#x107;</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Hayes</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Pr&#x17e;ulj</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Topological Network Alignment Uncovers Biological Function and Phylogeny</article-title>. <source>J.&#x20;R. Soc. Interf.</source> <volume>7</volume>, <fpage>1341</fpage>&#x2013;<lpage>1354</lpage>. <pub-id pub-id-type="doi">10.1098/rsif.2010.0063</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liben-Nowell</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Kleinberg</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>The Link-Prediction Problem for Social Networks</article-title>. <source>J.&#x20;Am. Soc. Inf. Sci.</source> <volume>58</volume>, <fpage>1019</fpage>&#x2013;<lpage>1031</lpage>. <pub-id pub-id-type="doi">10.1002/asi.20591</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>L&#xfc;</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Link Prediction in Complex Networks: A Survey</article-title>. <source>Physica A: Stat. Mech. its Appl.</source> <volume>390</volume>, <fpage>1150</fpage>&#x2013;<lpage>1170</lpage>. <pub-id pub-id-type="doi">10.1016/j.physa.2010.11.027</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Newman</surname>
<given-names>M. E. J.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Clustering and Preferential Attachment in Growing Networks</article-title>. <source>Phys. Rev. E</source> <volume>64</volume>, <fpage>025102</fpage>. <pub-id pub-id-type="doi">10.1103/PhysRevE.64.025102</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ofran</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Punta</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Rost</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Beyond Annotation Transfer by Homology: Novel Protein-Function Prediction Methods to Assist Drug Discovery</article-title>. <source>Drug Discov. Today</source> <volume>10</volume>, <fpage>1475</fpage>&#x2013;<lpage>1482</lpage>. <pub-id pub-id-type="doi">10.1016/S1359-6446(05)03621-4</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Patel</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tripathi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kumari</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Varadwaj</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deepinteract: Deep Neural Network Based Protein-Protein Interaction Prediction Tool</article-title>. <source>Cbio</source> <volume>12</volume>, <fpage>551</fpage>&#x2013;<lpage>557</lpage>. <pub-id pub-id-type="doi">10.2174/1574893611666160815150746</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schwikowski</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Uetz</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fields</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>A Network of Protein-Protein Interactions in Yeast</article-title>. <source>Nat. Biotechnol.</source> <volume>18</volume>, <fpage>1257</fpage>&#x2013;<lpage>1261</lpage>. <pub-id pub-id-type="doi">10.1038/82360</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sen</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Namata</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bilgic</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Getoor</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Galligher</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Eliassi-Rad</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Collective Classification in Network Data</article-title>. <source>AIMag</source> <volume>29</volume>, <fpage>93</fpage>&#x2013;<lpage>106</lpage>. <pub-id pub-id-type="doi">10.1609/aimag.v29i3.2157</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sharan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ulitsky</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Shamir</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Network&#x2010;based Prediction of Protein Function</article-title>. <source>Mol. Syst. Biol.</source> <volume>3</volume>, <fpage>88</fpage>. <pub-id pub-id-type="doi">10.1038/msb4100129</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Singh</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Berger</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Global Alignment of Multiple Protein Interaction Networks with Application to Functional Orthology Detection</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>105</volume>, <fpage>12763</fpage>&#x2013;<lpage>12768</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0806627105</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sleator</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Walsh</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>An Overview of In Silico Protein Function Prediction</article-title>. <source>Arch. Microbiol.</source> <volume>192</volume>, <fpage>151</fpage>&#x2013;<lpage>155</lpage>. <pub-id pub-id-type="doi">10.1007/s00203-010-0549-9</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Solava</surname>
<given-names>R. W.</given-names>
</name>
<name>
<surname>Michaels</surname>
<given-names>R. P.</given-names>
</name>
<name>
<surname>Milenkovi&#x107;</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Graphlet-based Edge Clustering Reveals Pathogen-Interacting Proteins</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>i480</fpage>&#x2013;<lpage>i486</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts376</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stark</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Breitkreutz</surname>
<given-names>B.-J.</given-names>
</name>
<name>
<surname>Chatr-Aryamontri</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Boucher</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Oughtred</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Livstone</surname>
<given-names>M. S.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>The Biogrid Interaction Database: 2011 Update</article-title>. <source>Nucleic Acids Res.</source> <volume>39</volume>, <fpage>D698</fpage>&#x2013;<lpage>D704</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkq1116</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Szklarczyk</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Franceschini</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kuhn</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Simonovic</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Roth</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Minguez</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>The String Database in 2011: Functional Interaction Networks of Proteins, Globally Integrated and Scored</article-title>. <source>Nucleic Acids Res.</source> <volume>39</volume>, <fpage>D561</fpage>&#x2013;<lpage>D568</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkq973</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>T&#xe4;ubig</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Buchner</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Griebsch</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Past: Fast Structure-Based Searching in the PDB</article-title>. <source>Nucleic Acids Res.</source> <volume>34</volume>, <fpage>W20</fpage>&#x2013;<lpage>W23</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkl273</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tong</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Faloutsos</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>J.-Y.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Random Walk with Restart: Fast Solutions and Applications</article-title>. <source>Knowl Inf. Syst.</source> <volume>14</volume>, <fpage>327</fpage>&#x2013;<lpage>346</lpage>. <pub-id pub-id-type="doi">10.1007/s10115-007-0094-2</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Waiho</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Afiqah&#x2010;Aleng</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Iryani</surname>
<given-names>M. T. M.</given-names>
</name>
<name>
<surname>Fazhan</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Protein-protein Interaction Network: an Emerging Tool for Understanding Fish Disease in Aquaculture</article-title>. <source>Rev. Aquacult.</source> <volume>13</volume>, <fpage>156</fpage>&#x2013;<lpage>177</lpage>. <pub-id pub-id-type="doi">10.1111/raq.12468</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wallace</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Laskowski</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Thornton</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>1996</year>). <article-title>Derivation of 3D Coordinate Templates for Searching Structural Databases: Application to Ser-His-Asp Catalytic Triads in the Serine Proteinases and Lipases</article-title>. <source>Protein Sci.</source> <volume>5</volume>, <fpage>1001</fpage>&#x2013;<lpage>1013</lpage>. <pub-id pub-id-type="doi">10.1002/pro.5560050603</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ng</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>S.-S.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Collective Prediction of Protein Functions from Protein-Protein Interaction Networks</article-title>. <source>BMC Bioinformatics</source> <volume>15</volume>, <fpage>S9</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-15-S2-S9</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Liao</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Comprehensive Review and Evaluation of Computational Methods for Identifying Protein Complexes from Protein-Protein Interaction Networks</article-title>. <source>Brief. Bioinform.</source> <volume>21</volume>, <fpage>1531</fpage>&#x2013;<lpage>1548</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbz085</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Protein Function Prediction by Collective Classification with Explicit and Implicit Edges in Protein-Protein Interaction Networks</article-title>. <source>BMC Bioinformatics</source> <volume>14</volume>, <fpage>S4</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-14-S12-S4</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xiong</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Active Learning for Protein Function Prediction in Protein-Protein Interaction Networks</article-title>. <source>Neurocomputing</source> <volume>145</volume>, <fpage>44</fpage>&#x2013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1016/j.neucom.2014.05.075</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>X.-M.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>PhosIDN: an Integrated Deep Neural Network for Improving Protein Phosphorylation Site Prediction by Combining Sequence and Protein-Protein Interaction Information</article-title>. <source>Bioinformatics</source>, <fpage>btab551</fpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab551</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Godzik</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>FATCAT: a Web Server for Flexible Structure Comparison and Structure Similarity Searching</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume>, <fpage>W582</fpage>&#x2013;<lpage>W585</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh430</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Du</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Convsppis: Identifying Protein-Protein Interaction Sites by an Ensemble Convolutional Neural Network with Feature Graph</article-title>. <source>Cbio</source> <volume>15</volume>, <fpage>368</fpage>&#x2013;<lpage>378</lpage>. <pub-id pub-id-type="doi">10.2174/1574893614666191105155713</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>