<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Bioinform.</journal-id>
<journal-title>Frontiers in Bioinformatics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Bioinform.</abbrev-journal-title>
<issn pub-type="epub">2673-7647</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">896295</article-id>
<article-id pub-id-type="doi">10.3389/fbinf.2022.896295</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Bioinformatics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>ContactPFP: Protein Function Prediction Using Predicted Contact Information</article-title>
<alt-title alt-title-type="left-running-head">Kagaya et al.</alt-title>
<alt-title alt-title-type="right-running-head">Contact Map-Based Function Prediction</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Kagaya</surname>
<given-names>Yuki</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1720894/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Flannery</surname>
<given-names>Sean T.</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Jain</surname>
<given-names>Aashish</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Kihara</surname>
<given-names>Daisuke</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/211526/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Biological Sciences</institution>, <institution>Purdue University</institution>, <addr-line>West Lafayette</addr-line>, <addr-line>IN</addr-line>, <country>United States</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Computer Science</institution>, <institution>Purdue University</institution>, <addr-line>West Lafayette</addr-line>, <addr-line>IN</addr-line>, <country>United States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1213731/overview">Andrzej Kloczkowski</ext-link>, The Research Institute at Nationwide Children&#x2019;s Hospital, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/979477/overview">Yaoqi Zhou</ext-link>, Griffith University, Australia</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1112712/overview">Castrense Savojardo</ext-link>, University of Bologna, Italy</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Daisuke Kihara, <email>dkihara@purude.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Protein Bioinformatics, a section of the journal Frontiers in Bioinformatics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>06</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>2</volume>
<elocation-id>896295</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>03</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>05</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Kagaya, Flannery, Jain and Kihara.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Kagaya, Flannery, Jain and Kihara</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Computational function prediction is one of the most important problems in bioinformatics as elucidating the function of genes is a central task in molecular biology and genomics. Most of the existing function prediction methods use protein sequences as the primary source of input information because the sequence is the most available information for query proteins. There are attempts to consider other attributes of query proteins. Among these attributes, the three-dimensional (3D) structure of proteins is known to be very useful in identifying the evolutionary relationship of proteins, from which functional similarity can be inferred. Here, we report a novel protein function prediction method, ContactPFP, which uses predicted residue-residue contact maps as input structural features of query proteins. Although 3D structure information is known to be useful, it has not been routinely used in function prediction because the 3D structure is not experimentally determined for many proteins. In ContactPFP, we overcome this limitation by using residue-residue contact prediction, which has become increasingly accurate due to rapid development in the protein structure prediction field. ContactPFP takes a query protein sequence as input and uses predicted residue-residue contact as a proxy for the 3D protein structure. To characterize how predicted contacts contribute to function prediction accuracy, we compared the performance of ContactPFP with several well-established sequence-based function prediction methods. The comparative study revealed the advantages and weaknesses of ContactPFP compared to contemporary sequence-based methods. There were many cases where it showed higher prediction accuracy. We examined factors that affected the accuracy of ContactPFP using several illustrative cases that highlight the strength of our method.</p>
</abstract>
<kwd-group>
<kwd>function prediction</kwd>
<kwd>residue contact prediction</kwd>
<kwd>gene function</kwd>
<kwd>functional genomics</kwd>
<kwd>protein structure</kwd>
<kwd>PFP</kwd>
</kwd-group>
<contract-num rid="cn001">R01GM123055 R01GM133840</contract-num>
<contract-num rid="cn002">DBI2003635 CMMI1825941 MCB1925643 DBI2146026</contract-num>
<contract-sponsor id="cn001">National Institutes of Health<named-content content-type="fundref-id">10.13039/100000002</named-content>
</contract-sponsor>
<contract-sponsor id="cn002">National Science Foundation<named-content content-type="fundref-id">10.13039/100000001</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Proteins are working molecules in a cell. Virtually all cellular functions are carried out mainly by proteins. Therefore, elucidating the biological function of proteins is a central problem in molecular biology, biochemistry, genetics, and genomics. Ultimately, the function of proteins needs to be determined by experiments. However, in the process of experimental elucidation of protein function, computational function prediction is very useful for guiding experiments by, for example, helping biologists construct hypotheses in designing experiments.</p>
<p>As sequencing the whole genome has become a standard experimental protocol for studying an organism, many protein sequences are now available in various databases (<xref ref-type="bibr" rid="B50">Sayers et al., 2021</xref>), and many of them remain unannotated. Thus, there is an increasing need for computational function prediction. Indeed, computational function prediction has been one of the most extensively studied topics in bioinformatics (<xref ref-type="bibr" rid="B20">Hawkins and Kihara, 2007</xref>). Conventionally, protein function annotation has been performed through sequence similarity search tools, which use BLAST (<xref ref-type="bibr" rid="B3">Altschul et al., 1990</xref>) or FASTA (<xref ref-type="bibr" rid="B35">Lipman and Pearson, 1985</xref>), and motif searches (<xref ref-type="bibr" rid="B6">Bairoch and Bucher, 1994</xref>; <xref ref-type="bibr" rid="B39">Mistry et al., 2021</xref>). In addition to such sequence-based methods (<xref ref-type="bibr" rid="B21">Hawkins et al., 2006</xref>; <xref ref-type="bibr" rid="B10">Chitale et al., 2009</xref>; <xref ref-type="bibr" rid="B24">Jain and Kihara, 2019</xref>), other approaches have been explored, which use omics-data (<xref ref-type="bibr" rid="B42">Obayashi et al., 2019</xref>; <xref ref-type="bibr" rid="B59">Szklarczyk et al., 2019</xref>), phylogenetic profiles (<xref ref-type="bibr" rid="B44">Pellegrini et al., 1999</xref>), and 3D structures of proteins (<xref ref-type="bibr" rid="B49">Sael and Kihara, 2010</xref>; <xref ref-type="bibr" rid="B47">Sael and Kihara, 2012</xref>; <xref ref-type="bibr" rid="B70">Zhu et al., 2015</xref>). As also observed in recent community-wide assessments for computational function prediction, the Critical Assessment of Function Annotation (CAFA), methods that combine different information sources by machine learning often showed relatively strong prediction performance (<xref ref-type="bibr" rid="B29">Khan et al., 2019</xref>; <xref ref-type="bibr" rid="B66">You et al., 2019</xref>).</p>
<p>In this work, we used protein 3D structure information for inferring the function of proteins. It has been long known that the 3D structures are better conserved than protein sequences during evolution (<xref ref-type="bibr" rid="B71">Chothia and Lesk, 1986</xref>), and thus they help capture distant functional relationships of proteins (<xref ref-type="bibr" rid="B11">Das et al., 2021</xref>). However, the 3D structure information has not been much used in practice in function prediction because the 3D structure has not been determined experimentally for many proteins. However, the situation has been changing due to recent progress in the protein structure prediction field, which has made significant improvements in amino acid residue contact and distance map prediction (<xref ref-type="bibr" rid="B16">Greener et al., 2019</xref>; <xref ref-type="bibr" rid="B63">Xu, 2019</xref>; <xref ref-type="bibr" rid="B25">Jain et al., 2021</xref>; <xref ref-type="bibr" rid="B36">Maddhuri Venkata Subramaniya et al., 2021</xref>). It may be noted that the accuracy of models by Alphafold (<xref ref-type="bibr" rid="B28">Jumper et al., 2021</xref>), the top-ranked structure prediction method in the recent Critical Assessment of techniques in protein Structure Prediction (CASP) (<xref ref-type="bibr" rid="B1">Abriata et al., 2019</xref>), often reach the level of experiments, such as X-ray crystallography. By using such a recent protein structure prediction method, it is now possible to compensate for the limited availability of the structural information of proteins. Thus, we now have an unprecedented opportunity for structure-based functional inference for nearly all proteins in genomes since their amino acid sequences are available, even if their 3D structures are not.</p>
<p>Here, we explore how protein structure information, particularly amino acid contact information, can contribute to the accuracy of function prediction. To do so, we developed a new protein function prediction method, ContactPFP. In ContactPFP, instead of performing sequence-based database search, a query protein is compared with proteins in a database in terms of predicted contact maps. Since an amino acid contact map is, in principle, sufficient to build a 3D structure of the protein, using contact maps is conceptually equivalent to considering 3D structure similarity. We benchmarked ContactPFP on a dataset of 9,642 proteins and compared its performance with sequence-based function prediction methods that performed among the top in CAFA. The benchmark revealed the strengths and weaknesses of ContactPFP. We report the performance of ContactPFP relative to several key parameters. Also, to characterize ContactPFP&#x2019;s performance, we discuss examples where ContactPFP showed its strengths and cases where ContactPFP did not perform as well as the sequence-based methods that were compared against.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Overview of the ContactPFP Method</title>
<p>
<xref ref-type="fig" rid="F1">Figure 1</xref> shows the workflow of ContactPFP. For a query protein sequence, ContactPFP constructs a multiple sequence alignment (MSA) using HHblits (<xref ref-type="bibr" rid="B55">Steinegger et al., 2019</xref>), that is, run against the Uniclust30 database (<xref ref-type="bibr" rid="B38">Mirdita et al., 2017</xref>) with a parameter set of &#x201c;-n 3 -id 99 -cov 50 -diff inf&#x201d;. Then, using the MSA, residue-residue distance prediction (a distance map) is computed for a query using trRosetta (<xref ref-type="bibr" rid="B64">Yang et al., 2020</xref>). The predicted distance map of the query is then compared with contact maps of proteins in the reference database using the GR-Align algorithm (<xref ref-type="bibr" rid="B37">Malod-Dognin and Pr&#x1e91;ulj, 2014</xref>). Since GR-Align compares contact maps of proteins, predicted distance maps were converted to contact maps, or contact graphs, where nodes represent amino acid residues and edges connect residue pairs that are closer than a distance cutoff value. As the distance cutoff values to define a contact, we used 8, 10, and 12&#xa0;&#x212b; between between C&#x3b2; atoms. For a given contact distance cutoff (e.g., 8&#xa0;&#x212b;), we define a residue pair as &#x201c;in contact&#x201d; if the probability that the pair has the cutoff distance or closer between each other is 0.5 or larger.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Overview of ContactPFP. From an input protein sequence, residue-residue contact information is predicted with trRosetta, which is represented as a graph. Then, the graph is compared with contact map graphs in a database using GR-Align. Proteins in the database are sorted by graph similarity to the query and GO terms are extracted from top hits.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g001.tif"/>
</fig>
<p>The reference database was constructed from Swiss-Prot (<xref ref-type="bibr" rid="B60">The UniProt Consortium, 2021</xref>). Sequences shorter than 20 residues and longer than 2000 residues, which made up 1.1% of the all sequences, were excluded. For the remaining 555,378 proteins (98.9%), contact graphs were computed as described above. Using GR-Align, the query contact graph is compared with all contact graphs in the reference database, which are then ranked by graph similarity to the query. In GR-Align, two contact graphs are aligned so that the similarity score of the graphs, which considers graphlet distribution similarity of mapped nodes, is maximized. This graphlet degree similarity can capture the local similarity of contact graphs. A graph similarity score by GR-Align ranges from 0 to 1.0, with 1.0 indicating an exact match in graphlet distributions. Proteins in the reference database that have a similarity score of 0.5 or higher by GR-Align were considered as hits. GO terms from hits were collected and weighted by the sum of the graph similarity scores of hits that have the GO terms. The score of a predicted GO term <italic>i</italic> is computed as follows:<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>O</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>n</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>s</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mi>G</mml:mi>
<mml:mi>O</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mi>G</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>h</mml:mi>
<mml:mo>_</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(Eq. 1)</label>
</disp-formula>where k is a hit (i.e., graph similarity score &#x3e; &#x3d; 0.5 to the query) that has the GO term <italic>i</italic> in its annotation. Finally, predicted GO terms for a query is normalized by the highest score among them, so that the most confident GO term has a score of 1.0.</p>
</sec>
<sec id="s2-2">
<title>2.2 Constructing the Function Annotation Database</title>
<p>GO terms for proteins in the reference database were compiled from 12 data sources. The primary database used was the UniProtKB/Swiss-Prot. GO terms with IEA (Inferred from Electronic Annotation) evidence code (<xref ref-type="bibr" rid="B7">Boutet et al., 2016</xref>) were also included because considering IEA gave a higher function prediction accuracy than excluding them in our previous works (<xref ref-type="bibr" rid="B10">Chitale et al., 2009</xref>; <xref ref-type="bibr" rid="B19">Hawkins et al., 2009</xref>). In addition to UniProtKB/Swiss-Prot, we integrated annotations from UniPathway (<xref ref-type="bibr" rid="B40">Morgat et al., 2012</xref>), TIGRFAMs (<xref ref-type="bibr" rid="B17">Haft et al., 2013</xref>), SMART (<xref ref-type="bibr" rid="B34">Letunic et al., 2021</xref>), Reactome (<xref ref-type="bibr" rid="B26">Jassal et al., 2020</xref>), PROSITE (<xref ref-type="bibr" rid="B53">Sigrist et al., 2013</xref>), ProDom (<xref ref-type="bibr" rid="B8">Bru et al., 2005</xref>), PRINTS (<xref ref-type="bibr" rid="B5">Attwood et al., 2012</xref>), PIRSF (<xref ref-type="bibr" rid="B41">Nikolskaya et al., 2006</xref>), Pfam (<xref ref-type="bibr" rid="B14">Finn et al., 2016</xref>), InterPro (<xref ref-type="bibr" rid="B13">Finn et al., 2017</xref>) and HAMAP (<xref ref-type="bibr" rid="B43">Pedruzzi et al., 2015</xref>). This database extension contributed an additional 8,727 GO terms to our reference database. We showed in our previous work that this annotation expansion has a positive effect on function prediction performance (<xref ref-type="bibr" rid="B30">Khan et al., 2015</xref>).</p>
</sec>
<sec id="s2-3">
<title>2.3 Constructing the Benchmark Dataset</title>
<p>We started from representative sequences in UniRef50 (<xref ref-type="bibr" rid="B58">Suzek et al., 2015</xref>) (22 October 2019 version). We filter out entries that do not fall within lengths of 100&#x2013;2000. We only considered the entry names that existed in the 11 November 2019 version of Swiss-Prot. We kept only entries with at least one experimentally verified GO Term in all three categories. To remove the potential redundancy of annotations in the dataset, we only kept one protein from homologous proteins from different organisms. For example, among the two entries, ZRT1_SCHPO and ZRT1_YEAST, which were originally included in the representative sequences, we kept only ZRT1_SCHPO. The ortholog proteins were identified by the common mnemonic protein identification code, e.g.,&#x201c;ZRTI&#x201d;. Further, to remove the sequence redundancy, we performed sequence clustering by MMseqs2 (<xref ref-type="bibr" rid="B56">Steinegger and S&#xf6;ding, 2017</xref>) with a 25% identity and a 80% coverage (--min-seq-id 0.25 -c 0.8). As a result, we had 9,642 sequences in the benchmark dataset.</p>
</sec>
<sec id="s2-4">
<title>2.4 Existing Methods Used as Reference</title>
<p>To characterize the performance of ContactPFP, we compared it with four sequence-based methods, PSI-BLAST (<xref ref-type="bibr" rid="B4">Altschul et al., 1997</xref>), PFP (<xref ref-type="bibr" rid="B19">Hawkins et al., 2009</xref>), ESG (<xref ref-type="bibr" rid="B10">Chitale et al., 2009</xref>), and Phylo-PFP (<xref ref-type="bibr" rid="B24">Jain and Kihara, 2019</xref>). PSI-BLAST is considered the baseline of function prediction methods. We selected PFP, ESG, and Phylo-PFP because they are sequence-based methods that performed well in CAFA challenges (<xref ref-type="bibr" rid="B45">Radivojac et al., 2013</xref>; <xref ref-type="bibr" rid="B27">Jiang et al., 2016</xref>; <xref ref-type="bibr" rid="B68">Zhou et al., 2019</xref>). Our group, who used these three methods, were among the best teams in the series of CAFA challenges. PFP, ESG, and Phylo-PFP use PSI-BLAST search results in different elaborate ways: In PFP, GO terms extracted from PSI-BLAST hits are scored by the sum of the negative logarithm of E-value of the hits. Thus, GO terms from hits with smaller E-values are scored higher. Up to 20,000 hits were considered. In ESG, the top hits of the first PSI-BLAST run are used to perform a second round of database search. In Phylo-PFP, retrieved sequences were ranked by considering both raw E-value and the edge distance on a phylogenetic tree constructed for the sequence. In ESG and Phylo-PFP, raw scores of predicted GO terms are normalized to a range between 0 and 1.0 by the highest score observed for the target protein.</p>
<p>These sequence-based methods identify the query itself as the top hit in a database search. To avoid taking GO terms from the query itself for PFP, ESG, and Phylo-PFP, we removed the query and all hits that had an E-value of 0.0 in the last round of PSI-BLAST before extracting GO terms. These excluded proteins were also removed from the hit list of ContactPFP. For PSI-BLAST, we extracted GO terms from the top 10 hits (except for the query itself and hits with 0&#xa0;E-value) in the third iteration of a PSI-BLAST run and assigned a score of 1.0 to all the predicted terms.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Effect of the Residue-Contact Definition and the Fold Similarity</title>
<p>To start with, we examined how two important hyperparameters in ContactPFP, the definition of residue-residue contacts and choices of top hits from a database search, affect the function prediction accuracy (<xref ref-type="fig" rid="F2">Figure 2</xref>). When constructing a graph from residue distance prediction, a choice of residue distance cutoff needs to be made. A larger distance connects more residue pairs making more edges in a contact graph, while a smaller cutoff would highlight densely connected domains (<xref ref-type="bibr" rid="B67">Yuan et al., 2012</xref>). We tested three distance cutoffs between C&#x3b2; atoms, 8, 10, and 12&#xa0;&#x212b;, to define residue-residue contacts. The latter parameter, the choice of top hits, decides which proteins from the database search to use for extracting GO terms for annotating the query protein. Increasing this similarity level reduces the number of hits to consider while decreasing the cutoff leads to an increased number of hits.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Influence of parameters on the prediction performance of ContactPFP. Parameters were examined that determine the definition of hits in the contact map database search. The y-axis shown is the average Fmax score computed for the four test sets in the four-fold cross validation. In each plot, three distance cutoffs, 8, 10, 12&#xa0;&#x212b;, were used that defined residue contacts. The bar indicates the standard deviation calculated from four-fold cross validation. <bold>(A)</bold> Raw contact map graph comparison score. From a database search result, we only considered retrieved proteins with a specified graph similarity score or higher. The average standard deviation was 0.002. <bold>(B)</bold> Selecting top N hits by the raw score. In this scheme, we only selected top N hits as specified on the x-axis regardless of their scores. The average standard deviation was 0.002. <bold>(C)</bold> Z-score of the contact map graph comparison score. In this scheme, we chose hits to consider by the Z-score of the graph similarity score relative to the score distribution of the entire reference database. The average standard deviation was 0.002.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g002.tif"/>
</fig>
<p>To examine the effect of the parameters, we performed a four-fold cross validation. In combination with the distance cutoffs, we examined three different ways to select top hits from a search (<xref ref-type="fig" rid="F2">Figure 2</xref>). The Fmax score shown in the panels in <xref ref-type="fig" rid="F2">Figure 2</xref> are the average values of the four test sets used in the cross validation. In <xref ref-type="fig" rid="F2">Figure 2A</xref>, we used the raw graph similarity score from GR-Align to select hits from a database search. Retrieved proteins in a search that have a similarity score lower than a specified cutoff were discarded. Among the four scores examined, 0.3, 0.5, 0.7, and 0.9, the highest Fmax score of 0.555 was observed when a graph similarity score of 0.7 was used in combination with a distance cutoff of 12&#xa0;&#x212b;. In <xref ref-type="fig" rid="F2">Figure 2B</xref>, we used the top N hits from a database search to extract GO terms regardless of their graph similarity score. The highest Fmax score, 0.638, was achieved when the first two hits were used (i.e., <italic>n</italic> &#x3d; 2) with a distance cutoff of 12&#xa0;&#x212b;. In the last panel, <xref ref-type="fig" rid="F2">Figure 2C</xref>, we considered the Z-score of the graph similarity score to select top hits. The highest Fmax score, 0.571, was achieved with a Z-score of 7 using a residue distance cutoff of 12&#xa0;&#x212b;. In each panel in <xref ref-type="fig" rid="F2">Figure 2</xref>, the standard deviations from the four-fold validation were small, 0.002, and the best parameter combinations were consistent across the four-fold.</p>
<p>Overall, the combination of 12&#xa0;&#x212b; for the distance cutoff and using the top 2 hits showed the best performance among the conditions tested. Therefore, we report the results with this condition in the subsequent sections. Regarding the contact distance cutoff, 12&#xa0;&#x212b; showed the best performance in all three panels, which is consistent with what the GR-Align paper reported (<xref ref-type="bibr" rid="B37">Malod-Dognin and Pr&#x1e91;ulj, 2014</xref>).</p>
</sec>
<sec id="s3-2">
<title>3.2 GO Term Prediction Performance of ContactPFP</title>
<p>We now report the overall performance of ContactPFP in <xref ref-type="table" rid="T1">Table 1</xref> in comparison to the other four methods. Two values are reported in <xref ref-type="table" rid="T1">Table 1</xref>. In terms of the average Fmax score (the left column), ContactPFP&#x2019;s performance was the second-highest, slightly lower than Phylo-PFP. ESG, PFP, PSI-BLAST followed in this order. Breakdown of the performance in the three GO categories (<xref ref-type="table" rid="T2">Table 2</xref>) showed essentially the same trend. ContactPFP was the second in Cellular Component, the third in Molecular Function, and the second in the Biological Process.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>The average Fmax score of ContactPFP and the other four methods on the benchmark dataset.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Method</th>
<th align="center">Fmax</th>
<th align="center">Wins by ContactPFP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">ContactPFP</td>
<td align="char" char=".">0.638</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">Phylo-PFP</td>
<td align="char" char=".">0.662</td>
<td align="char" char="(">5,357 (55.6%)</td>
</tr>
<tr>
<td align="left">ESG</td>
<td align="char" char=".">0.634</td>
<td align="char" char="(">5,452 (56.5%)</td>
</tr>
<tr>
<td align="left">PFP</td>
<td align="char" char=".">0.586</td>
<td align="char" char="(">5,940 (61.6%)</td>
</tr>
<tr>
<td align="left">PSI-BLAST</td>
<td align="char" char=".">0.574</td>
<td align="char" char="(">6,386 (66.2%)</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The count of benchmark proteins in which ContactPFP performed better than the other methods is shown in the third column. For ContactPFP, the top 2 hits from a search were used to extract GO terms.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>The average Fmax score and Smin score in the three GO categories.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Method</th>
<th colspan="4" align="center">Fmax</th>
<th colspan="4" align="center">Smin</th>
</tr>
<tr>
<th align="center">ALL</th>
<th align="center">CC</th>
<th align="center">MF</th>
<th align="center">BP</th>
<th align="center">ALL</th>
<th align="center">CC</th>
<th align="center">MF</th>
<th align="center">BP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">ContactPFP</td>
<td align="char" char=".">0.638</td>
<td align="char" char=".">0.718</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.606</td>
<td align="char" char=".">95.042</td>
<td align="char" char=".">16.294</td>
<td align="char" char=".">12.319</td>
<td align="char" char=".">66.341</td>
</tr>
<tr>
<td align="left">Phylo-PFP</td>
<td align="char" char=".">0.662</td>
<td align="char" char=".">0.75</td>
<td align="char" char=".">0.759</td>
<td align="char" char=".">0.641</td>
<td align="char" char=".">98.985</td>
<td align="char" char=".">15.067</td>
<td align="char" char=".">13.376</td>
<td align="char" char=".">70.48</td>
</tr>
<tr>
<td align="left">ESG</td>
<td align="char" char=".">0.634</td>
<td align="char" char=".">0.714</td>
<td align="char" char=".">0.746</td>
<td align="char" char=".">0.598</td>
<td align="char" char=".">106.368</td>
<td align="char" char=".">16.541</td>
<td align="char" char=".">13.227</td>
<td align="char" char=".">76.673</td>
</tr>
<tr>
<td align="left">PFP</td>
<td align="char" char=".">0.586</td>
<td align="char" char=".">0.698</td>
<td align="char" char=".">0.689</td>
<td align="char" char=".">0.562</td>
<td align="char" char=".">117.773</td>
<td align="char" char=".">18.055</td>
<td align="char" char=".">16.365</td>
<td align="char" char=".">82.991</td>
</tr>
<tr>
<td align="left">PSI-BLAST</td>
<td align="char" char=".">0.574</td>
<td align="char" char=".">0.655</td>
<td align="char" char=".">0.678</td>
<td align="char" char=".">0.544</td>
<td align="char" char=".">281.223</td>
<td align="char" char=".">44.163</td>
<td align="char" char=".">38.382</td>
<td align="char" char=".">198.695</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>CC, cellular component; MF, molecular function; BP, biological process.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>On the other hand, when predictions given to individual target proteins were compared between two methods (<xref ref-type="table" rid="T1">Table 1</xref>, the right column), more than half of the proteins (55.6%) had predictions with a higher Fmax score by ContactPFP than Phylo-PFP. ContactPFP also had more wins over ESG, PFP, and PSI-BLAST. Thus, in this head-to-head comparison, ContactPFP was the best. To understand how the performance of the methods differ, we showed the Fmax score of individual target proteins by ContactPFP and each of the other methods in scatter plots (<xref ref-type="fig" rid="F3">Figure 3</xref>). The scores seem not to distribute randomly. Rather, they show an interesting pattern of a &#x201c;mirror-imaged N-shape&#x201d;, where there are a substantial number of targets with around 1.0 Fmax as well as other targets with around 0.1 by ContactPFP. This distribution implies that ContactPFP may have both a characteristic strength and weakness when compared with the other methods.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Comparison of Fmax score of individual target proteins. To be precise, they are F-score of each protein using the score cutoff that yielded the Fmax score of the benchmark dataset. Each point represents a protein in the benchmark dataset. <bold>(A)</bold> Comparison between ContactPFP and Phylo-PFP; <bold>(B)</bold> Comparison between ContactPFP and ESG; <bold>(C)</bold> comparison between ContactPFP and PFP; <bold>(D)</bold> Comparison between ContactPFP vs. PSI-BLAST.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g003.tif"/>
</fig>
<p>In <xref ref-type="table" rid="T2">Table 2</xref>, we also presented the Smin score of the five methods. Smin evaluates remaining uncertainty/missing information from predicted GO terms (<xref ref-type="bibr" rid="B27">Jiang et al., 2016</xref>). The lower, the better prediction. In terms of Smin, ContactPFP is the best among the all five methods when all GO categories, MF, or BP are considered. ContactPFP was the second in the CC category following Phylo-PFP.</p>
</sec>
<sec id="s3-3">
<title>3.3 Effect of the Contact Prediction Accuracy</title>
<p>We examined how contact prediction accuracy affects to the GO prediction accuracy in ContactPFP. For this analysis, we used 1,029 targets in the benchmark dataset, which have an experimentally determined protein structure that covers more than 80% of the residues in the target protein. If we consider precision of all predicted contacts (<xref ref-type="fig" rid="F4">Figure 4A</xref>), all the targets fall into contact precision around 0.8 (the average: 0.801), and we found no correlation between the Fmax score. The conclusion was the same when we only considered long-range contacts predicted within the top L/5 scores (<xref ref-type="fig" rid="F4">Figure 4B</xref>); no correlation was observed.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Function prediction accuracy relative to structural features of target proteins. <bold>(A)</bold> and <bold>(B)</bold>, Fmax score of ContactPFP relative to the precision of contact prediction. Each point is corresponding to a protein which has an experimentally determined structure. There were 1,029 proteins of them. <bold>(A)</bold> Fmax score relative to the precision of all predicted contacts. Contacts are defined for residue pairs that have a C&#x3b2; distance within 12&#xa0;&#x212b; from each other. The average precision was 0.801. <bold>(B)</bold> Fmax score relative to the precision when we considered the top L/5 predicted long-range contacts, which were defined as contacts that are 24 residues or more apart on the sequence. L is the length of a protein. Contacts were defined as residue pairs that have their C&#x3b2; atoms placed within 8&#xa0;&#x212b;. The average precision L/5 long precision was 0.908. <bold>(C)</bold> Comparison of the performance between ContactPFP using predicted contacts and ContactPFP that uses accurate contacts taken from the experimentally determined structures. Contacts are defined for residue pairs that have C&#x3b2; atoms within 12&#xa0;&#x212b; from each other. Fmax scores of the 1,029 targets that have PDB structures were compared. <bold>(D)</bold> The effect of the fraction of disordered regions in proteins to the Fmax score. We used fldpnn to predict residues in disordered regions.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g004.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F4">Figure 4C</xref>, we further examined what would happen if we used completely accurate contacts for targets that were taken from experimentally determined structures. Interestingly, the Fmax score was higher when predicted contacts were used. The Fmax score of ContactPFP using predicted contacts and accurate contacts were 0.744 and 0.658, respectively. This is mostly because we use the reference database of proteins with predicted contacts. Similar proteins are likely to have similar predicted contact patterns, either accurate or inaccurate, and the similarity can be captured by contact graph comparison.</p>
</sec>
<sec id="s3-4">
<title>3.4 Effect of Disordered Regions</title>
<p>We were also curious how ContactPFP performs for proteins that have intrinsic disordered regions (IDRs) because an IDR does not usually form residue contacts. In <xref ref-type="fig" rid="F4">Figure 4D</xref>, we examined Fmax scores of target proteins relative to the fraction of IDRs in a protein. IDRs were predicted with fldpnn (<xref ref-type="bibr" rid="B23">Hu et al., 2021</xref>). To make disorder predictions more reliable, we used SPIDER3-single (<xref ref-type="bibr" rid="B22">Heffernan et al., 2018</xref>) to predict secondary structure of proteins and only considered residues which were also predicted as loops (class C) SPIDER3-single as the final disorder residues. We did not observe clear correlation between Fmax scores and the fraction of IDRs.</p>
</sec>
<sec id="s3-5">
<title>3.5 Case Studies</title>
<p>In this section, we discuss cases that illustrate ContactPFP&#x2019;s performance relative to sequence-based methods.</p>
<sec id="s3-5-1">
<title>3.5.1 Case 1: Outer Membrane Porin G</title>
<p>The first example shows a successful prediction by ContactPFP for outer membrane porin G (UniProt ID: P76045). This protein is present in the outer membrane of <italic>E. coli</italic>, for which GO terms such as &#x201c;cell outer membrane&#x201d; (GO: 0009279) are annotated in the CC category. This protein is transmembrane and transports sugars from outside to inside the cell. This corresponds to &#x201c;maltose transporting porin activity&#x201d; (GO: 0015481). For this protein, ContactPFP showed a high prediction accuracy, a Fmax score of 0.754, while it was 0.083, 0.140, and 0.085 for PFP, ESG, and Phylo-PFP, respectively. <xref ref-type="table" rid="T3">Table 3</xref> shows predicted correct and incorrect GO terms by the methods. ContactPFP was able to predict three out of four correct GO terms with the highest score of 1.0 and the remaining term (pore complex) with a score about half of the highest score (0.473). In contrast, Phylo-PFP, ESG, and PFP assigned low scores to the correct terms and instead selected wrong GO terms that do not exist in the UniProt entry with the highest (1.0) or with high scores over 0.9.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Predicted GO terms for outer membrane porin G (UniProt ID: P76045).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" colspan="2" align="left"/>
<th rowspan="2" align="center">Correct GO terms</th>
<th colspan="4" align="center">Confidence Score</th>
</tr>
<tr>
<th align="center">ContactPFP</th>
<th align="center">Phylo-PFP</th>
<th align="center">ESG</th>
<th align="center">PFP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">MF</td>
<td align="left">GO:0015481</td>
<td align="left">Maltose transporting porin activity</td>
<td align="center">-</td>
<td align="center">-</td>
<td align="center">-</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0015478</td>
<td align="left">oligosaccharide transporting porin activity</td>
<td align="center">-</td>
<td align="center">-</td>
<td align="char" char=".">0.001</td>
<td align="center">-</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0015288</td>
<td align="left">porin activity</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.283</td>
<td align="char" char=".">0.090</td>
<td align="char" char=".">0.250</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0034219</td>
<td align="left">carbohydrate transmembrane transport</td>
<td align="center">-</td>
<td align="char" char=".">0.432</td>
<td align="char" char=".">0.053</td>
<td align="char" char=".">0.490</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0006811</td>
<td align="left">ion transport</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.458</td>
<td align="char" char=".">0.087</td>
<td align="char" char=".">0.510</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0009279</td>
<td align="left">cell outer membrane</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.345</td>
<td align="char" char=".">0.092</td>
<td align="char" char=".">0.380</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0045203</td>
<td align="left">integral component of cell outer membrane</td>
<td align="center">-</td>
<td align="char" char=".">0.099</td>
<td align="char" char=".">0.062</td>
<td align="char" char=".">0.110</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0046930</td>
<td align="left">pore complex</td>
<td align="char" char=".">0.473</td>
<td align="char" char=".">0.099</td>
<td align="char" char=".">0.090</td>
<td align="char" char=".">0.110</td>
</tr>
<tr>
<td colspan="2" align="left"/>
<td align="left">
<bold>Incorrect GO terms</bold>
</td>
<td align="center">
<bold>ContactPFP</bold>
</td>
<td align="center">
<bold>Phylo-PFP</bold>
</td>
<td align="center">
<bold>ESG</bold>
</td>
<td align="center">
<bold>PFP</bold>
</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0046872</td>
<td align="left">metal ion binding</td>
<td align="center">-</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.272</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0007155</td>
<td align="left">cell adhesion</td>
<td align="center">-</td>
<td align="char" char=".">0.975</td>
<td align="char" char=".">0.161</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005737</td>
<td align="left">cytoplasm</td>
<td align="center">-</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.097</td>
<td align="char" char=".">0.920</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>All correct GO terms assigned in the UniProt entry are listed. For incorrect GO terms, GO terms that illustrate the difference between ContactPFP and the other methods are shown. Scores assigned to GO terms by the methods were normalized by the highest GO term score for this target protein. Thus, 1.0 means it is the top (i.e., most confident) prediction by the method for this protein.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>In <xref ref-type="fig" rid="F5">Figure 5</xref>, we analyzed how the correct GO prediction was possible by ContactPFP. The query protein has the &#x3b2; barrel fold, which was well predicted by trRosetta (<xref ref-type="fig" rid="F5">Figures 5A,D</xref>). With the accurate contact prediction of the query, ContactPFP was able to identify two other outer membrane proteins, YaiO (YAIO_ECOLI) and probable N-acetylneuraminic acid outer membrane channel protein NanC (NANC_ECOL6) from <italic>E. coli</italic> O6:H1, that also have a &#x3b2; barrel fold (<xref ref-type="fig" rid="F5">Figures 5B&#x2013;F</xref>). These two proteins have correct GO terms, GO:0009279 (cell outer membrane), GO:0015288 (porin activity), GO:0006811 (ion transport), GO:0046930 (pore complex), and GO:0008643 (carbohydrate transport), which is a parent, more-general term of a correct term, GO:0034219 (carbohydrate transmembrane transport). Since these top 2 most similar structures were used for GO term transfer, ContactPFP was able to make the correct GO predictions in <xref ref-type="table" rid="T3">Table 3</xref>. The query protein has been reported to have low sequence similarity with other porin proteins with similar functions (<xref ref-type="bibr" rid="B57">Subbarao and van den Berg, 2006</xref>). Indeed, the E-value of these two proteins to the query was over 125 and thus they were not able to be detected by PSI-BLAST. <xref ref-type="fig" rid="F5">Figure 5G</xref> shows the top 50 hits by PSI-BLAST. As shown, this query does not have similar sequences in Swiss-Prot. All the hits have almost 0 -log (E-value) scores. To conclude, in this example functionally related proteins were only retrieved by the similarity of structure but not by sequence.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>GO prediction by ContactPFP for outer membrane porin G (P76045). The first three panels <bold>(A&#x2013;C)</bold> and the subsequent panels in the second row <bold>(D&#x2013;F)</bold> are predicted residue contacts and resulting protein structure models. <bold>(A)</bold> The predicted contact maps of the query, OMPG_ECOLI (P76045). Residue pairs predicted to be in contact are shown in yellow. <bold>(B)</bold> The predicted contact maps of YAIO_ECOLI (Q47534), the most similar contact map with the GR-align score of 0.733 <bold>(C)</bold> The predicted contact maps of NANC_ECOL6 (P69856). The second closest contact map with the GR-align score of 0.658. GO terms of these two proteins were used for the prediction. <bold>(D)</bold> The predicted structure of OMPG_ECOLI (P76045) was generated by trRosetta (rainbow) superimposed with PDB structure 2X9K (gray). The root mean square deviation (RMSD) of the model to the native is 3.63&#xa0;&#xc5;. <bold>(E)</bold> The predicted structure of YAIO_ECOLI (Q47534) was generated by trRosetta (rainbow). For this protein, no experimental structure has been reported. <bold>(F)</bold> The predicted structure of NANC_ECOL6 (P69856) was generated by trRosetta (rainbow). No experimental structure was reported for this protein. <bold>(G)</bold> The top hits for OMPG_ECOLI by PSI-BLAST search against Swiss-Prot. The query itself is shown in the first position. Funsim functional similarity scores (<xref ref-type="bibr" rid="B51">Schlicker et al., 2006</xref>; <xref ref-type="bibr" rid="B19">Hawkins et al., 2009</xref>). The three categories of each protein compared with the query are shown in the top row in a color scale. The y-axis shows the sequence similarity in the form of -log<sub>10</sub> (E-value). The proteins that have incorrect GO terms listed in <xref ref-type="table" rid="T3">Table 3</xref> are marked with symbols: &#x2a;, &#x201c;metal ion binding&#x201d; (GO: 0046872); &#x23;, &#x201c;cell adhesion&#x201d; (GO: 0007155); and &#x2020;, &#x201c;cytoplasm&#x201d; (GO: 0005737).</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g005.tif"/>
</fig>
</sec>
<sec id="s3-5-2">
<title>3.5.2 Case 2: Leucine-Rich Repeat-Containing Protein 10</title>
<p>This is another successful example of ContactPFP, where it predicted more accurately than the sequence-based methods. The reason for this success was different from the first case. The query is a Leucine-rich repeat-containing protein 10 of mice (UniProt ID: Q8K3W2). The function of this protein includes &#x201c;actin binding&#x201d; (GO: 0003779), &#x201c;alpha-actinin binding&#x201d; (GO: 0051393), &#x201c;cardiac muscle cell development&#x201d; (GO: 0055013), and the protein is localized in the &#x201c;nucleus&#x201d; (GO: 0005634), &#x201c;cytoskeleton&#x201d; (GO: 0005856), &#x201c;mitochondrion&#x201d; (GO: 0005739), &#x201c;sarcomere&#x201d; (GO: 0030017), and &#x201c;myofibril&#x201d; (GO: 0030016). As shown in <xref ref-type="table" rid="T4">Table 4</xref>, ContactPFP predicted all GO term correctly with the highest confidence score, 1.0. For this protein, ContactPFP showed a high prediction accuracy, a Fmax score of 1.000, while it was 0.131, 0.115, and 0.160 by PFP, ESG, and Phylo-PFP, respectively.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Predicted GO terms for Leucine-rich repeat-containing protein 10 (UniProt ID: Q8K3W2).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" colspan="2" align="left"/>
<th rowspan="2" align="center">Correct GO terms</th>
<th colspan="4" align="center">Confidence Score</th>
</tr>
<tr>
<th align="center">ContactPFP</th>
<th align="center">Phylo-PFP</th>
<th align="center">ESG</th>
<th align="center">PFP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">MF</td>
<td align="left">GO:0003779</td>
<td align="left">actin binding</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.582</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.070</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0051393</td>
<td align="left">alpha-actinin binding</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.521</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.010</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0055013</td>
<td align="left">cardiac muscle cell development</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.521</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.020</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005634</td>
<td align="left">nucleus</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.315</td>
<td align="char" char=".">0.257</td>
<td align="char" char=".">0.180</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005856</td>
<td align="left">cytoskeleton</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.201</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.050</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005739</td>
<td align="left">mitochondrion</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.210</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.040</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0030017</td>
<td align="left">sarcomere</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.178</td>
<td align="char" char=".">0.129</td>
<td align="char" char=".">0.010</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0030016</td>
<td align="left">myofibril</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.155</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.010</td>
</tr>
<tr>
<td colspan="2" align="left"/>
<td align="left">
<bold>Incorrect GO terms</bold>
</td>
<td align="center">
<bold>ContactPFP</bold>
</td>
<td align="center">
<bold>Phylo-PFP</bold>
</td>
<td align="center">
<bold>ESG</bold>
</td>
<td align="center">
<bold>PFP</bold>
</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0005524</td>
<td align="left">ATP binding</td>
<td align="char" char=".">-</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">-</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0006952</td>
<td align="left">defense response</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.478</td>
<td align="char" char=".">0.127</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0030054</td>
<td align="left">cell junction</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.360</td>
<td align="char" char=".">0.861</td>
<td align="char" char=".">0.130</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>See the caption in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The correct GO term predictions by ContactPFP were transferred from the two most structurally similar proteins shown in <xref ref-type="fig" rid="F6">Figure 6B,C,E,F</xref>. The query and these two proteins have a horse-shoe fold, a typical fold for proteins with Leucine-rich repeats. These two structures have high graph similarity scores of 0.917 (LRC10_BOVIN) and 0.914 (LRC10_HUMAN), respectively. These two proteins also have significant sequence similarities with E-value of 8e-59 and 1e-59, with the third and the second hits as shown in <xref ref-type="fig" rid="F6">Figure 6G</xref>. However, the poor prediction accuracy by the sequence-based methods occurred because there are many other proteins with significant sequence similarity, which do not have common GO terms with the query. As shown in <xref ref-type="fig" rid="F6">Figure 6G</xref>, all top 50 hits have an E-value of 10<sup>&#x2013;40</sup> or smaller, but only a few of them have correct GO terms. As a result, incorrect GO terms (shown in symbols) that frequently appear in the top 50 hits accumulated higher scores. Thus, in this case, the structure information was able to select the most functionally relevant proteins among proteins that are similar in sequence but not in function.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Illustration of GO term predictions by ContactPFP for Leucine-rich repeat-containing protein 10 from mouse (LRC10_MOUSE, Q8K3W2). Residue pair contact prediction of <bold>(A)</bold> The query, LRC10_MOUSE; <bold>(B)</bold> Leucine-rich repeat-containing protein from bovine, LRC10_BOVIN (Q24K06), and <bold>(C)</bold> Leucine-rich repeat-containing protein from human LRC10_HUMAN (Q5BKY1). The following three panels, <bold>(D&#x2013;F)</bold>, are the corresponding predicted structures of these three proteins, respectively. The color shows the orientation of the proteins from the N-terminus to the C-terminus from blue to red. There are no experimentally determined structures for these proteins. <bold>(G)</bold> The top 50 hits for the query by PSI-BLAST against Swiss-Prot. Funsim scores compared with the query protein are shown in the top row in a color scale. The leftmost column is the query itself. The protein names associated with the &#x201c;incorrect GO terms&#x201d; listed in <xref ref-type="table" rid="T4">Table 4</xref> are marked with the corresponding symbols, &#x2a;, &#x201c;ATP binding&#x201d; (GO:0005524); &#x23;, &#x201c;defense response&#x201d; (GO: 0006952); and &#x2020; &#x201c;cell junction&#x201d; (GO:0030054).</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g006.tif"/>
</fig>
</sec>
<sec id="s3-5-3">
<title>3.5.3 Case 3: Cyclin-dependent Kinase Inhibitor 4</title>
<p>The last one is the opposite case where ContactPFP did not perform as well as the other sequence-based methods. The query is Cyclin-dependent kinase inhibitor 4 in Arabidopsis (KRP4_ARATH, Q8GYJ3). This protein has GO annotations of &#x201c;cyclin-dependent protein serine/threonine kinase inhibitor activity&#x201d; (GO: 0004861), &#x201c;negative regulation of cell cycle&#x201d; (GO: 0045786), &#x201c;negative regulation of cyclin-dependent protein serine/threonine kinase activity&#x201d; (GO: 0045736), &#x201c;nucleoplasm&#x201d; (GO: 0005654), &#x201c;nucleus&#x201d; (GO: 0005634), and &#x201c;cytoplasm&#x201d; (GO: 0005737) (<xref ref-type="table" rid="T5">Table 5</xref>). As shown in <xref ref-type="table" rid="T5">Table 5</xref>, ContactPFP predicted only one term among the correct terms and instead predicted wrong terms, including actin filament binding and organization (GO:0051015, GO:0007015), and microtubule (GO:0005874), which was worse than the other three sequence-based methods. The Fmax score of ContactPFP was 0.127, while Phylo-PFP, ESG, and PFP had a high Fmax score of 0.995.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>The detail of predicted GO terms for Cyclin-dependent kinase inhibitor 4 (UniProt ID: Q8GYJ3).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" colspan="2" align="left"/>
<th rowspan="2" align="center">Correct GO terms</th>
<th colspan="4" align="center">Confidence Score</th>
</tr>
<tr>
<th align="center">ContactPFP</th>
<th align="center">Phylo-PFP</th>
<th align="center">ESG</th>
<th align="center">PFP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">MF</td>
<td align="left">GO:0004861</td>
<td align="left">cyclin-dependent protein serine/threonine kinase inhibitor activity</td>
<td align="char" char=".">-</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.923</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0007049</td>
<td align="left">cell cycle</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.141</td>
<td align="char" char=".">0.923</td>
<td align="char" char=".">0.100</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0045786</td>
<td align="left">negative regulation of cell cycle</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.274</td>
<td align="char" char=".">0.923</td>
<td align="char" char=".">0.300</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0045736</td>
<td align="left">negative regulation of cyclin-dependent protein serine/threonine kinase activity</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.292</td>
<td align="char" char=".">0.346</td>
<td align="char" char=".">0.410</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005654</td>
<td align="left">nucleoplasm</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.486</td>
<td align="char" char=".">0.553</td>
<td align="char" char=".">0.640</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005634</td>
<td align="left">nucleus</td>
<td align="char" char=".">-</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.617</td>
<td align="char" char=".">1.000</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005737</td>
<td align="left">cytoplasm</td>
<td align="char" char=".">0.988</td>
<td align="char" char=".">0.053</td>
<td align="char" char=".">0.005</td>
<td align="char" char=".">0.180</td>
</tr>
<tr>
<td colspan="2" align="left"/>
<td align="left">
<bold>Incorrect GO terms</bold>
</td>
<td align="center">
<bold>ContactPFP</bold>
</td>
<td align="center">
<bold>Phylo-PFP</bold>
</td>
<td align="center">
<bold>ESG</bold>
</td>
<td align="center">
<bold>PFP</bold>
</td>
</tr>
<tr>
<td align="left">MF</td>
<td align="left">GO:0051015</td>
<td align="left">actin filament binding</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.004</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.010</td>
</tr>
<tr>
<td align="left">BP</td>
<td align="left">GO:0007015</td>
<td align="left">actin filament organization</td>
<td align="char" char=".">1.000</td>
<td align="char" char=".">0.001</td>
<td align="char" char=".">-</td>
<td align="char" char=".">-</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">GO:0005874</td>
<td align="left">microtubule</td>
<td align="char" char=".">0.988</td>
<td align="char" char=".">0.002</td>
<td align="char" char=".">-</td>
<td align="char" char=".">0.010</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>See the caption in <xref ref-type="table" rid="T3">Table 3</xref>.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>According to UniProt, more than half of residues are annotated as disordered. Therefore, it is highly likely that predicted contacts (<xref ref-type="fig" rid="F7">Figure 7A</xref>) and structures (<xref ref-type="fig" rid="F7">Figure 7D</xref>) are incorrect. Furthermore, the top two structures selected by ContactPFP have a long, straight helical structure, which is not similar overall to the predicted structure of the query. Indeed, these two retrieved proteins, TPM3_HUMAN (<xref ref-type="fig" rid="F7">Figures 7B,E</xref>) and TPM_CHAFE (<xref ref-type="fig" rid="F7">Figures 7C,F</xref>) do not have any common GO terms with the query protein. In contrast, about a dozen top hits by sequence similarity search are functionally highly similar to the query, which is reflected in the high Fmax scores of Phylo-PFP, ESG, and PFP (<xref ref-type="fig" rid="F7">Figure 7G</xref>). Thus, to conclude, this is an example where incorrect protein structure prediction led to the failure of ContactPFP&#x2019;s function prediction.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Illustration of GO term predictions by ContactPFP for Cyclin-dependent kinase inhibitor 4 (KRP4_ARATH, Q8GYJ3). The first three panels are predicted contacts for the query, KRP4_ARATH <bold>(A)</bold>, and the two most similar proteins in terms of the contact pattern, <bold>(B)</bold> TPM3_HUMAN and <bold>(C)</bold> TPM_CHAFE. The graph similarity scores by GR-align were 0.677 and 0.673, respectively. Panel D, E, F are predicted structures of these three proteins by trRosetta in the same order as the first row. <bold>(G)</bold>, The top 50 hits for the query by PSI-BLAST against Swiss-Prot. Funsim scores compared with the query protein are shown in the top row in a color scale. The most left column is the query itself. The proteins that have incorrect GO terms listed in <xref ref-type="table" rid="T5">Table 5</xref> are marked with symbols, &#x2a;, &#x201c;actin filament binding&#x201d; (GO: 0051015); &#x23;, &#x201c;actin filament organization&#x201d; (GO: 0007015); and &#x2020;, &#x201c;microtubule&#x201d; (GO: 0005874).</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g007.tif"/>
</fig>
</sec>
</sec>
<sec id="s3-6">
<title>3.6 Ensemble Methods</title>
<p>ContactPFP&#x2019;s characteristic prediction performance discussed in <xref ref-type="fig" rid="F3">Figure 3</xref> motivated us to develop ensemble methods. Particularly, the primary focus is to improve predictions for proteins where ContactPFP did not perform well but other conventional methods achieved higher Fmax scores. Combining ContactPFP with other methods also makes sense from a biological point of view because the former uses protein structure information, which is complementary with the latter that uses sequence information.</p>
<p>We constructed ensemble methods of all possible combinations of the methods starting from single methods to a combination of all five methods (<xref ref-type="fig" rid="F8">Figure 8A</xref>; <xref ref-type="table" rid="T6">Table 6</xref>). When multiple methods are combined, scores of GO terms from the combined methods were simply averaged. The highest average Fmax score, 0.699, was achieved by a combination of three methods, ContactPFP, Phylo-PFP, and PSI-BLAST. Compared with the Fmax of the lone ContactPFP, 0.638, it is a 9.6% improvement. From <xref ref-type="fig" rid="F4">Figure 4A</xref>, we can see that the top methods all include ContactPFP as its ensemble component, which implies it is complementary to the other methods.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The prediction performance of ensemble methods with ContactPFP. <bold>(A)</bold> The average Fmax score of ensemble methods with ContactPFP. Results of all the combinations of 1&#x2013;5 methods are shown. Patterns in the bar graphs show the number of methods combined. The bars are sorted by their Fmax scores. CPFP, ContactPFP; PHYLO, Phylo-PFP; BLAST, PSI-BLAST. <bold>(B)</bold> Fmax score distribution of the ensemble method with ContactPFP, Phylo-PFP, and PSI-BLAST, the combination with the highest Fmax score, and distributions of individual methods shown in violin plots. The three horizontal bars in a plot indicate the maximum, median, and minimum values. <bold>(C)</bold> Comparison of Fmax scores of individual target proteins by the best ensemble method and Phylo-PFP. Each point represents a target protein in the benchmark dataset. <bold>(D)</bold> Comparison of Fmax scores of individual target proteins by the ContactPFP &#x2b; PhyloPFP and ContactPFP &#x2b; PSI-BLAST.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g008.tif"/>
</fig>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>The function prediction performance of ensemble methods.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">&#x23; of methods</th>
<th rowspan="2" align="center">Ensemble name</th>
<th colspan="4" align="center">Fmax</th>
</tr>
<tr>
<th align="center">ALL</th>
<th align="center">CC</th>
<th align="center">MF</th>
<th align="center">BP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">3</td>
<td align="left">CPFP-PHYLO-BLAST</td>
<td align="char" char=".">0.699</td>
<td align="char" char=".">0.778</td>
<td align="char" char=".">0.799</td>
<td align="char" char=".">0.679</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CPFP-PHYLO-ESG-BLAST</td>
<td align="char" char=".">0.694</td>
<td align="char" char=".">0.774</td>
<td align="char" char=".">0.795</td>
<td align="char" char=".">0.672</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CPFP-PHYLO-PFP-BLAST</td>
<td align="char" char=".">0.693</td>
<td align="char" char=".">0.774</td>
<td align="char" char=".">0.786</td>
<td align="char" char=".">0.673</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">CPFP-BLAST</td>
<td align="char" char=".">0.688</td>
<td align="char" char=".">0.758</td>
<td align="char" char=".">0.789</td>
<td align="char" char=".">0.671</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CPFP-PFP-BLAST</td>
<td align="char" char=".">0.688</td>
<td align="char" char=".">0.77</td>
<td align="char" char=".">0.788</td>
<td align="char" char=".">0.67</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">CPFP-PHYLO-ESG-PFP-BLAST</td>
<td align="char" char=".">0.687</td>
<td align="char" char=".">0.771</td>
<td align="char" char=".">0.784</td>
<td align="char" char=".">0.666</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CPFP-ESG-BLAST</td>
<td align="char" char=".">0.684</td>
<td align="char" char=".">0.765</td>
<td align="char" char=".">0.795</td>
<td align="char" char=".">0.665</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CPFP-PHYLO-ESG</td>
<td align="char" char=".">0.683</td>
<td align="char" char=".">0.758</td>
<td align="char" char=".">0.782</td>
<td align="char" char=".">0.661</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CPFP-ESG-PFP-BLAST</td>
<td align="char" char=".">0.682</td>
<td align="char" char=".">0.767</td>
<td align="char" char=".">0.784</td>
<td align="char" char=".">0.662</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">CPFP-PHYLO</td>
<td align="char" char=".">0.68</td>
<td align="char" char=".">0.749</td>
<td align="char" char=".">0.775</td>
<td align="char" char=".">0.657</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">CPFP-PHYLO-ESG-PFP</td>
<td align="char" char=".">0.675</td>
<td align="char" char=".">0.762</td>
<td align="char" char=".">0.774</td>
<td align="char" char=".">0.652</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">CPFP-ESG</td>
<td align="char" char=".">0.674</td>
<td align="char" char=".">0.742</td>
<td align="char" char=".">0.779</td>
<td align="char" char=".">0.647</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">PHYLO-BLAST</td>
<td align="char" char=".">0.673</td>
<td align="char" char=".">0.758</td>
<td align="char" char=".">0.779</td>
<td align="char" char=".">0.651</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">PHYLO-ESG-BLAST</td>
<td align="char" char=".">0.669</td>
<td align="char" char=".">0.753</td>
<td align="char" char=".">0.777</td>
<td align="char" char=".">0.646</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CPFP-PHYLO-PFP</td>
<td align="char" char=".">0.668</td>
<td align="char" char=".">0.751</td>
<td align="char" char=".">0.762</td>
<td align="char" char=".">0.646</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">PHYLO-ESG</td>
<td align="char" char=".">0.668</td>
<td align="char" char=".">0.753</td>
<td align="char" char=".">0.768</td>
<td align="char" char=".">0.644</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">CPFP-ESG-PFP</td>
<td align="char" char=".">0.666</td>
<td align="char" char=".">0.748</td>
<td align="char" char=".">0.764</td>
<td align="char" char=".">0.644</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">PHYLO</td>
<td align="char" char=".">0.662</td>
<td align="char" char=".">0.75</td>
<td align="char" char=".">0.759</td>
<td align="char" char=".">0.641</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">PHYLO-PFP-BLAST</td>
<td align="char" char=".">0.662</td>
<td align="char" char=".">0.749</td>
<td align="char" char=".">0.762</td>
<td align="char" char=".">0.641</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">PHYLO-ESG-PFP-BLAST</td>
<td align="char" char=".">0.66</td>
<td align="char" char=".">0.75</td>
<td align="char" char=".">0.765</td>
<td align="char" char=".">0.638</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">CPFP-PFP</td>
<td align="char" char=".">0.659</td>
<td align="char" char=".">0.739</td>
<td align="char" char=".">0.752</td>
<td align="char" char=".">0.637</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">PHYLO-ESG-PFP</td>
<td align="char" char=".">0.651</td>
<td align="char" char=".">0.745</td>
<td align="char" char=".">0.747</td>
<td align="char" char=".">0.626</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">ESG-BLAST</td>
<td align="char" char=".">0.64</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.76</td>
<td align="char" char=".">0.612</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">ESG-PFP-BLAST</td>
<td align="char" char=".">0.639</td>
<td align="char" char=".">0.732</td>
<td align="char" char=".">0.753</td>
<td align="char" char=".">0.617</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">CPFP</td>
<td align="char" char=".">0.638</td>
<td align="char" char=".">0.718</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.606</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">PHYLO-PFP</td>
<td align="char" char=".">0.637</td>
<td align="char" char=".">0.736</td>
<td align="char" char=".">0.732</td>
<td align="char" char=".">0.614</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">ESG</td>
<td align="char" char=".">0.634</td>
<td align="char" char=".">0.714</td>
<td align="char" char=".">0.746</td>
<td align="char" char=".">0.598</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">PFP-ESG</td>
<td align="char" char=".">0.623</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.597</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">PFP-BLAST</td>
<td align="char" char=".">0.623</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.597</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">PFP</td>
<td align="char" char=".">0.586</td>
<td align="char" char=".">0.698</td>
<td align="char" char=".">0.689</td>
<td align="char" char=".">0.562</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">BLAST</td>
<td align="char" char=".">0.574</td>
<td align="char" char=".">0.655</td>
<td align="char" char=".">0.678</td>
<td align="char" char=".">0.544</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>All possible combinations are listed. In the column of ensemble name, CPFP, PHYLO, and BLAST correspond to ContactPFP, Phylo-PFP, and PSI-BLAST, respectively. CC, MF, BP corresponds to Cellular component, Molecular Function, and Biological Process, respectively.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>
<xref ref-type="fig" rid="F8">Figure 8B</xref> shows the Fmax score distribution of the best ensemble method and the five individual methods. ContactPFP&#x2019;s distribution has a characteristic peak at a low Fmax score of around 0.1. This peak disappeared when combined with other methods, which is the main reason why the performance improved by the ensemble method. In the score comparison of individual proteins (<xref ref-type="fig" rid="F8">Figure 8C</xref>), we can also see that many proteins with a low score around 0.1 by ContactPFP improved by the ensemble method.</p>
<p>One thing which drew our attention is that the top combination included PSI-BLAST, which performed the worst among the five non-ensembled methods. To examine why adding PSI-BLAST improved the performance, we compared the ensemble of ContactPFP with Phylo-PFP with the ensemble of ContactPFP with PSI-BLAST (<xref ref-type="fig" rid="F8">Figure 8D</xref>). We compared PSI-BLAST with Phylo-PFP because the latter was the best single method among the methods we compared (<xref ref-type="fig" rid="F3">Figure 3A</xref>). From this plot, we selected two target proteins where ContactPFP &#x2b; PSI-BLAST performed significantly better than ContactPFP &#x2b; Phylo-PFP. APOE_HUMAN, Apolipoprotein E of Human, is one such example. The Fmax score of ContactPFP &#x2b; Phylo-PFP was 0.234, while that of ContactPFP &#x2b; PSI-BLAST was 0.801. This protein has 164 GO annotations. By PSI-BLAST, most of the GO terms were found at least once in the top 10 sequences thus all the GO terms had a score of 0.931 (because this is how GO terms are scored in PSI-BLAST). These GO terms were found also in the sequences retrieved by Phylo-PFP and ContactPFP. However, the GO terms were found infrequently in the retrieved sequences and thus received a low score. The same mechanism was observed for DNJA3_HUMAN, DnaJ homolog subfamily A member 3, which had an Fmax score of 0.213 by ContactPFP &#x2b; Phylo-PFP and 0.988 by ContactPFP &#x2b; PSI-BLAST.</p>
</sec>
<sec id="s3-7">
<title>3.7 Computational Time</title>
<p>In <xref ref-type="fig" rid="F9">Figure 9</xref> we show the computation time of ContactPFP. In our computational environment, the entire ContactPFP pipeline took approximately 15&#xa0;min for a 500 residue-long protein. <xref ref-type="fig" rid="F8">Figure 8</xref> also shows the breakdown of the time needed for five steps in ContactPFP. The computational time for running trRosetta and reference database search by GR-Align grows as the protein length increases. Particularly, the computational time for using trRosetta grows sharply, and it exceeds the time for the reference database search when the query protein is longer than 500 residues.</p>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>The cumulative computational time of ContactPFP. The time is decomposed into five steps: The database search with HHblits, the distance map prediction by trRosetta, converting a predicted distance map to a contact graph, contact graph comparison against the reference database by GR-Align, constructing GO term list from a hit list, are reported. The times are reported in the wall-clock time (seconds). All computations were performed on CPU, 2 AMD EPYC 7252 cores (16 cores in total) with 128GB RAM. The following 13 proteins were used, which have a length between 100 and 800 amino acids. The length of each protein is shown in the parenthesis. P0CM71 (98), P9WF14 (150), P69162 (200), B1W5S5 (250), A1YG61 (300), Q6Q972 (350), Q550G0 (400), C5A1K9 (450), Q00456 (500), A1DHW5 (550), A5DX93 (600), Q96QV1 (700), and Q54WZ0 (800). These proteins were chosen because they hit the same number of sequences, 500 (&#xb1; 10) sequences, by HHblits.</p>
</caption>
<graphic xlink:href="fbinf-02-896295-g009.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>ContactPFP developed in this work identifies proteins with similar contact maps and transfers their functions to the query. Despite the knowledge that protein structure and function are closely related, protein structure information has not been effectively used for automatic protein function prediction mainly because of the low coverage of experimentally determined structure information for proteins. It is now possible to employ structure prediction methods to cover structure information of remaining proteins that have no experimentally determined structures. ContactPFP showed a slightly lower average Fmax score than one of the best sequence-based methods, Phylo-PFP, but had more wins over Phylo-PFP when predictions for individual proteins were counted. Thus, overall, we could say ContactPFP performed on par with Phylo-PFP. Combining ContactPFP in ensemble methods with sequence-based methods successfully achieved higher accuracy than the individual methods. In the current work we used simple averaging to ensemble scores from different methods. It would be worthwhile to explore other approaches beyond averaging, such as learning-to-rank, or even other signals that might indicate that a given query protein will benefit more from sequence similarity instead of structural similarity.</p>
<p>Since ContactPFP is based on database search, it can predict any GO terms, including very rare ones, as long as proteins found by a search have such GO annotations. This is very important for practical use of a function prediction method and is different from recent machine learning-based methods (<xref ref-type="bibr" rid="B32">Kulmanov and Hoehndorf, 2020</xref>; <xref ref-type="bibr" rid="B62">Wan and Jones, 2020</xref>; <xref ref-type="bibr" rid="B65">You et al., 2021</xref>), which need training on a dataset of proteins with a limited set of abundant GO terms.</p>
<p>Besides proving practical usefulness, we have a couple of important findings. Through comparison with sequence-based methods, we observed the strengths and the weakness of ContactPFP, which would apply to any function prediction methods that use predicted protein structures. The notable strength is that, as illustrated in the first case study, ContactPFP can often identify distantly related proteins by considering structural similarity, which leads to more accurate function prediction than sequence-based methods. ContactPFP was also able to select the most relevant proteins among proteins that are similar in the sequence (the second case study). On the other hand, weaknesses originate from the accuracy and current challenges of protein structure prediction. Structure-based retrieval does not work well when predicted structures (predicted contact maps) are not accurate. Also, handling intrinsically disordered proteins is a challenge because all disordered proteins look alike, and it is hard to distinguish functionally similar ones by their structures. An interesting finding is that the structure-based approach showed complementary strengths from sequence-based methods (<xref ref-type="fig" rid="F3">Figure 3</xref>), and thus it is effective to construct an ensemble approach with other methods.</p>
<p>We compared simplified protein structure representations, residue contact maps as graphs, instead of directly using the three-dimensional structures. The comparison of contact maps made it possible to scan the reference database within a realistic amount of time, although it still took about 20&#xa0;min for prediction on one query protein. A further speed up will be possible by using a different, efficient structure representation, such as the 3D Zernike descriptor (<xref ref-type="bibr" rid="B61">Venkatraman et al., 2009</xref>; <xref ref-type="bibr" rid="B31">Kihara et al., 2011</xref>) as it was successfully applied for real-time protein structure database search (<xref ref-type="bibr" rid="B48">Sael et al., 2008</xref>; <xref ref-type="bibr" rid="B33">La et al., 2009</xref>; <xref ref-type="bibr" rid="B12">Esquivel-Rodr&#xed;guez et al., 2015</xref>; <xref ref-type="bibr" rid="B18">Han et al., 2017</xref>; <xref ref-type="bibr" rid="B2">Aderinwale et al., 2022</xref>).</p>
<p>The development of various bioinformatics tools using predicted protein structures will progress further in the future as a more recent method, Alphafold2 (<xref ref-type="bibr" rid="B28">Jumper et al., 2021</xref>) made significant improvements in the modeling accuracy. Function prediction (<xref ref-type="bibr" rid="B46">Sael et al., 2012</xref>) from predicted structures will be one such major application (<xref ref-type="bibr" rid="B15">Gligorijevi&#x107; et al., 2021</xref>). Here, we showed an approach using global protein structure comparison, but with structures, we can also identify local functional sites of proteins (<xref ref-type="bibr" rid="B9">Chikhi et al., 2010</xref>; <xref ref-type="bibr" rid="B70">Zhu et al., 2015</xref>; <xref ref-type="bibr" rid="B54">Sit et al., 2019</xref>) and predict binding ligands (<xref ref-type="bibr" rid="B52">Shin et al., 2016</xref>; <xref ref-type="bibr" rid="B69">Zhu et al., 2016</xref>).</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.6525075">https://doi.org/10.5281/zenodo.6525075</ext-link>, <ext-link ext-link-type="uri" xlink:href="https://github.com/kiharalab/contactpfp">https://github.com/kiharalab/contactpfp</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>DK conceived the study. YK developed the ContactPFP pipeline and conducted the experiments. SF generated the benchmark dataset. SF and AJ ran three existing function prediction methods. YK analyzed the data and wrote the initial draft of the manuscript. DK critically edited it.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>DK acknowledges support from the National Institutes of Health (R01GM133840, R01GM123055, and 3R01GM133840-02S1), the National Science Foundation (DBI2003635, DBI2146026, CMMI1825941, and MCB1925643). YK was supported in part by the Top Global University Project from the Ministry of Education, Culture, Sports, Science and Technology of Japan (MEXT).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>Computations were mainly performed on the NIG supercomputer at ROIS National Institute of Genetics, Japan.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abriata</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Tam&#xf2;</surname>
<given-names>G. E.</given-names>
</name>
<name>
<surname>Dal Peraro</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>A Further Leap of Improvement in Tertiary Structure Prediction in CASP13 Prompts New Routes for Future Assessments</article-title>. <source>Proteins</source> <volume>87</volume>, <fpage>1100</fpage>&#x2013;<lpage>1112</lpage>. <pub-id pub-id-type="doi">10.1002/prot.25787</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aderinwale</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Bharadwaj</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Christoffer</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Terashi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Jahandideh</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Real-Time Structure Search and Structure Classification for AlphaFold Protein Models</article-title>. <source>Commun. Biol.</source> <volume>5</volume> (<issue>1</issue>), <fpage>316</fpage>. <pub-id pub-id-type="doi">10.1038/s42003-022-03261-8</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Gish</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Myers</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Basic Local Alignment Search Tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume>, <fpage>403</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Madden</surname>
<given-names>T. L.</given-names>
</name>
<name>
<surname>Sch&#xe4;ffer</surname>
<given-names>A. A.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>1997</year>). <article-title>Gapped BLAST and PSI-BLAST: A New Generation of Protein Database Search Programs</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume>, <fpage>3389</fpage>&#x2013;<lpage>3402</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.17.3389</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Attwood</surname>
<given-names>T. K.</given-names>
</name>
<name>
<surname>Coletta</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Muirhead</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Pavlopoulou</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Philippou</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Popov</surname>
<given-names>I.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>The PRINTS Database: A Fine-Grained Protein Sequence Annotation and Analysis Resource&#x2014;Its Status in 2012</article-title>. <source>Database</source> <volume>2012</volume>, <fpage>bas019</fpage>. <pub-id pub-id-type="doi">10.1093/database/bas019</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bairoch</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bucher</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>PROSITE: Recent Developments</article-title>. <source>Nucleic Acids Res.</source> <volume>22</volume>, <fpage>3583</fpage>&#x2013;<lpage>3589</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Boutet</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Lieberherr</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tognolli</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Schneider</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bansal</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bridge</surname>
<given-names>A. J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). &#x201c;<article-title>Uniprotkb/swiss-prot, the Manually Annotated Section of the Uniprot Knowledgebase: How to Use the Entry View</article-title>,&#x201d; in <source>Methods in Molecular Biology</source>. Editor <person-group person-group-type="editor">
<name>
<surname>Edwards</surname>
<given-names>D.</given-names>
</name>
</person-group> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>23</fpage>&#x2013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-3167-5_2</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bru</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Courcelle</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Carr&#xe8;re</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Beausse</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Dalmar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kahn</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>The ProDom Database of Protein Domain Families: More Emphasis on 3D</article-title>. <source>Nucleic Acids Res.</source> <volume>33</volume>, <fpage>D212</fpage>&#x2013;<lpage>D215</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki034</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chikhi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Real-Time Ligand Binding Pocket Database Search Using Local Surface Descriptors</article-title>. <source>Proteins</source> <volume>78</volume>, <fpage>2007</fpage>&#x2013;<lpage>2028</lpage>. <pub-id pub-id-type="doi">10.1002/PROT.22715</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chitale</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hawkins</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>ESG: Extended Similarity Group Method for Automated Protein Function Prediction</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>1739</fpage>&#x2013;<lpage>1745</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp309</pub-id> </citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chothia</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lesk</surname>
<given-names>A. M.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>The Relation Between the Divergence of Sequence and Structure in Proteins</article-title>. <source>EMBO J.</source> <volume>5</volume> (<issue>4</issue>), <fpage>823</fpage>&#x2013;<lpage>826</lpage>. <pub-id pub-id-type="doi">10.1002/j.1460-2075.1986.tb04288.x</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Das</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Scholes</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Sen</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Orengo</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>CATH Functional Families Predict Functional Sites in Proteins</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>1099</fpage>&#x2013;<lpage>1106</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa937</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Esquivel-Rodr&#xed;guez</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Guang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Christoffer</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Navigating 3D Electron Microscopy Maps with EM-SURFER</article-title>. <source>BMC Bioinforma.</source> <volume>16</volume>, <fpage>181</fpage>. <pub-id pub-id-type="doi">10.1186/S12859-015-0580-6</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Finn</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Attwood</surname>
<given-names>T. K.</given-names>
</name>
<name>
<surname>Babbitt</surname>
<given-names>P. C.</given-names>
</name>
<name>
<surname>Bateman</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bork</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Bridge</surname>
<given-names>A. J.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>InterPro in 2017-Beyond Protein Family and Domain Annotations</article-title>. <source>Nucleic Acids Res.</source> <volume>45</volume>, <fpage>D190</fpage>&#x2013;<lpage>D199</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1107</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Finn</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Coggill</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Eberhardt</surname>
<given-names>R. Y.</given-names>
</name>
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Mistry</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Mitchell</surname>
<given-names>A. L.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>The Pfam Protein Families Database: Towards a More Sustainable Future</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>D279</fpage>&#x2013;<lpage>D285</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1344</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gligorijevi&#x107;</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Renfrew</surname>
<given-names>P. D.</given-names>
</name>
<name>
<surname>Kosciolek</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Leman</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Berenberg</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Vatanen</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Structure-Based Protein Function Prediction Using Graph Convolutional Networks</article-title>. <source>Nat. Commun.</source> <volume>12</volume> (<issue>1</issue>), <fpage>3168</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-23303-9</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Greener</surname>
<given-names>J. G.</given-names>
</name>
<name>
<surname>Kandathil</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>D. T.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Deep Learning Extends De Novo Protein Modelling Coverage of Genomes Using Iteratively Predicted Structural Constraints</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>3977</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-11994-0</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Haft</surname>
<given-names>D. H.</given-names>
</name>
<name>
<surname>Selengut</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Richter</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Harkins</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Basu</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Beck</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>TIGRFAMs and Genome Properties in 2013</article-title>. <source>Nucleic Acids Res.</source> <volume>41</volume>, <fpage>D387</fpage>&#x2013;<lpage>D395</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks1234</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Han</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Protein 3D Structure and Electron Microscopy Map Retrieval Using 3D-SURFER2.0 and EM-SURFER</article-title>. <source>Curr. Protoc. Bioinforma.</source> <volume>60</volume>, <fpage>3.14.1</fpage>&#x2013;<lpage>3.14.15</lpage>. <pub-id pub-id-type="doi">10.1002/CPBI.37</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hawkins</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Chitale</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Luban</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>PFP: Automated Prediction of Gene Ontology Functional Annotations with Confidence Scores Using Protein Sequence Data</article-title>. <source>Proteins</source> <volume>74</volume>, <fpage>566</fpage>&#x2013;<lpage>582</lpage>. <pub-id pub-id-type="doi">10.1002/prot.22172</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hawkins</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Function Prediction of Uncharacterized Proteins</article-title>. <source>J. Bioinform Comput. Biol.</source> <volume>5</volume>, <fpage>1</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1142/S0219720007002503</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hawkins</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Luban</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Enhanced Automated Function Prediction Using Distantly Related Sequences and Contextual Association by PFP</article-title>. <source>Protein Sci.</source> <volume>15</volume>, <fpage>1550</fpage>&#x2013;<lpage>1556</lpage>. <pub-id pub-id-type="doi">10.1110/ps.062153506</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heffernan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Paliwal</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lyons</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Singh</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Single-Sequence-Based Prediction of Protein Secondary Structures and Solvent Accessibility by Deep Whole-Sequence Learning</article-title>. <source>J. Comput. Chem.</source> <volume>39</volume>, <fpage>2210</fpage>&#x2013;<lpage>2216</lpage>. <pub-id pub-id-type="doi">10.1002/JCC.25534</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Katuwawala</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ghadermarzi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>flDPnn: Accurate Intrinsic Disorder Prediction with Putative Propensities of Disorder Functions</article-title>. <source>Nat. Commun.</source> <volume>12</volume> (<issue>1</issue>), <fpage>4438</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-24773-7</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jain</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Phylo-PFP: Improved Automated Protein Function Prediction Using Phylogenetic Distance of Distantly Related Sequences</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>753</fpage>&#x2013;<lpage>759</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty704</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jain</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Terashi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Kagaya</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Subramaniya</surname>
<given-names>S. R. M. V.</given-names>
</name>
<name>
<surname>Christoffer</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Analyzing Effect of Quadruple Multiple Sequence Alignments on Deep Learning Based Protein Inter-Residue Distance Prediction</article-title>. <source>Sci. Rep.</source> <volume>11</volume>, <fpage>7574</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-87204-z</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jassal</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Matthews</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Viteri</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gong</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Lorente</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fabregat</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>The Reactome Pathway Knowledgebase</article-title>. <source>Nucleic Acids Res.</source> <volume>48</volume>, <fpage>D498</fpage>&#x2013;<lpage>D503</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz1031</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Oron</surname>
<given-names>T. R.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>W. T.</given-names>
</name>
<name>
<surname>Bankapur</surname>
<given-names>A. R.</given-names>
</name>
<name>
<surname>D&#x27;Andrea</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Lepore</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>An Expanded Evaluation of Protein Function Prediction Methods Shows an Improvement in Accuracy</article-title>. <source>Genome Biol.</source> <volume>17</volume>, <fpage>184</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-016-1037-6</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumper</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Evans</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Pritzel</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Green</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Figurnov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ronneberger</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Highly Accurate Protein Structure Prediction with AlphaFold</article-title>. <source>Nature</source> <volume>596</volume>, <fpage>583</fpage>&#x2013;<lpage>589</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khan</surname>
<given-names>I. K.</given-names>
</name>
<name>
<surname>Jain</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Rawi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bensmail</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Prediction of Protein Group Function by Iterative Classification on Functional Relevance Network</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>1388</fpage>&#x2013;<lpage>1394</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty787</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Khan</surname>
<given-names>I. K.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Chapman</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kc</surname>
<given-names>D. B.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>The PFP and ESG Protein Function Prediction Methods in 2014: Effect of Database Updates and Ensemble Approaches</article-title>. <source>Gigascience</source> <volume>4</volume>, <fpage>43</fpage>. <pub-id pub-id-type="doi">10.1186/s13742-015-0083-4</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chikhi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Esquivel-Rodriguez</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Molecular Surface Representation Using 3D Zernike Descriptors for Protein Shape Comparison and Docking</article-title>. <source>Curr. Protein Pept. Sci.</source> <volume>12</volume>, <fpage>520</fpage>&#x2013;<lpage>530</lpage>. <pub-id pub-id-type="doi">10.2174/138920311796957612</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kulmanov</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hoehndorf</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>DeepGOPlus: Improved Protein Function Prediction from Sequence</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>422</fpage>&#x2013;<lpage>429</lpage>. <pub-id pub-id-type="doi">10.1093/BIOINFORMATICS/BTZ595</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>La</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Esquivel-Rodr&#xed;guez</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Venkatraman</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ueng</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>3D-SURFER: Software for High-Throughput Protein Surface Comparison and Analysis</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2843</fpage>&#x2013;<lpage>2844</lpage>. <pub-id pub-id-type="doi">10.1093/BIOINFORMATICS/BTP542</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Letunic</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Khedkar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bork</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>SMART: Recent Updates, New Developments and Status in 2020</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D458</fpage>&#x2013;<lpage>D460</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa937</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Pearson</surname>
<given-names>W. R.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Rapid and Sensitive Protein Similarity Searches</article-title>. <source>Science (1979)</source> <volume>227</volume>, <fpage>1435</fpage>&#x2013;<lpage>1441</lpage>. <pub-id pub-id-type="doi">10.1126/science.2983426</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maddhuri Venkata Subramaniya</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Terashi</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Jain</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Kagaya</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Protein Contact Map Refinement for Improving Structure Prediction Using Generative Adversarial Networks</article-title>. <source>Bioinformatics</source> <volume>37</volume> (<issue>19</issue>), <fpage>3168</fpage>&#x2013;<lpage>3174</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab220</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Malod-Dognin</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Pr&#x17e;ulj</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>GR-Align: Fast and Flexible Alignment of Protein 3D Structures Using Graphlet Degree Similarity</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>1259</fpage>&#x2013;<lpage>1265</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu020</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mirdita</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>von Den Driesch</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Galiez</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>S&#xf6;ding</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Steinegger</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Uniclust Databases of Clustered and Deeply Annotated Protein Sequences and Alignments</article-title>. <source>Nucleic Acids Res.</source> <volume>45</volume>, <fpage>D170</fpage>&#x2013;<lpage>D176</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw1081</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mistry</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chuguransky</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Qureshi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salazar</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Sonnhammer</surname>
<given-names>E. L. L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Pfam: The Protein Families Database in 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D412</fpage>&#x2013;<lpage>D419</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa913</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morgat</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Coissac</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Coudert</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Axelsen</surname>
<given-names>K. B.</given-names>
</name>
<name>
<surname>Keller</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bairoch</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>UniPathway: A Resource for the Exploration and Annotation of Metabolic Pathways</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>D761</fpage>&#x2013;<lpage>D769</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr1023</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nikolskaya</surname>
<given-names>A. N.</given-names>
</name>
<name>
<surname>Arighi</surname>
<given-names>C. N.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Barker</surname>
<given-names>W. C.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C. H.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>PIRSF Family Classification System for Protein Functional and Evolutionary Analysis</article-title>. <source>Evol. Bioinform Online</source> <volume>2</volume>, <fpage>197</fpage>&#x2013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.1177/117693430600200033</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Obayashi</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kagaya</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Aoki</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Tadaka</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kinoshita</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>COXPRESdb V7: A Gene Coexpression Database for 11 Animal Species Supported by 23 Coexpression Platforms for Technical Evaluation and Evolutionary Inference</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>D55</fpage>&#x2013;<lpage>D62</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1155</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pedruzzi</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Rivoire</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Auchincloss</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>Coudert</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Keller</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>de Castro</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>HAMAP in 2015: Updates to the Protein Family Classification and Annotation System</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>D1064</fpage>&#x2013;<lpage>D1070</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku1002</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pellegrini</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Marcotte</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Thompson</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Eisenberg</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Yeates</surname>
<given-names>T. O.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Assigning Protein Functions by Comparative Genome Analysis: Protein Phylogenetic Profiles</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>96</volume>, <fpage>4285</fpage>&#x2013;<lpage>4288</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.96.8.4285</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Radivojac</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>W. T.</given-names>
</name>
<name>
<surname>Oron</surname>
<given-names>T. R.</given-names>
</name>
<name>
<surname>Schnoes</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Wittkop</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sokolov</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>A Large-Scale Evaluation of Computational Protein Function Prediction</article-title>. <source>Nat. Methods</source> <volume>10</volume>, <fpage>221</fpage>&#x2013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2340</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chitale</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Structure- and Sequence-Based Function Prediction for Non-Homologous Proteins</article-title>. <source>J. Struct. Funct. Genomics</source> <volume>13</volume>, <fpage>111</fpage>&#x2013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1007/S10969-012-9126-6</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Detecting Local Ligand-Binding Site Similarity in Nonhomologous Proteins by Surface Patch Comparison</article-title>. <source>Proteins</source> <volume>80</volume>, <fpage>1177</fpage>&#x2013;<lpage>1195</lpage>. <pub-id pub-id-type="doi">10.1002/PROT.24018</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>La</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ramani</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Rustamov</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>Fast Protein Tertiary Structure Retrieval Based on Global Surface Shape Similarity</article-title>. <source>Proteins</source> <volume>72</volume>, <fpage>1259</fpage>&#x2013;<lpage>1273</lpage>. <pub-id pub-id-type="doi">10.1002/PROT.22030</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Characterization and Classification of Local Protein Surfaces Using Self-Organizing Map</article-title>. <source>Int. J. Knowl. Discov. Bioinforma. (IJKDB)</source> <volume>1</volume>, <fpage>32</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.4018/jkdb.2010100203</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sayers</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Beck</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bolton</surname>
<given-names>E. E.</given-names>
</name>
<name>
<surname>Bourexis</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Brister</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Canese</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Database Resources of the National Center for Biotechnology Information</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D10</fpage>&#x2013;<lpage>D17</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa892</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schlicker</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Domingues</surname>
<given-names>F. S.</given-names>
</name>
<name>
<surname>Rahnenf&#xfc;hrer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lengauer</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>A New Measure for Functional Similarity of Gene Products Based on Gene Ontology</article-title>. <source>BMC Bioinforma.</source> <volume>7</volume>, <fpage>302</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-7-302</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shin</surname>
<given-names>W. H.</given-names>
</name>
<name>
<surname>Christoffer</surname>
<given-names>C. W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>PL-PatchSurfer2: Improved Local Surface Matching-Based Virtual Screening Method that Is Tolerant to Target and Ligand Structure Variation</article-title>. <source>J. Chem. Inf. Model</source> <volume>56</volume>, <fpage>1676</fpage>&#x2013;<lpage>1691</lpage>. <pub-id pub-id-type="doi">10.1021/ACS.JCIM.6B00163</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sigrist</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>de Castro</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Cerutti</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cuche</surname>
<given-names>B. A.</given-names>
</name>
<name>
<surname>Hulo</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bridge</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>New and Continuing Developments at PROSITE</article-title>. <source>Nucleic Acids Res.</source> <volume>41</volume>, <fpage>D344</fpage>&#x2013;<lpage>D347</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks1067</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sit</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>W. H.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Three-Dimensional Krawtchouk Descriptors for Protein Local Surface Shape Comparison</article-title>. <source>Pattern Recognit.</source> <volume>93</volume>, <fpage>534</fpage>&#x2013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1016/J.PATCOG.2019.05.019</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Steinegger</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Meier</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mirdita</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>V&#xf6;hringer</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Haunsberger</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>S&#xf6;ding</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>HH-Suite3 for Fast Remote Homology Detection and Deep Protein Annotation</article-title>. <source>BMC Bioinforma.</source> <volume>20</volume>, <fpage>473</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-019-3019-7</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Steinegger</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>S&#xf6;ding</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>MMseqs2 Enables Sensitive Protein Sequence Searching for the Analysis of Massive Data Sets</article-title>. <source>Nat. Biotechnol.</source> <volume>35</volume> (<issue>11</issue>), <fpage>1026</fpage>&#x2013;<lpage>1028</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3988</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Subbarao</surname>
<given-names>G. V.</given-names>
</name>
<name>
<surname>van den Berg</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Crystal Structure of the Monomeric Porin OmpG</article-title>. <source>J. Mol. Biol.</source> <volume>360</volume>, <fpage>750</fpage>&#x2013;<lpage>759</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmb.2006.05.045</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzek</surname>
<given-names>B. E.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>McGarvey</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C. H.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>UniRef Clusters: A Comprehensive and Scalable Alternative for Improving Sequence Similarity Searches</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>926</fpage>&#x2013;<lpage>932</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu739</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Szklarczyk</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gable</surname>
<given-names>A. L.</given-names>
</name>
<name>
<surname>Lyon</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Junge</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wyder</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Huerta-Cepas</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>STRING V11: Protein-Protein Association Networks with Increased Coverage, Supporting Functional Discovery in Genome-Wide Experimental Datasets</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>D607</fpage>&#x2013;<lpage>D613</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1131</pub-id> </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<collab>The UniProt Consortium</collab> (<year>2021</year>). <article-title>UniProt: The Universal Protein Knowledgebase in 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D480</fpage>&#x2013;<lpage>D489</lpage>. <pub-id pub-id-type="doi">10.1093/NAR/GKAA1100</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Venkatraman</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Sael</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Potential for Protein Surface Shape Analysis Using Spherical Harmonics and 3D Zernike Descriptors</article-title>. <source>Cell Biochem. Biophys.</source> <volume>54</volume>, <fpage>23</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1007/S12013-009-9051-X</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wan</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Jones</surname>
<given-names>D. T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Protein Function Prediction Is Improved by Creating Synthetic Feature Samples with Generative Adversarial Networks</article-title>. <source>Nat. Mach. Intell.</source> <volume>2</volume> (<issue>9</issue>), <fpage>540</fpage>&#x2013;<lpage>550</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-0222-1</pub-id> </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Distance-Based Protein Folding Powered by Deep Learning</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>116</volume>, <fpage>16856</fpage>&#x2013;<lpage>16865</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1821309116</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Anishchenko</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ovchinnikov</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Baker</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Improved Protein Structure Prediction Using Predicted Interresidue Orientations</article-title>. <source>Proc. Natl. Acad. Sci. U. S. A.</source> <volume>117</volume>, <fpage>1496</fpage>&#x2013;<lpage>1503</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1914677117</pub-id> </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>You</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Mamitsuka</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>DeepGraphGO: Graph Neural Network for Large-Scale, Multispecies Protein Function Prediction</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>i262</fpage>&#x2013;<lpage>i271</lpage>. <pub-id pub-id-type="doi">10.1093/BIOINFORMATICS/BTAB270</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>You</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mamitsuka</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>NetGO: Improving Large-Scale Protein Function Prediction with Massive Network Information</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>W379</fpage>&#x2013;<lpage>W387</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz388</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Effective Inter-Residue Contact Definitions for Accurate Protein Fold Recognition</article-title>. <source>BMC Bioinforma.</source> <volume>13</volume>, <fpage>292</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-13-292</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Bergquist</surname>
<given-names>T. R.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Kacsoh</surname>
<given-names>B. Z.</given-names>
</name>
<name>
<surname>Crocker</surname>
<given-names>A. W.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>The CAFA Challenge Reports Improved Protein Function Prediction and New Functional Annotations for Hundreds of Genes through Experimental Screens</article-title>. <source>Genome Biol.</source> <volume>20</volume>, <fpage>244</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1835-8</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Shin</surname>
<given-names>W. H.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Combined Approach of Patch-Surfer and PL-PatchSurfer for Protein-Ligand Binding Prediction in CSAR 2013 and 2014</article-title>. <source>J. Chem. Inf. Model</source> <volume>56</volume>, <fpage>1088</fpage>&#x2013;<lpage>1099</lpage>. <pub-id pub-id-type="doi">10.1021/ACS.JCIM.5B00625</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Xiong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kihara</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Large-Scale Binding Ligand Prediction by Improved Patch-Based Method Patch-Surfer2.0</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>707</fpage>&#x2013;<lpage>713</lpage>. <pub-id pub-id-type="doi">10.1093/BIOINFORMATICS/BTU724</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>