<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">839453</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.839453</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>PPA-GCN: A Efficient GCN Framework for Prokaryotic Pathways Assignment</article-title>
<alt-title alt-title-type="left-running-head">Lu et al.</alt-title>
<alt-title alt-title-type="right-running-head">GCN Framework Improve Pathway Assignment</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Lu</surname>
<given-names>Yuntao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1562152/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Qi</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/232110/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Li</surname>
<given-names>Tao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/232687/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Key Laboratory of Freshwater Ecology and Biotechnology</institution>, <institution>Institute of Hydrobiology</institution>, <institution>Chinese Academy of Sciences</institution>, <addr-line>Wuhan</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>College of Advanced Agricultural Sciences</institution>, <institution>University of Chinese Academy of Sciences</institution>, <addr-line>Beijing</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/124337/overview">Marco Pellegrini</ext-link>, Italian National Research Council, Italy</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1572964/overview">Massimo La Rosa</ext-link>, National Research Council (CNR), Italy</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/958003/overview">Beifang Niu</ext-link>, Chinese Academy of Sciences (CAS), China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Qi Li, <email>liqi@ihb.ac.cn</email>; Tao Li, <email>litao@ihb.ac.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>839453</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>17</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Lu, Li and Li.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Lu, Li and Li</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>With the rapid development of sequencing technology, completed genomes of microbes have explosively emerged. For a newly sequenced prokaryotic genome, gene functional annotation and metabolism pathway assignment are important foundations for all subsequent research work. However, the assignment rate for gene metabolism pathways is lower than 48% on the whole. It is even lower for newly sequenced prokaryotic genomes, which has become a bottleneck for subsequent research. Thus, the development of a high-precision metabolic pathway assignment framework is urgently needed. Here, we developed PPA-GCN, a prokaryotic pathways assignment framework based on graph convolutional network, to assist functional pathway assignments using KEGG information and genomic characteristics. In the framework, genomic gene synteny information was used to construct a network, and ideas of self-supervised learning were inspired to enhance the framework&#x2019;s learning ability. Our framework is applicable to the genera of microbe with sufficient whole genome sequences. To evaluate the assignment rate, genomes from three different genera (<italic>Flavobacterium</italic> (65 genomes) and <italic>Pseudomonas</italic> (100 genomes), <italic>Staphylococcus</italic> (500 genomes)) were used. The initial functional pathway assignment rate of the three test genera were 27.7% (<italic>Flavobacterium</italic>), 49.5% (<italic>Pseudomonas</italic>) and 30.1% (<italic>Staphylococcus</italic>). PPA-GCN achieved excellence performance of 84.8% (<italic>Flavobacterium</italic>), 77.0% (<italic>Pseudomonas</italic>) and 71.0% (<italic>Staphylococcus</italic>) for assignment rate. At the same time, PPA-GCN was proved to have strong fault tolerance. The framework provides novel insights into assignment for metabolism pathways and is likely to inform future deep learning applications for interpreting functional annotations and extends to all prokaryotic genera with sufficient genomes.</p>
</abstract>
<kwd-group>
<kwd>graph convolution network</kwd>
<kwd>prokaryotic genome</kwd>
<kwd>metabolic pathway</kwd>
<kwd>deep learning</kwd>
<kwd>self supervised</kwd>
</kwd-group>
<contract-num rid="cn001">No. 91851118 No.31900096</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>Introduction</title>
<p>With the rapid development of sequencing technology, the number of newly released prokaryotic genomes has exploded, providing an important foundation for subsequent research work (<xref ref-type="bibr" rid="B10">Doerks et al., 2004</xref>). Functional annotation and pathway assignment are important components of understanding the details of metabolism. Accordingly, a series of reference genome databases and functional annotation platforms have been developed (<xref ref-type="bibr" rid="B5">Benson et al., 2012</xref>; <xref ref-type="bibr" rid="B14">Federhen, 2012</xref>; <xref ref-type="bibr" rid="B22">Keegan et al., 2016</xref>; <xref ref-type="bibr" rid="B7">Chen et al., 2019</xref>; <xref ref-type="bibr" rid="B4">Bazgir et al., 2020</xref>). The Kyoto Encyclopedia of Genes and Genomes (KEGG) is one of the most widely used and reliable functional platforms, and it provides three annotation software tools, namely, BlastKOALA, GhostKOALA, and KofamKOALA, for functional annotation (<xref ref-type="bibr" rid="B37">Suzuki et al<italic>.</italic>, 2014</xref>; <xref ref-type="bibr" rid="B21">Kanehisa et al., 2016a</xref>; <xref ref-type="bibr" rid="B20">Kanehisa et al., 2016b</xref>; <xref ref-type="bibr" rid="B3">Aramaki et al<italic>.</italic>, 2020</xref>). Currently, only 48% of the protein sequences are assigned to pathways in the KEGG GENES database (<xref ref-type="bibr" rid="B3">Aramaki et al<italic>.</italic>, 2020</xref>). It is even lower for newly sequenced prokaryotic genomes, which has become a bottleneck for subsequent research (<xref ref-type="bibr" rid="B37">Suzuki et al<italic>.</italic>, 2014</xref>). Thus, the development of a high-precision metabolic pathway assignment framework is urgently needed.</p>
<p>Here, we propose PPA-GCN, a framework based on graph convolutional network (GCN) that uses genomic gene synteny information within specific genus, from which the graph topological pattern and gene node characteristics can be learned, to disseminate node attributes in the network and provide assistance to the assignment of metabolic pathways. Synteny is defined as two or more pairs of homologous genes occupying the same chromosomal segment, where homologous loci are defined based on the similarity of function of the products of the corresponding genes (<xref ref-type="bibr" rid="B28">Nadeau and Taylor, 1984</xref>). Analyzing synteny can provide insight regarding the evolution and function of genes (<xref ref-type="bibr" rid="B42">Zhang et al., 2016</xref>). As an inherent biological attribute, bacteria of different genera have different synteny patterns. In general, bacterial genomes have two different pan-genome types. The pan-genome refers to all genes detected in a whole group of genomes (<xref ref-type="bibr" rid="B39">Wang L. et al., 2020</xref>). Some prokaryotes have genomes with highly conserved gene content (closed pan-genomes), while others are more flexible (open pan-genomes). Since the concept of a &#x201c;pan-genome&#x201d; was first proposed in 2005, pan-genome analysis has revealed the diversity and evolution of bacterial genomes (<xref ref-type="bibr" rid="B38">Tettelin et al., 2005</xref>). In present, there is currently no deep learning framework for direct assignment of functional pathways against KEGG database. To evaluate PPA-GCN, genome datasets of three different genus were used, and on all of them, the proposed framework had achieved excellent performance. PPA-GCN enables novel insights into assignment for functional pathways and is likely to inform future deep learning applications for interpreting functional annotations.</p>
</sec>
<sec id="s2">
<title>Related Work</title>
<p>The study of gene location in the genome is one of the classic fields of genetics (<xref ref-type="bibr" rid="B31">Rogozin et al<italic>.,</italic> 2004</xref>). In prokaryotes, genes encoding functional linked proteins are usually organized into gene clusters (<xref ref-type="bibr" rid="B36">Shmakov et al.<italic>,</italic> 2019</xref>). There were methods assign protein function using neighborhood properties (<xref ref-type="bibr" rid="B32">Saha et al.<italic>,</italic> 2012</xref>; <xref ref-type="bibr" rid="B19">Jun et al., 2017</xref>; <xref ref-type="bibr" rid="B33">Saha et al<italic>.</italic>, 2018</xref>). It has been shown that the neighborhood milieu of genes in a network can assist in predicting the probable function of a gene for which no function is known (<xref ref-type="bibr" rid="B17">Hao et al., 2012</xref>). However, there is almost no method to assign KEGG pathways using gene neighborhood information.</p>
<p>In recent years, deep learning has been widely used in the field of life science, for example, for identifying and interpreting the contextual features of transcription factors (<xref ref-type="bibr" rid="B44">Zheng et al., 2021</xref>), generating functional protein sequences (<xref ref-type="bibr" rid="B29">Repecka et al., 2021</xref>), and identifying cell types (<xref ref-type="bibr" rid="B26">Lukassen et al<italic>.</italic>, 2020</xref>; <xref ref-type="bibr" rid="B40">Wang M. et al., 2020</xref>). At present, the applications of graph neural networks in the medical and biology fields show strong representation and integration capabilities (<xref ref-type="bibr" rid="B41">Wu et al<italic>.</italic>, 2020</xref>), including neuroimage analysis (<xref ref-type="bibr" rid="B43">Zhang et al., 2018</xref>), disease gene identification (<xref ref-type="bibr" rid="B24">Li et al., 2019</xref>; <xref ref-type="bibr" rid="B34">Schulte-Sasse et al., 2021</xref>), drug combination synergy prediction (<xref ref-type="bibr" rid="B45">Zitnik et al., 2018</xref>; <xref ref-type="bibr" rid="B18">Jiang et al<italic>.</italic>, 2020</xref>; <xref ref-type="bibr" rid="B12">Manoochehri and Nourani, 2020</xref>), discovery of disease pathways (<xref ref-type="bibr" rid="B1">Agrawal et al<italic>.</italic>, 2018</xref>), prediction of tissue cell function (<xref ref-type="bibr" rid="B46">Zitnik et al., 2017</xref>), pseudogene function prediction (<xref ref-type="bibr" rid="B13">Fan and Zhang, 2020</xref>), conducting taxonomic classification for phage contigs (<xref ref-type="bibr" rid="B35">Shang et al., 2021</xref>) and identifying missing protein&#x2013;phenotype associations (<xref ref-type="bibr" rid="B25">Liu et al., 2021</xref>). The graph convolutional network (GCN) is a type of graph neural network that can learn the structure of a graph. This network model was originally proposed for semi-supervised classification (<xref ref-type="bibr" rid="B23">Kipf et al., 2016</xref>). A GCN model can extensively integrate graph topological features and node information by defining each node as a computational graph and using neural networks to integrate neighbor node information.</p>
</sec>
<sec sec-type="materials|methods" id="s3">
<title>Materials and Methods</title>
<sec id="s3-1">
<title>Problem Statement</title>
<p>Given an undirected graph G &#x3d; (V<sub>tr</sub>, V<sub>te</sub>, E), where V<sub>tr</sub> is the set of nodes that assigned function pathway, V<sub>te</sub> is the set of nodes that unassigned function pathway, V &#x3d; {V<sub>tr</sub>, V<sub>te</sub>}. E is the set of edges and the edge represents two genes belonging to different nodes are connected in the genome. A label set L &#x3d; {l<sub>1</sub>, l<sub>2</sub>...l<sub>k</sub>} is formed according to the KEGG secondary class. The relationship between the node set and the label set is represented by a matrix Y<sub>NxK</sub>. Y<sub>ij</sub> &#x3d; 1, if there is a gene in node i has assigned to label j. Our goal is to assign the possible pathway labels to those nodes that have no labels.</p>
</sec>
<sec id="s3-2">
<title>Framework</title>
<p>PPA-GCN is a deep learning framework based on a graph convolutional model <bold>(</bold>
<xref ref-type="fig" rid="F1">Figure 1</xref>). Gene synteny information from the selected genome is used to construct edges in a network, while genes sharing high sequence similarity and cover ratio are grouped into nodes. All node and edge information are used to construct the gene synteny network. PPA-GCN applies a three-layer graph convolutional architecture. Input features include node encoding, node scale and adjacency probability matrix. The KEGG metabolic pathway information of the secondary class is used as the node labels for initial training. Improve performance with inspiration from self-supervised learning. The final outputs are ranked in accordance with the stability of the assignment during the training process.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>PPA-GCN architecture. The input to the framework is the metabolic pathway network extracted from the KEGG metabolic pathways and the gene synteny network composed of the prokaryotic genomes. The graph convolutional layer attempts to construct a mapping relationship between the two input networks and iteratively uses the training results to update the input inspired by self-supervised learning until a steady state is reached and the final assignment output is obtained.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g001.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>Graph Construction</title>
<sec id="s3-3-1">
<title>Node Construction</title>
<p>Blast (<xref ref-type="bibr" rid="B2">Altschul et al., 1990</xref>) was used to compare the sequence similarity of all genome genes in one genus. In order to quickly and strictly find the similar genes, we directly adopted the reciprocal best hits comparison and controlled the identities and cover ratios to 65%. Taking <italic>Flavobacterium</italic> as an example, a total of 16,830 orthologs were obtained using OrthoFinder 2.0 (<xref ref-type="bibr" rid="B11">Emms and Kelly, 2019</xref>), and 51,247 nodes were obtained using our method, of which 50,998 nodes contained only one orthologs (99.5%). Therefore, our method is stricter than directly using orthologs. Node2vec algorithm (<xref ref-type="bibr" rid="B16">Grover and Leskovec, 2016</xref>) was used to generate graph embeddings for each node.</p>
</sec>
<sec id="s3-3-2">
<title>Edge Construction</title>
<p>Positional relationship pairs between two genes from each genome were constructed using the data of coding DNA sequence (CDS) (<xref ref-type="fig" rid="F2">Figure 2</xref>). Through the correspondence between genes and nodes, all positional pairs were connected into a single gene synteny network, in which there could be more than one connection between two nodes. The adjacency matrix was constructed in accordance with the number of connections between nodes.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Schematic diagram of the use of multiple genomes to construct a gene synteny network. First, all genomic genes are compared for sequence similarity, and genes that share high reciprocal similarity and cover ratios are assigned the same node id. Then, positional relationship pairs between two genes from each genome were constructed. Finally, all gene position relationship pairs are connected into a gene synteny network.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g002.tif"/>
</fig>
</sec>
</sec>
<sec id="s3-4">
<title>Construction of the Adjacency Probability Matrix</title>
<p>The adjacency probability is defined as the probability that two nodes form a certain number of connections in the network. First, the degree of each node in the gene synteny network (the number of connections by which a node is directly connected to surrounding nodes) was calculated. Then, the probability <italic>P</italic>
<sub>
<italic>i</italic>
</sub> that an edge is connected to a specific node <italic>i</italic> was calculated. Finally, the probability that there are <italic>k</italic> edges between node <italic>i</italic> and node <italic>j</italic> was defined as:<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mstyle displaystyle="true">
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mstyle>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m2">
<mml:mrow>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:msubsup>
<mml:mi>P</mml:mi>
<mml:mi>j</mml:mi>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>P</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>g</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <italic>N</italic> is the total number of nodes in the gene synteny graph and <italic>degree(i)</italic> is the degree of node <italic>i</italic>, <italic>C</italic> is the combination symbol.</p>
<p>After the adjacency probabilities of all nodes had been formed into an <italic>N&#x2a;N</italic> adjacency probability matrix, because there are no connections between most nodes, the node2vec algorithm was used to densify the adjacency probability matrix.</p>
</sec>
<sec id="s3-5">
<title>The GCN Model</title>
<sec id="s3-5-1">
<title>Framework Architecture</title>
<p>Given an undirected graph with node feature matrix <italic>X</italic> and adjacency matrix <italic>A</italic>, the graph convolution operation (Kipf et al<italic>.</italic>, 2016) is defined as:<disp-formula id="e3">
<mml:math id="m3">
<mml:mrow>
<mml:mi>H</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>A</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mi>X</mml:mi>
<mml:mi>W</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>
<disp-formula id="e4">
<mml:math id="m4">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mi>A</mml:mi>
<mml:mo>&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>I</mml:mi>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>j</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:msub>
<mml:mi>A</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="true">&#x5e;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>I</italic> is the identity matrix, <italic>W</italic> is the matrix of trainable weights in the neural network, <italic>X</italic> is the feature matrix before the update, <italic>H</italic> is the feature matrix after the update, and <italic>&#x3c3;</italic> is the activation function (ReLU). The graph convolution operation iteratively calculates the weighted average of the node attributes of the neighbors of the current node to obtain the new feature matrix of the node. In this framework, the features of unlabeled nodes (nodes without assigned functional pathways) and the features of nearby labeled nodes (nodes with assigned functional pathways) are mixed to be propagated through the synteny network diagram. If two nodes have the same neighbor structure and neighbor features, their embedded feature matrix <italic>H</italic> will be exactly the same.</p>
<p>Python&#x2019;s PyTorch Geometric Module was used to implement PPA-GCN. Multiple graph convolutional layers can be stacked to enable learning on a larger domain structure. After testing, a three-layer stack was found to perform the best. The two-class cross entropy was used as the loss function because of the multilabel nature of the problem.</p>
</sec>
<sec id="s3-5-2">
<title>Self-Supervised Learning Inspiration</title>
<p>The original input was fed into the framework, and 50 epochs of random sampling verification training were performed with the test set. The nodes with an average cross-validation accuracy rate of less than 30% are removed from the training set, and nodes and labels with a assignment stability of 90% in the test set (that is, the same label is assigned more than 45&#xa0;times) are added to the training set. After many iterations, when the number of nodes in the training set reached more than 90% of the total number of nodes in the gene synteny network, the training was considered to have reached a stable state, and the final assignment results were output.</p>
</sec>
</sec>
<sec id="s3-6">
<title>Topological Analysis</title>
<sec id="s3-6-1">
<title>Degree and Degree Distribution</title>
<p>The degree is defined as the number of all edge connections of a node in a graph, describing the first-order connection degree of the node. The degree distribution is an overall description of the nodes in a network, that is, the probability distribution or statistical distribution of the node degrees.</p>
</sec>
<sec id="s3-6-2">
<title>Clustering Coefficient</title>
<p>The clustering coefficient is used to describe the degree of clumping among the vertices of a network. Specifically, it is the degree of interconnection among the adjacent nodes of a node, describing the second-order connection degree of the node. For node <italic>i</italic> with degree <italic>k</italic>
<sub>
<italic>i</italic>
</sub>, the local clustering coefficient is defined as:<disp-formula id="e5">
<mml:math id="m5">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mi>L</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>k</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where <italic>L</italic>
<sub>
<italic>i</italic>
</sub> is the number of connections among the <italic>k</italic>
<sub>
<italic>i</italic>
</sub> neighbors of node <italic>i</italic>. The overall aggregation coefficient of the network is characterized as the average value of the aggregation coefficients of all nodes.</p>
</sec>
</sec>
</sec>
<sec sec-type="results" id="s4">
<title>Result</title>
<sec id="s4-1">
<title>Data</title>
<p>All training genomes were downloaded from the National Center for Biotechnology Information (NCBI) database in June 2021 (<ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/genome/browse">https://www.ncbi.nlm.nih.gov/genome/browse&#x23;!/overview/</ext-link>). The datasets include <italic>Flavobacterium</italic> (Gram-negative, 65 genomes), <italic>Pseudomonas</italic> (Gram-negative, 100 genomes) and <italic>Staphylococcus</italic> (Gram-positive, 500 genomes). <italic>Staphylococcus</italic> has a closed pan-genome. The 500 genomes selected for this study contain 1,332,382 genes grouped into a gene synteny network of 10,074 nodes. <italic>Flavobacterium</italic> and <italic>Pseudomonas</italic> have open pan-genomes. The 65 <italic>Flavobacterium</italic> genomes and 100 <italic>Pseudomonas</italic> genomes selected for this study contain 243,834 and 550,752 genes grouped into 51,247 and 79,941 nodes, respectively.</p>
<p>KEGG internal annotation tool KofamKOALA (version 100.0, updated October 1, 2021) was used to assign genes to functional pathways. The pathway labels belonging to the global and overview maps category were removed. <italic>Staphylococcus</italic> had 400,478 genes (1,324 nodes) assigned to metabolic pathways, <italic>Flavobacterium</italic> had 67,529 genes (3,694 nodes) assigned to metabolic pathways, and <italic>Pseudomonas</italic> had 272,388 genes (12,429 nodes) assigned to metabolic pathways (<xref ref-type="sec" rid="s11">Supplementary Table S1</xref>). The original assignment rates for the three genera were 7.2% (<italic>Flavobacterium</italic>), 15.5% (<italic>Pseudomonas</italic>) and 13.1% (<italic>Staphylococcus</italic>).</p>
<p>In order to verify the performance of the model, the new genome data of the three genera were downloaded from the National Center for Biotechnology Information (NCBI) database in October 2021 (newly released genomes were downloaded first). The datasets include <italic>Flavobacterium</italic> (30 genomes), <italic>Pseudomonas</italic> (50 genomes) and <italic>Staphylococcus</italic> (200 genomes).</p>
</sec>
<sec id="s4-2">
<title>Evaluation Metrics</title>
<p>Pathway label assignment is essentially a multilabel classification problem. Hence, some commonly used evaluation indicators for binary classification problems are not suitable for PPA-GCN. We use six indicators to measure the effectiveness of the framework:</p>
<sec id="s4-2-1">
<title>Prediction Rate of Assignment</title>
<p>PRA is the accuracy at the node level and is defined as the proportion of genes with at least one label assigned correctly.</p>
</sec>
<sec id="s4-2-2">
<title>Total Label Prediction Rate</title>
<p>The TLPR is the accuracy at the label level and is defined as the number of correctly assigned labels divided by the total number of labels.</p>
</sec>
<sec id="s4-2-3">
<title>Weighted Prediction Rate of Assignment</title>
<p>When a label is predicted for a node, we assign weights in accordance with the assignment probability, sum the WPRA of each label of a node to obtain the <italic>WPRA</italic> of that node, and divide by the total number of nodes to obtain the overall WPRA:<disp-formula id="e6">
<mml:math id="m6">
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <italic>N</italic> is the total number of nodes, <italic>I</italic> is the number of labels for node <italic>i</italic>, <italic>T</italic>
<sub>
<italic>i</italic>
</sub> is the order of the correct label probabilities assigned for node <italic>i</italic> (from large to small), and <italic>k</italic> is the k-th ranked probability label that was assigned correctly.</p>
</sec>
<sec id="s4-2-4">
<title>Kappa Coefficient</title>
<p>The kappa coefficient is often used for testing consistency, that is, whether the assignment effect of the model is consistent with the actual classification effect. Its value is between -1 and 1. When the value is greater than 0.6, it is considered substantial, and when it is greater than 0.8, it is considered almost perfect. The calculation of the kappa coefficient is based on the confusion matrix:<disp-formula id="e7">
<mml:math id="m7">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>a</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>
<disp-formula id="e8">
<mml:math id="m8">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mn>0</mml:mn>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>.</mml:mo>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mo>.</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>M</italic> is the confusion matrix of the assignment results.</p>
</sec>
<sec id="s4-2-5">
<title>Hamming Distance</title>
<p>The Hamming distance is measure of the distance between the assigned and real labels, with a value between 0 and 1. A distance of 0 means that the assigned results are exactly the same as the real results, and a distance of 1 means that the model&#x2019;s results are completely opposite to the desired results. This indicator is calculated as the number of erroneously assigned labels divided by the total number of labels.</p>
</sec>
<sec id="s4-2-6">
<title>Jaccard Similarity Coefficient</title>
<p>This coefficient is an indicator for comparing the similarity of two finite sets, defined as the size of the intersection of two label sets (the true label set and the assigned label set) divided by the size of the union. When this coefficient is 1, the assigned results are completely consistent with the actual situation; when the coefficient is 0, the assigned results are completely inconsistent with the actual situation.</p>
</sec>
</sec>
<sec id="s4-3">
<title>Results of Experiments</title>
<sec id="s4-3-1">
<title>Results of Cross-Validation</title>
<p>We tested PPA-GCN with 5-fold cross-validation on three data sets. PPA-GCN achieved prediction rates of assignment (PRAs) of 84.8% (<italic>Flavobacterium</italic>), 77.0% (<italic>Pseudomonas</italic>) and 71.0% (<italic>Staphylococcus</italic>) on the three prokaryotic bacterial genera (<xref ref-type="fig" rid="F3">Figure 3</xref>). According to the evaluation index results (<xref ref-type="table" rid="T1">Table 1</xref>), PPA-GCN is well adapted to all three genera.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The performance of PPA-GCN on three genera (in terms of the PRA) and the node scale distribution of the node set at each PRA level (10% as one level). From left to right are <italic>Flavobacterium, Pseudomonas,</italic> and <italic>Staphylococcus</italic>.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g003.tif"/>
</fig>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Performance under 5-fold cross-validation for the three genera.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Species</th>
<th align="center">PRA</th>
<th align="center">TLPR</th>
<th align="center">WPRA</th>
<th align="center">KC</th>
<th align="center">HD</th>
<th align="center">JS</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<italic>Flavobacterium</italic>
</td>
<td align="char" char=".">0.848</td>
<td align="char" char=".">0.846</td>
<td align="char" char=".">0.829</td>
<td align="char" char=".">0.842</td>
<td align="char" char=".">0.008</td>
<td align="char" char=".">0.751</td>
</tr>
<tr>
<td align="left">
<italic>Pseudomonas</italic>
</td>
<td align="char" char=".">0.770</td>
<td align="char" char=".">0.728</td>
<td align="char" char=".">0.736</td>
<td align="char" char=".">0.721</td>
<td align="char" char=".">0.014</td>
<td align="char" char=".">0.609</td>
</tr>
<tr>
<td align="left">
<italic>Staphylococcus</italic>
</td>
<td align="char" char=".">0.710</td>
<td align="char" char=".">0.691</td>
<td align="char" char=".">0.698</td>
<td align="char" char=".">0.689</td>
<td align="char" char=".">0.008</td>
<td align="char" char=".">0.651</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, we compared PPA-GCN with five other machine learning methods. deepNF (<xref ref-type="bibr" rid="B15">Gligorijevic et al., 2018</xref>), Mashup (<xref ref-type="bibr" rid="B8">Cho et al., 2016</xref>) and Pseudo2GO (<xref ref-type="bibr" rid="B13">Fan and Zhang, 2020</xref>) are three deep learning methods that use graph information for function prediction. Support vector machines (SVM) and deep neural networks (DNN) are two machine learning models that are not based on graph information. Using the <italic>Staphylococcus</italic> genome as the test data set, all methods use the same features in PPA-GCN as input, and use 5-fold cross-validation to test performance. The results (<xref ref-type="table" rid="T2">Table 2</xref>) show that, PPA-GCN achieves the best performance among all indicators.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Performance comparison under 5-fold cross-validation</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Methods</th>
<th align="center">PRA</th>
<th align="center">TLPR</th>
<th align="center">WPRA</th>
<th align="center">KC</th>
<th align="center">HD</th>
<th align="center">JS</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">deepNF</td>
<td align="char" char=".">0.562</td>
<td align="char" char=".">0.365</td>
<td align="char" char=".">0.339</td>
<td align="char" char=".">0.511</td>
<td align="char" char=".">0.273</td>
<td align="char" char=".">0.379</td>
</tr>
<tr>
<td align="left">Mashup</td>
<td align="char" char=".">0.562</td>
<td align="char" char=".">0.446</td>
<td align="char" char=".">0.479</td>
<td align="char" char=".">0.529</td>
<td align="char" char=".">0.108</td>
<td align="char" char=".">0.450</td>
</tr>
<tr>
<td align="left">Pseudo2GO</td>
<td align="char" char=".">0.578</td>
<td align="char" char=".">0.470</td>
<td align="char" char=".">0.466</td>
<td align="char" char=".">0.513</td>
<td align="char" char=".">0.051</td>
<td align="char" char=".">0.433</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="char" char=".">0.483</td>
<td align="char" char=".">0.304</td>
<td align="char" char=".">0.319</td>
<td align="char" char=".">0.506</td>
<td align="char" char=".">0.118</td>
<td align="char" char=".">0.414</td>
</tr>
<tr>
<td align="left">DNN</td>
<td align="char" char=".">0.402</td>
<td align="char" char=".">0.365</td>
<td align="char" char=".">0.339</td>
<td align="char" char=".">0.501</td>
<td align="char" char=".">0.063</td>
<td align="char" char=".">0.429</td>
</tr>
<tr>
<td align="left">PPA-GCN (without self-supervised learning)</td>
<td align="char" char=".">0.607</td>
<td align="char" char=".">0.570</td>
<td align="char" char=".">0.539</td>
<td align="char" char=".">0.522</td>
<td align="char" char=".">0.034</td>
<td align="char" char=".">0.402</td>
</tr>
<tr>
<td align="left">PPA-GCN</td>
<td align="char" char=".">0.710</td>
<td align="char" char=".">0.691</td>
<td align="char" char=".">0.698</td>
<td align="char" char=".">0.689</td>
<td align="char" char=".">0.008</td>
<td align="char" char=".">0.651</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-3-2">
<title>Results of Test</title>
<p>In order to evaluate the adaptability of PPA-GCN to new data, the genes of the new genome were classified into network nodes. The test set node of the newly assigned functional path label in the network was used as the evaluation object, and the difference between the assigned output label and the real label is directly compared. The results are shown in <xref ref-type="table" rid="T3">Table 3</xref>, which proves that the results of PPA-GCN is reliable.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Performance under the new data set.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Metrics</th>
<th align="center">
<italic>Flavobacterium</italic>
</th>
<th align="center">
<italic>Pseudomonas</italic>
</th>
<th align="center">
<italic>Staphylococcus</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PRA</td>
<td align="char" char=".">0.637</td>
<td align="char" char=".">0.613</td>
<td align="char" char=".">0.798</td>
</tr>
<tr>
<td align="left">TLPR</td>
<td align="char" char=".">0.606</td>
<td align="char" char=".">0.538</td>
<td align="char" char=".">0.723</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-3-3">
<title>Fault Tolerance Evaluation</title>
<p>Because functional pathway assignment for bacterial genomes is still in the development stage, there will inevitably be some false pathway labels on the bacterial genes. Hence, we needed to test the fault tolerance of PPA-GCN. All assigned labels were assumed to be correct. In each epoch of training, some unlabeled nodes were given random labels to also participate in the training process. Two sets of experiments were conducted. In one, a certain percentage (5&#x2013;20%) of incorrectly labeled samples were added in each epoch independently, and in the other, incorrectly labeled samples were added accumulatively. The PRA without the addition of incorrect labels was taken as the standard, and the PRA after the addition of incorrect labels was divided by the standard PRA to serve as the performance indicator. The results (<xref ref-type="fig" rid="F4">Figure 4</xref>, <xref ref-type="sec" rid="s11">Supplementary Table S2</xref>) show that PPA-GCN can still maintain more than 75% performance with the addition of incorrect labels at a rate of up to 100% (that is, the incorrectly labeled samples compose up to 50% of the training set). Because the distribution of wrong labels is random, and the distribution of correct labels is ordered, the influence of correct labels on the training results is greater than that of wrong labels, which enhances the fault tolerance of the framework. With an increasing proportion of incorrect labels, the efficiency of the framework did not drop sharply. This result shows that PPA-GCN has strong fault tolerance.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Framework fault tolerance evaluation. On the three datasets, the performance was tested with the accumulation of 5&#x2013;20% incorrectly labeled data in each epoch; the horizontal axis is the number of iteration, and the vertical axis is the performance indicator (current PRA/original PRA). This result shows that PPA-GCN has strong fault tolerance.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g004.tif"/>
</fig>
</sec>
<sec id="s4-3-4">
<title>Feature Importance Test</title>
<p>A graph neural network can achieve excellent prediction accuracy, but it is difficult to give practical meaning to features. To evaluate the importance of the selected features, the PRAs before and after feature removal were compared (<xref ref-type="fig" rid="F5">Figure 5</xref>, <xref ref-type="sec" rid="s11">Supplementary Table S3</xref>). There are three important features in the PPA-GCN input: the node scale, the adjacency probability matrix and the gene synteny network.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>
<bold>(A)</bold> Feature importance assessment of the node scale. Comparison of performance changes before (standard) and after removing node scale (no scale). <bold>(B)</bold> Feature importance evaluation of the probability adjacency matrix. Comparison of performance changes before (standard) and after removing probability adjacency matrix (no pro). <bold>(C)</bold> Feature importance assessment of the gene synteny network. The horizontal axis represents standard training and training using random networks generated with three strategies: not including any adjacency probability matrix (with no pro), including the adjacency probability matrix of the newly generated network (with pro), and including the adjacency probability matrix of the real network (with true pro). The three graphs all use the prediction rate of assignment (PRA) as the evaluation index. The results show that node scale, adjacency probability matrix and network are very important features of PPA-GCN.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g005.tif"/>
</fig>
<p>The node scale is defined as the number of genes grouped into one node. The node scale was selected as an input feature because it can reflect the characteristics of a group of genomes. <italic>Staphylococcus</italic> has a closed pan-genome with an average node scale of 132.3, that is, an average of approximately 132 genes grouped into one node. <italic>Flavobacterium</italic> and <italic>Pseudomonas</italic> have open pan-genomes with average node scales of only 4.8 and 6.9, respectively. The node scale was one of the major observed differences between the labeled (training set) and unlabeled (test set) node sets in the gene synteny network. PPA-GCN showed no significant difference in performance when the node scale information was removed from the input (<xref ref-type="fig" rid="F5">Figure 5A</xref>). The node scale has no effect on framework training, and this is beneficial for the applicability of the framework to unlabeled nodes.</p>
<p>The locations of genes in genomes are often specific, and the gene synteny network extracted from the same genus could reflect the intrinsic properties of the genus. The adjacency probability matrix is defined as the probability that two specific nodes can achieve a certain number of connections in a specific genome synteny network. Adding the adjacency probability matrix to the input was found to greatly improve the performance of the framework (<xref ref-type="fig" rid="F5">Figure 5B</xref>). The adjacency probability matrix provides PPA-GCN with an information dissemination pattern for a specific bacterial genus in the gene synteny network.</p>
<p>Since the adjacency probability matrix can be used to extract synteny information patterns for specific microbial species, we wished to verify whether the gene synteny network could be replaced. Two types of random networks were designed while keeping the degree distribution constant. In one case, the arrangement of the gene positions in each sample genome was disrupted, and in the other, the positional relationships of all genomes were disrupted. Three strategies were considered for feature selection: not including any adjacency probability matrix, including the adjacency probability matrix of the newly generated network, and including the adjacency probability matrix of the real network. The training results show that (<xref ref-type="fig" rid="F5">Figure 5C</xref>), regardless of which random network was used, the training performance when using a random network was much lower than that achieved using the real network. Interestingly, the true probability adjacency matrix can improve the framework training performance, while including the matrix of a random network actually impairs performance. This further shows that the adjacency probability matrix can capture specific information patterns of bacterial genomes. The gene synteny network and the adjacency probability matrix can provide the framework with different information patterns, and neither can replace the other.</p>
</sec>
</sec>
<sec id="s4-4">
<title>Effectiveness of Self-Supervised Learning Inspiration</title>
<p>Currently, the assignment rate for gene metabolism pathways is lower than 50% in the KEGG GENES database. For the tested genera of three prokaryotes, the assignment rate for metabolic pathways is less than 20% of all nodes in the network, which greatly limits the training performance. The inspiration of self-supervised learning was adopted to extend the training set. Nodes with low PRAs in the validation set were temporarily excluded from the training set, and nodes with highly stable assigned labels in the test set were temporarily added to the training set. After several iterations, the performance eventually stabilized and showed a great improvement over the initial performance (<xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Self-supervised learning iteration results for the three genera. In each iteration, training was performed for 50 epochs, nodes with an average PRA of less than 30% were removed from the training set, and nodes with stably assigned labels in the test set (with assignment consistency over more than 90% of epochs) were added to the training set. When the proportion of the total number of nodes included in the training set exceeded 90%, the final results were output.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g006.tif"/>
</fig>
<p>We speculate that PPA-GCN&#x2019;s performance could be significantly improved because labeled nodes spread node attributes in a certain pattern, ultimately causing the entire gene synteny network to present a genus-specific information pattern. The question of whether this kind of propagation can be universally applied to different types of gene synteny networks or is suitable only for network structures with a more &#x201c;uniform&#x201d; topology should be considered. Labeled and unlabeled nodes were extracted to construct training and test networks, respectively, and the topological structures of the two new networks were compared. Because PPA-GCN iteratively extracts information from the first- and second-order neighbors of nodes, the tightness of the first- and second-order connections in the network, as measured in terms of the degree distribution and clustering coefficient, need to be considered. The results (<xref ref-type="fig" rid="F7">Figure 7</xref>, <xref ref-type="sec" rid="s11">Supplementary Table S4</xref>) show that the degree distributions of the initial training set and the test set for the three genera are different, reflecting the genome characteristics of each genus to a certain extent. The degree distribution curves and clustering coefficients for the closed pan-genome (<italic>Staphylococcus</italic>) are not significantly different between the initial training set and the test set; in contrast, the initial training set networks of the open pan-genomes (<italic>Flavobacterium</italic> and <italic>Pseudomonas</italic>) are more closely connected than the test set networks, and the overall networks exhibit some level of inhomogeneity. These findings show that the self-supervised inspiration can effectively adapt to gene synteny networks with different topologies.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Degree distribution curves describing the degree distributions of the overall network, the training set network and the test set network for each of the three genera [<bold>(A)</bold> <italic>Flavobacterium,</italic> <bold>(B)</bold> <italic>Pseudomonas</italic> <bold>(C)</bold> <italic>Staphylococcus</italic>]. The horizontal axis is the degree (truncated to 100), and the vertical axis is the probability distribution (The sum of the probabilities is 1). The results show that the degree distributions of the initial training set and the test set for the three genera are different, reflecting the genome characteristics of each genus to a certain extent.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g007.tif"/>
</fig>
</sec>
<sec id="s4-5">
<title>The Impact of Different Types of Genomes on Training</title>
<p>Synteny has been used to filter, organize and process local similarities between genome sequences of related organisms to build a coherent global chromosomal context (<xref ref-type="bibr" rid="B9">Deb et al., 2020</xref>). Each genus of prokaryotes possesses characteristic genomic gene synteny information, and its patterns are broadly associated with many bacterial functional traits (<xref ref-type="bibr" rid="B6">Brbi&#x107; et al., 2016</xref>). Integrating gene synteny data from one genus can provide assistance to the functional pathway assignments of all genes.</p>
<p>Whether different types of genomes would affect training results should be considered. In addition to the node scale, the run number of self-supervised iterations needed to reach convergence can also reflect differences between different types of genomes. <italic>Staphylococcus</italic> requires more iterations to reach a steady state than <italic>Flavobacterium</italic> or <italic>Pseudomonas</italic>. This suggests that the information pattern of a closed pan-genome is relatively conservative and cannot be easily extended, while the information pattern of an open pan-genome is easier to spread. PPA-GCN could provide insights for judging genome types in accordance with the number of iterations needed for self-supervised learning when analyzing the genome of an unknown species.</p>
</sec>
<sec id="s4-6">
<title>The Role of Hyperlink Nodes in the Gene Synteny Network</title>
<p>There are several nodes with a &#x201c;super connection number&#x201d; in the gene synteny network of each genus. Further analysis revealed that these hyperlinked nodes have certain similarities in function. A large proportion of such nodes is assigned to mobile genetic elements (MGEs), which have the potential to disrupt the synteny of the involved genomes and are considered to cause gradual changes (sometimes mutations) in biological genes and promote biological evolution (<xref ref-type="bibr" rid="B27">Muszewska et al., 2019</xref>; <xref ref-type="bibr" rid="B30">Richards et al., 2019</xref>).</p>
<p>We investigated whether the insertion of MGEs into the genomes is random and has an impact on the pattern of functional labels. Two sets of experiments were designed. In the first set, all MGE nodes were removed from the gene synteny network to verify whether the insertion of the MGEs disrupted the information pattern of the original gene synteny networks. In the second set, all MGE nodes were added to the training set as negative samples to verify whether the intervention of the MGEs affected the distribution of functional labels. The results show (<xref ref-type="fig" rid="F8">Figure 8</xref>) that when the MGE nodes are removed from the networks, the performance of PPA-GCN is significantly reduced. When they are used as negative samples, the performance of the framework is only slightly reduced. This indicates that from the perspective of gene location, MGEs may constitute an important part of the gene synteny network of a specific genus, and removing them will destroy the information pattern of the existing gene synteny network. Moreover, MGEs do not interfere with the distribution pattern of gene function.</p>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>The impact of MGEs on PPA-GCN performance. The horizontal axis represents standard training for the three genera (left), training with the MGEs as negative samples (middle) and training with the MGEs removed from the gene synteny network (right). The vertical axis uses PRA as an evaluation index. The results show that when the MGE nodes are removed from the networks, the performance of PPA-GCN is significantly reduced. When they are used as negative samples, the performance of the framework is only slightly reduced.</p>
</caption>
<graphic xlink:href="fgene-13-839453-g008.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>Discussion</title>
<p>In present, PPA-GCN is the first deep learning framework that uses genomic structure information to directly assist metabolic pathway assignments of prokaryotic genomes against KEGG information. Datasets representing three genera (<italic>Flavobacterium</italic>, <italic>Pseudomonas and Staphylococcus</italic>) were used to evaluate the assignment rate of the framework, and on all of them, good performance and strong fault tolerance were achieved. These results support the broad application of PPA-GCN to prokaryotic genomic research. For example, it can provide support for the mechanism research of pathogenic bacteria and the design of synthetic biology elements, modules and pathways.</p>
<p>Although all bacterial genome had been fragmented and shuffled by the endless genomic reconstruction and horizontal gene transfer, the localized genome structure was conserved within specific genus of bacteria. Gene synteny structure is intrinsic and stable under genus level and PPA-GCN relies on it. PPA-GCN captures the graph structure and node attributes from the gene synteny information through a graph convolutional network. To maximize the given pathway information of genomes of a genus, PPA-GCN obtains and mines as many possibilities for label assignment through the network as possible. Then PPA-GCN constructs the adjacency probability matrix to evaluate all possibilities, improving the certainty of all assigned labels. The idea of self-supervised learning is adopted to expand the training set and reinforce the training process.</p>
<p>PPA-GCN has the potential for further improvement. The runtime and memory usage of PPA-GCN will be optimized (<xref ref-type="sec" rid="s11">Supplementary Table S5</xref>). At present, only one kind of graph information (the gene synteny network) is used to make assignments. In the future, some other information networks could be incorporated to improve the performance of PPA-GCN, potentially providing the perfect complement to the existing framework, such as a protein-protein interaction network and gene co-expression network.</p>
<p>PPA-GCN exhibits good performance and shows promise to help guide experimental verification and provide considerable additional space for downstream analysis. PPA-GCN could be applied to more genera of prokaryotes with sufficient whole genome sequences and used to build a database of consensus sequences from the perspective of functional pathway assignment, that could describe the differences in prokaryotes of various genera. In short, we present a deep learning framework with great potential to explain the relationship between gene synteny and KEGG pathway information in prokaryotes, which can provide novel insights into functional pathways assignments and is likely to inform future deep learning applications for interpreting functional annotations.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s11">Supplementary Material</xref>, further inquiries can be directed to the corresponding authors.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>YL developed the framework, TL conceived and supervised the project. YL and QL wrote the manuscript.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This research was supported by the National Key R&#x26;D Program of China (Grant No. 2020YFA0907402), the National Natural Science Foundation of China (Grant No. 91851118 and No.31900096), and the Science and Technology Basic Resources Investigation Program of China (Grant No. 2017FY100300).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors would like to thank the support from Yan Lin.</p>
</ack>
<sec id="s11">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.839453/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2022.839453/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.docx" id="SM1" mimetype="application/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Agrawal</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zitnik</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Leskovec</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). &#x201c;<article-title>Large-scale Analysis of Disease Pathways in the Human Interactome</article-title>,&#x201d; in <conf-name>PACIFIC SYMPOSIUM ON BIOCOMPUTING 2018: Proceedings of the Pacific Symposium</conf-name>, <fpage>111</fpage>&#x2013;<lpage>122</lpage>. <pub-id pub-id-type="doi">10.1142/9789813235533_0011</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Gish</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Myers</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Basic Local Alignment Search Tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume> (<issue>3</issue>), <fpage>403</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/s0022-2836(05)80360-2</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Aramaki</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Blanc-Mathieu</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Endo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ohkubo</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kanehisa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Goto</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>KofamKOALA: KEGG Ortholog Assignment Based on Profile HMM and Adaptive Score Threshold</article-title>. <source>Bioinformatics</source> <volume>36</volume> (<issue>7</issue>), <fpage>2251</fpage>&#x2013;<lpage>2252</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz859</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bazgir</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Dhruba</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Rahman</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ghosh</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Pal</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Representation of Features as Images with Neighborhood Dependencies for Compatibility with Convolutional Neural Networks</article-title>. <source>Nat. Commun.</source> <volume>11</volume> (<issue>1</issue>), <fpage>4391</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-020-18197-y</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Benson</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Cavanaugh</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Karsch-Mizrachi</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Ostell</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>GenBank</article-title>. <source>Genbank. Nucleic Acids Research</source> <volume>41</volume> (<issue>D1</issue>), <fpage>D36</fpage>&#x2013;<lpage>D42</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gks1195</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brbi&#x107;</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Pi&#x161;korec</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Vidulin</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Kri&#x161;ko</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>&#x160;muc</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Supek</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>The Landscape of Microbial Phenotypic Traits and Associated Genes</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>10074</fpage>&#x2013;<lpage>10090</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw964</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>I.-M. A.</given-names>
</name>
<name>
<surname>Chu</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Palaniappan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Pillay</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ratner</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>IMG/M v.5.0: an Integrated Data Management and Comparative Analysis System for Microbial Genomes and Microbiomes</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume> (<issue>D1</issue>), <fpage>D666</fpage>&#x2013;<lpage>D677</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky901</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cho</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Berger</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Peng</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Compact Integration of Multi-Network Topology for Functional Analysis of Genes</article-title>. <source>Cel Syst.</source> <volume>3</volume> (<issue>6</issue>), <fpage>540</fpage>&#x2013;<lpage>548</lpage>. <pub-id pub-id-type="doi">10.1016/j.cels.2016.10.017</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deb</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Jayaprasad</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ravi</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Rao</surname>
<given-names>K. R.</given-names>
</name>
<name>
<surname>Whadgar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hariharan</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Classification of Grain Amaranths Using Chromosome-Level Genome Assembly of Ramdana, A. Hypochondriacus</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>, <fpage>579529</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2020.579529</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Doerks</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Von Mering</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bork</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Functional Clues for Hypothetical Proteins Based on Genomic Context Analysis in Prokaryotes</article-title>. <source>Nucleic Acids Res.</source> <volume>32</volume> (<issue>21</issue>), <fpage>6321</fpage>&#x2013;<lpage>6326</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkh973</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Emms</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Kelly</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>OrthoFinder: Phylogenetic Orthology Inference for Comparative Genomics</article-title>. <source>Genome Biol.</source> <volume>20</volume> (<issue>1</issue>), <fpage>238</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1832-y</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eslami Manoochehri</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Nourani</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Drug-target Interaction Prediction Using Semi-bipartite Graph Model and Deep Learning</article-title>. <source>BMC bioinformatics</source> <volume>21</volume> (<issue>4</issue>), <fpage>248</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3518-6</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fan</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Pseudo2GO: a Graph-Based Deep Learning Method for Pseudogene Function Prediction by Borrowing Information from Coding Genes</article-title>. <source>Front. Genet.</source> <volume>11</volume>, <fpage>807</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2020.00807</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Federhen</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The NCBI Taxonomy Database</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume> (<issue>D1</issue>), <fpage>D136</fpage>&#x2013;<lpage>D143</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr1178</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gligorijevi&#x107;</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Barot</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bonneau</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>deepNF: Deep Network Fusion for Protein Function Prediction</article-title>. <source>Bioinformatics</source> <volume>34</volume> (<issue>22</issue>), <fpage>3873</fpage>&#x2013;<lpage>3881</lpage>. </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grover</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Leskovec</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>node2vec: Scalable Feature Learning for Networks</article-title>,&#x201d; in <conf-name>Proceedings of the 22nd ACM SIGKDD international conference on Knowledge discovery and data mining</conf-name>, <fpage>855</fpage>&#x2013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1145/2939672.2939754</pub-id>
<source>KDD</source>
<volume>2016</volume> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hao</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Boss&#xe9;</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Nickle</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Par&#xe9;</surname>
<given-names>P. D.</given-names>
</name>
<name>
<surname>Postma</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Laviolette</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Lung eQTLs to Help Reveal the Molecular Underpinnings of Asthma</article-title>. <source>Plos Genet.</source> <volume>8</volume> (<issue>11</issue>), <fpage>e1003029</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1003029</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Lakowski</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Deep Graph Embedding for Prioritizing Synergistic Anticancer Drug Combinations</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>18</volume>, <fpage>427</fpage>&#x2013;<lpage>438</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2020.02.006</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jun</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Nookaew</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Hauser</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gorin</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Assessment of Genome Annotation Using Gene Function Similarity within the Gene Neighborhood</article-title>. <source>BMC bioinformatics</source> <volume>18</volume> (<issue>1</issue>), <fpage>345</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1761-2</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kanehisa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sato</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kawashima</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Furumichi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tanabe</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016b</year>). <article-title>KEGG as a Reference Resource for Gene and Protein Annotation</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume> (<issue>D1</issue>), <fpage>D457</fpage>&#x2013;<lpage>D462</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1070</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kanehisa</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sato</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Morishima</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2016a</year>). <article-title>BlastKOALA and GhostKOALA: KEGG Tools for Functional Characterization of Genome and Metagenome Sequences</article-title>. <source>J. Mol. Biol.</source> <volume>428</volume> (<issue>4</issue>), <fpage>726</fpage>&#x2013;<lpage>731</lpage>. <pub-id pub-id-type="doi">10.1016/j.jmb.2015.11.006</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Keegan</surname>
<given-names>K. P.</given-names>
</name>
<name>
<surname>Glass</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Meyer</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2016</year>). &#x201c;<article-title>MG-RAST, a Metagenomics Service for Analysis of Microbial Community Structure and Function</article-title>,&#x201d; in <source>Microbial Environmental Genomics (MEG)</source> (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Humana Press</publisher-name>), <fpage>207</fpage>&#x2013;<lpage>233</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-3369-3_13</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kipf</surname>
<given-names>T. N.</given-names>
</name>
<name>
<surname>Welling</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Semi-supervised Classification with Graph Convolutional Networks</article-title>. <source>arXiv preprint arXiv:1609.02907</source>. </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kuwahara</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Gao</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>PGCN: Disease Gene Prioritization by Disease and Gene Embedding through Graph Convolutional Neural Networks</article-title>. <source>bioRxiv</source> <volume>2019</volume>, <fpage>532226</fpage>. </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Mamitsuka</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>HPOFiller: Identifying Missing Protein-Phenotype Associations by Graph Convolutional Network</article-title>. <source>Bioinformatics</source> <volume>2021</volume>, <fpage>btab224</fpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab224</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lukassen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ten</surname>
<given-names>F. W.</given-names>
</name>
<name>
<surname>Adam</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Eils</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Conrad</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Gene Set Inference from Single-Cell Sequencing Data Using a Hybrid of Matrix Factorization and Variational Autoencoders</article-title>. <source>Nat. Mach Intell.</source> <volume>2</volume> (<issue>12</issue>), <fpage>800</fpage>&#x2013;<lpage>809</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-00269-9</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Muszewska</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Steczkiewicz</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Stepniewska-Dziubinska</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ginalski</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Transposable Elements Contribute to Fungal Genes and Impact Fungal Lifestyle</article-title>. <source>Sci. Rep.</source> <volume>9</volume> (<issue>1</issue>), <fpage>4307</fpage>&#x2013;<lpage>4310</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-019-40965-0</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nadeau</surname>
<given-names>J. H.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>B. A.</given-names>
</name>
</person-group> (<year>1984</year>). <article-title>Lengths of Chromosomal Segments Conserved since Divergence of Man and Mouse</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>81</volume> (<issue>3</issue>), <fpage>814</fpage>&#x2013;<lpage>818</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.81.3.814</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Repecka</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Jauniskis</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Karpus</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Rembeza</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Rokaitis</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Zrimec</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Expanding Functional Protein Sequence Spaces Using Generative Adversarial Networks</article-title>. <source>Nat. Mach Intell.</source> <volume>3</volume> (<issue>4</issue>), <fpage>324</fpage>&#x2013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-021-00310-5</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Richards</surname>
<given-names>V. P.</given-names>
</name>
<name>
<surname>Velsko</surname>
<given-names>I. M.</given-names>
</name>
<name>
<surname>Alam</surname>
<given-names>M. T.</given-names>
</name>
<name>
<surname>Zadoks</surname>
<given-names>R. N.</given-names>
</name>
<name>
<surname>Manning</surname>
<given-names>S. D.</given-names>
</name>
<name>
<surname>Pavinski Bitar</surname>
<given-names>P. D.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Population Gene Introgression and High Genome Plasticity for the Zoonotic Pathogen Streptococcus Agalactiae</article-title>. <source>Mol. Biol. Evol.</source> <volume>36</volume> (<issue>11</issue>), <fpage>2572</fpage>&#x2013;<lpage>2590</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/msz169</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rogozin</surname>
<given-names>I. B.</given-names>
</name>
<name>
<surname>Makarova</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Wolf</surname>
<given-names>Y. I.</given-names>
</name>
<name>
<surname>Koonin</surname>
<given-names>E. V.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Computational Approaches for the Analysis of Gene Neighbourhoods in Prokaryotic Genomes</article-title>. <source>Brief. Bioinformatics</source> <volume>5</volume> (<issue>2</issue>), <fpage>131</fpage>&#x2013;<lpage>149</lpage>. <pub-id pub-id-type="doi">10.1093/bib/5.2.131</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Saha</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chatterjee</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Basu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kundu</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nasipuri</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Improving Prediction of Protein Function from Protein Interaction Network Using Intelligent Neighborhood Approach</article-title>,&#x201d; in <conf-name>2012 International Conference on Communications, Devices and Intelligent Systems (CODIS)</conf-name> (<publisher-loc>Kolkata, India</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>584</fpage>&#x2013;<lpage>587</lpage>. <pub-id pub-id-type="doi">10.1109/codis.2012.6422270</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Saha</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Prasad</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chatterjee</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Basu</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Nasipuri</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Protein Function Prediction from Protein-Protein Interaction Network Using Gene Ontology Based Neighborhood Analysis and Physico-Chemical Features</article-title>. <source>J. Bioinform. Comput. Biol.</source> <volume>16</volume> (<issue>06</issue>), <fpage>1850025</fpage>. <pub-id pub-id-type="doi">10.1142/s0219720018500257</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schulte-Sasse</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Budach</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hnisz</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Marsico</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Integration of Multiomics Data with Graph Convolutional Networks to Identify New Cancer Genes and Their Associated Molecular Mechanisms</article-title>. <source>Nat. Mach Intell.</source> <volume>3</volume> (<issue>6</issue>), <fpage>513</fpage>&#x2013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-021-00325-y</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jiang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Bacteriophage Classification for Assembled Contigs Using Graph Convolutional Network</article-title>. <source>arXiv preprint arXiv:2102.03746</source>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab293</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shmakov</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Faure</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Makarova</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Wolf</surname>
<given-names>Y. I.</given-names>
</name>
<name>
<surname>Severinov</surname>
<given-names>K. V.</given-names>
</name>
<name>
<surname>Koonin</surname>
<given-names>E. V.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Systematic Prediction of Functionally Linked Genes in Bacterial and Archaeal Genomes</article-title>. <source>Nat. Protoc.</source> <volume>14</volume> (<issue>10</issue>), <fpage>3013</fpage>&#x2013;<lpage>3031</lpage>. <pub-id pub-id-type="doi">10.1038/s41596-019-0211-1</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzuki</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kakuta</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ishida</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Akiyama</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>GHOSTX: an Improved Sequence Homology Search Algorithm Using a Query Suffix Array and a Database Suffix Array</article-title>. <source>PloS one</source> <volume>9</volume> (<issue>8</issue>), <fpage>e103833</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0103833</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tettelin</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Masignani</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Cieslewicz</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Donati</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Medini</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ward</surname>
<given-names>N. L.</given-names>
</name>
<etal/>
</person-group> (<year>2005</year>). <article-title>Genome Analysis of Multiple Pathogenic Isolates of Streptococcus Agalactiae : Implications for the Microbial "Pan-Genome"</article-title>. <source>Proc. Natl. Acad. Sci. U.S.A.</source> <volume>102</volume> (<issue>39</issue>), <fpage>13950</fpage>&#x2013;<lpage>13955</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0506758102</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Nie</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Xin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>An Interpretable Deep-Learning Architecture of Capsule Networks for Identifying Cell-type Gene Expression Programs from Single-Cell RNA-Sequencing Data</article-title>. <source>Nat. Mach Intell.</source> <volume>2</volume> (<issue>11</issue>), <fpage>693</fpage>&#x2013;<lpage>703</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-00244-4</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang M.</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kong</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ma</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Pan-Genome Analyses of Geobacillus Spp. Reveal Genetic Characteristics and Composting Potential</article-title>. <source>Ijms</source> <volume>21</volume> (<issue>9</issue>), <fpage>3393</fpage>. <pub-id pub-id-type="doi">10.3390/ijms21093393</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>P. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>A Comprehensive Survey on Graph Neural Networks</article-title>. <source>IEEE Trans. Neural Netw. Learn. Syst.</source> <volume>32</volume> (<issue>1</issue>), <fpage>4</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.1109/TNNLS.2020.2978386</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Yue</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zou</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Characterization and Functional Analysis of Calmodulin and Calmodulin-like Genes in Fragaria Vesca</article-title>. <source>Front. Plant Sci.</source> <volume>7</volume>, <fpage>1820</fpage>. <pub-id pub-id-type="doi">10.3389/fpls.2016.01820</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Multi-View Graph Convolutional Network and its Applications on Neuroimage Analysis for Parkinson&#x27;s Disease</article-title>. <source>AMIA Annu. Symp. Proc.</source> <volume>2018</volume>, <fpage>1147</fpage>&#x2013;<lpage>1156</lpage>. </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lamkin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Gymrek</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Deep Neural Networks Identify Sequence Context Features Predictive of Transcription Factor Binding</article-title>. <source>Nat. Mach Intell.</source> <volume>3</volume> (<issue>2</issue>), <fpage>172</fpage>&#x2013;<lpage>180</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-00282-y</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zitnik</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Agrawal</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Leskovec</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Modeling Polypharmacy Side Effects with Graph Convolutional Networks</article-title>. <source>Bioinformatics</source> <volume>34</volume> (<issue>13</issue>), <fpage>i457</fpage>&#x2013;<lpage>i466</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty294</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zitnik</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Leskovec</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Predicting Multicellular Function through Multi-Layer Tissue Networks</article-title>. <source>Bioinformatics</source> <volume>33</volume> (<issue>14</issue>), <fpage>i190</fpage>&#x2013;<lpage>i198</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx252</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>