<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Cell Dev. Biol.</journal-id>
<journal-title>Frontiers in Cell and Developmental Biology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Cell Dev. Biol.</abbrev-journal-title>
<issn pub-type="epub">2296-634X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fcell.2021.739715</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Cell and Developmental Biology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Prediction of Gastric Cancer-Related Proteins Based on Graph Fusion Method</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Zhang</surname> <given-names>Hao</given-names></name>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Ruisi</given-names></name>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1403715/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Ding</surname> <given-names>Meng</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zhang</surname> <given-names>Ying</given-names></name>
<xref ref-type="corresp" rid="c002"><sup>&#x002A;</sup></xref>
</contrib>
</contrib-group>
<aff><institution>Endoscopy Center, China-Japan Union Hospital of Jilin University</institution>, <addr-line>Changchun</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Lei Deng, Central South University, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Ningyi Zhang, Harbin Institute of Technology, China; Hong Ju, Heilongjiang Vocational College of Biology Science and Technology, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Meng Ding, <email>309478506@qq.com</email></corresp>
<corresp id="c002">Ying Zhang, <email>515070789@qq.com</email></corresp>
<fn fn-type="equal" id="fn002"><p><sup>&#x2020;</sup>These authors have contributed equally to this work</p></fn>
<fn fn-type="other" id="fn004"><p>This article was submitted to Molecular and Cellular Pathology, a section of the journal Frontiers in Cell and Developmental Biology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>01</day>
<month>11</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>9</volume>
<elocation-id>739715</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>02</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2021 Zhang, Xu, Ding and Zhang.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Zhang, Xu, Ding and Zhang</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Gastric cancer is a common malignant tumor of the digestive system with no specific symptoms. Due to the limited knowledge of pathogenesis, patients are usually diagnosed in advanced stage and do not have effective treatment methods. Proteome has unique tissue and time specificity and can reflect the influence of external factors that has become a potential biomarker for early diagnosis. Therefore, discovering gastric cancer-related proteins could greatly help researchers design drugs and develop an early diagnosis kit. However, identifying gastric cancer-related proteins by biological experiments is time- and money-consuming. With the high speed increase of data, it has become a hot issue to mine the knowledge of proteomics data on a large scale through computational methods. Based on the hypothesis that the stronger the association between the two proteins, the more likely they are to be associated with the same disease, in this paper, we constructed both disease similarity network and protein interaction network. Then, Graph Convolutional Networks (GCN) was applied to extract topological features of these networks. Finally, Xgboost was used to identify the relationship between proteins and gastric cancer. Results of 10-cross validation experiments show high area under the curve (AUC) (0.85) and area under the precision recall (AUPR) curve (0.76) of our method, which proves the effectiveness of our method.</p>
</abstract>
<kwd-group>
<kwd>gastric cancer</kwd>
<kwd>protein</kwd>
<kwd>proteomics data</kwd>
<kwd>graph convolutional network</kwd>
<kwd>Xgboost</kwd>
</kwd-group>
<counts>
<fig-count count="2"/>
<table-count count="1"/>
<equation-count count="10"/>
<ref-count count="30"/>
<page-count count="6"/>
<word-count count="3962"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="S1">
<title>Introduction</title>
<p>Gastric cancer is a worldwide disease with high incidence rate and mortality rate, especially in East Asia (<xref ref-type="bibr" rid="B24">Villanueva, 2011</xref>). According to the data from GLOBOCAN in 2018, there were 1,033,701 new cases of and 782,685 deaths from gastric cancer in the world (<xref ref-type="bibr" rid="B2">Bray et al., 2018</xref>). At present, the early diagnosis of gastric cancer is limited; patients are usually diagnosed in advanced stage. Therefore, early diagnosis is the key to improve the prognosis of patients, which is also the goal pursued by many researchers (<xref ref-type="bibr" rid="B10">Jin et al., 2015</xref>). Biomarkers refer to the substances that can reflect the physiological, biochemical, immune, genetic, and other molecular changes in the organism (<xref ref-type="bibr" rid="B20">Sz&#x00E1;sz et al., 2016</xref>; <xref ref-type="bibr" rid="B11">Liang et al., 2019</xref>; <xref ref-type="bibr" rid="B4">Cheng et al., 2021</xref>). The levels of biomarkers in patients&#x2019; samples (such as blood, plasma, saliva, and urine) can reflect the health or disease status of patients, as well as the response to anticancer treatment. Due to the strong heterogeneity of gastric cancer, the use of proteomics technology to find new specific biomarkers will greatly improve the sensitivity and accuracy of the diagnosis of patients (<xref ref-type="bibr" rid="B7">Gullo et al., 2018</xref>).</p>
<p>Although many researchers tend to reveal diseases pathogenic mechanism by genomics (<xref ref-type="bibr" rid="B16">Peng and Zhao, 2020</xref>; <xref ref-type="bibr" rid="B26">Zhao et al., 2020b</xref>; <xref ref-type="bibr" rid="B30">Zhou et al., 2020</xref>), changes in protein quality in diseases reflect the progression of the disease and are also the product of genes (<xref ref-type="bibr" rid="B27">Zhao et al., 2020c</xref>). Unlike those studies that research diseases through gene expression (<xref ref-type="bibr" rid="B29">Zhao et al., 2021b</xref>), protein quantification is more accurate and has the potential to become a biomarker. Researchers have used various protein separation techniques, such as two-dimensional gel electrophoresis (2-DE) (<xref ref-type="bibr" rid="B8">Gygi et al., 2000</xref>), two-dimensional fluorescence difference gel electrophoresis (2D-DIGE) (<xref ref-type="bibr" rid="B21">Tannu and Hemby, 2006</xref>), isobaric tags for relative and absolute quantitation (iTRAQ), hydrophilic interaction liquid chromatography (HILIC) screening of potential target proteins of new gastric cancer biomarkers, and then Western blotting and enzyme-linked immunosorbent assay or immunohistochemistry (IHC) methods are further validated, and biomarkers that play a key role in the occurrence of malignant tumors can be discovered.</p>
<p><xref ref-type="bibr" rid="B17">Ryu et al. (2003)</xref> used the tumor proteomics technology of antibody microarrays to identify inflammatory protein markers of gastric cancer. They found that 14 proteins have different expressions between normal gastric mucosa and tumor gastric mucosa. The proteome can be regarded as the functional cell equivalent of the genome. Proteomics is useful in discovering biomarkers and improving the diagnostic efficiency of early gastric cancer and has obvious advantages. At present, the prognosis and treatment methods of gastric cancer are guided by genome. Surgical resection is still the most common strategy of gastric cancer, but due to the high risk of disease progression in stage II or III patients, it becomes important to increase adjuvant therapy. The strong heterogeneity of gastric cancer makes the therapeutic effect heterogeneous. Therefore, although the TNM system can help the prognosis of gastric cancer, many researchers tend to discover biomarkers to predict treatment outcomes more accurately (<xref ref-type="bibr" rid="B14">Pang et al., 2018</xref>). For example, <xref ref-type="bibr" rid="B1">Balluff et al. (2011)</xref> used matrix-assisted laser desorption/ionization (MALDI) imaging technology to analyze tissue samples and found that cysteine-rich intestinal protein 1 (CRIP1) and human neutrophil peptide-1 (HNP-1) were prognostic factors for gastric cancer. Human epidermal growth factor receptor 2 (HER2) is an important biomarker in gastric tumors, which can be specifically targeted for treatment with trastuzumab monoclonal antibody (mAb). For patients with advanced gastric cancer or gastroesophageal junction cancer, trastuzumab combined with chemotherapy can improve the survival rate of patients (<xref ref-type="bibr" rid="B15">Park et al., 2018</xref>).</p>
<p>There are still few proteins known to be related to gastric cancer. With the explosive growth of various types of omics data (<xref ref-type="bibr" rid="B13">Mo et al., 2020</xref>; <xref ref-type="bibr" rid="B25">Zhao et al., 2020a</xref>,<xref ref-type="bibr" rid="B28">2021a</xref>), computational methods are widely used to identify disease-related biomolecules. Mining disease-related molecules based on the protein interaction networks has become a universal method. <xref ref-type="bibr" rid="B18">Sang et al. (2011)</xref> discovered genes and pathways of ciliopathy disease based on protein network. <xref ref-type="bibr" rid="B19">Seyfried et al. (2017)</xref> constructed protein network to identify protein-specific co-expression in Alzheimer&#x2019;s disease. With the development of Graph Convolutional Networks (GCN), an increasing number of researchers tend to use this method to process the complex topological features of the biological network. Its core point of view is to make the entire graph converge through the dissemination of node information and then make predictions on the basis of it. It has been widely used in prediction of biomolecular interaction (<xref ref-type="bibr" rid="B22">Tianyi et al., 2020</xref>). Therefore, we proposed a GCN-based method in this paper, named &#x201C;GXGCP&#x201D; (Gcn-Xgboost for Gastric Cancer-related Proteins identification) to identify gastric cancer-related proteins.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<p>There are four steps to implement GXGCP. Step 1 is to construct disease similarity network and protein interaction network. Step 2 is using GCN to extract topological features of disease similarity network and protein interaction network, respectively. Step 3 is to reduce the dimension of protein and gastric cancer features by principal component analysis (PCA). Step 4 is to identify gastric cancer-related proteins based on the features of protein and gastric cancer by Xgboost. The workflow of GXGCP is shown in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>Workflow of GXGCP (Gcn-Xgboost for Gastric Cancer-related Proteins identification).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcell-09-739715-g001.tif"/>
</fig>
<sec id="S2.SS1">
<title>Construction of Network</title>
<p>We used SemFunsim (<xref ref-type="bibr" rid="B5">Cheng et al., 2014</xref>) to obtain diseases that are similar to gastric cancer. This method considers both disease semantic association and gene association. The detailed calculation process will not be repeated in this paper. A total of 327 diseases were found to be similar to gastric cancer. Based on the similarity, we constructed disease network, in which the edges are similarity and nodes are diseases. Therefore, the network has weight.</p>
<p>We downloaded protein interaction information from Search Tool for the Retrieval of Interacting Genes/Proteins (STRING) (<xref ref-type="bibr" rid="B12">Mering et al., 2003</xref>). Based on the interaction, we constructed a protein network. If a protein can interact with the other one, there would be an edge to connect each other. Since the intensity of interaction between different proteins is different, this network also has weight.</p>
</sec>
<sec id="S2.SS2">
<title>Extracting Topological Features by Graph Convolutional Networks</title>
<p>To fully extract topological features of protein and disease network, GCN was applied (<xref ref-type="bibr" rid="B9">Han et al., 2019</xref>). The aim to implement GCN is to convert network topology into a vector output:</p>
<disp-formula id="S2.E1">
<label>(1)</label>
<mml:math id="M1">
<mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>H</italic><sup>(0)</sup> is node&#x2019;s feature in the network.</p>
<p>First, Laplace transform should be done on the network:</p>
<disp-formula id="S2.E2">
<label>(2)</label>
<mml:math id="M2">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>-</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where D is the degree matrix of the network, and A is the adjacency matrix.</p>
<disp-formula id="S2.E3">
<label>(3)</label>
<mml:math id="M3">
<mml:mrow>
<mml:msub>
<mml:mover accent="true">
<mml:mtext>D</mml:mtext>
<mml:mo stretchy="false">^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mtext>ii</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>j</mml:mi>
</mml:munder>
<mml:msub>
<mml:mover accent="true">
<mml:mtext>A</mml:mtext>
<mml:mo stretchy="false">^</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Then, normalization should be implemented on the Laplacian matrix:</p>
<disp-formula id="S2.E4">
<label>(4)</label>
<mml:math id="M4">
<mml:mrow>
<mml:msup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mi>I</mml:mi>
<mml:mo>-</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p><italic>L</italic><sup><italic>s</italic><italic>y</italic><italic>m</italic></sup> is defined as:</p>
<disp-formula id="S2.E5">
<label>(5)</label>
<mml:math id="M5">
<mml:mrow>
<mml:msubsup>
<mml:mi>L</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable displaystyle="true" rowspacing="0pt">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo mathvariant="italic" separator="true">&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2002;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mpadded width="+2.8pt">
<mml:mi>j</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+2.8pt">
<mml:mi>d</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mi>deg</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mi>deg</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mi>deg</mml:mi>
<mml:mo>&#x2061;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mfrac>
</mml:mrow>
<mml:mo mathvariant="italic" separator="true">&#x2003;&#x2003;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mo>&#x2260;</mml:mo>
<mml:mrow>
<mml:mpadded width="+2.8pt">
<mml:mi>j</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+2.8pt">
<mml:mi>d</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+2.8pt">
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+2.8pt">
<mml:mi>t</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+2.8pt">
<mml:mi>o</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>v</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo mathvariant="italic" separator="true">&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2003;&#x2002;</mml:mo>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>t</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mi/>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>With the Laplacian matrix, we can perform spectral convolution on the network. We need to find a suitable convolution kernel so that <italic>f</italic>() can reduce the loss of classification after the convolution transformation of the convolution kernel. The core of the machine learning task on the graph is to find a convolution kernel that can reduce the loss, regard <italic>h</italic>(&#x03BB;<sub>1</sub>),&#x2026;<italic>h</italic>(&#x03BB;<sub><italic>n</italic></sub>) as the parameters of the model, and apply the gradient descent method to update these parameters.</p>
<p>The final formula of GCN would be:</p>
<disp-formula id="S2.E6">
<label>(6)</label>
<mml:math id="M6">
<mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mi>D</mml:mi>
<mml:mrow>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03C3;() is the activation function, and <italic>W</italic><sup>(<italic>l</italic>)</sup> is the parameter to be trained.</p>
</sec>
<sec id="S2.SS3">
<title>Reduction Dimension by Principal Component Analysis</title>
<p>Since the dimension of metabolites and gastric cancer features are large, we used PCA to reduce the dimension. There are four steps to apply PCA (<xref ref-type="bibr" rid="B23">Tipping and Bishop, 1999</xref>; <xref ref-type="bibr" rid="B6">Cheng et al., 2019</xref>). The first step is feature centralization. That is, the data of each dimension are subtracted from the mean value of that dimension, and the mean value of each dimension becomes 0 after the transformation. The second step is to calculate covariance matrix. The third step is to calculate the eigenvalues and eigenvectors of the covariance matrix. The last step is to select the feature vector corresponding to the large feature value to obtain a new data set.</p>
</sec>
<sec id="S2.SS4">
<title>Classification of Gastric Cancer-Related Proteins by Xgboost</title>
<p>Xgboost is a sparse perception algorithm that can be used for parallel tree learning (<xref ref-type="bibr" rid="B3">Chen and Guestrin, 2016</xref>). Since the features of gastric cancer and proteins are sparse, Xgboost is very suitable for the classification.</p>
<p>Xgboost is a tree ensemble model. It sums the results of K (the number of trees) as the final predicted value.</p>
<disp-formula id="S2.E7">
<label>(7)</label>
<mml:math id="M7">
<mml:mrow>
<mml:msub>
<mml:mover>
<mml:mi>y</mml:mi>
<mml:mo>&#x2322;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>&#x03D5;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:munderover>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>K</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo rspace="13.6pt">,</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Assuming that a given sample set has n samples and m features, then</p>
<disp-formula id="S2.E8">
<label>(8)</label>
<mml:math id="M8">
<mml:mrow>
<mml:mi>D</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>x</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>x</italic><sub><italic>i</italic></sub> represents the i-th sample, <italic>y</italic><sub><italic>i</italic></sub> represents the i-th category label, and the space F of the regression tree (CART tree) is:</p>
<disp-formula id="S2.E9">
<label>(9)</label>
<mml:math id="M9">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mi>f</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:mi>q</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where q represents the structure of each tree, it maps the sample to the corresponding leaf node; T is the number of leaf nodes of the corresponding tree; f(x) corresponds to the structure q of the tree and the leaf node weight w. Therefore, the predicted value of Xgboost is the sum of the values of the leaf nodes corresponding to each tree.</p>
<p>Our goal is to learn these k trees, so we minimize the following objective function with regular terms:</p>
<disp-formula id="S2.E10">
<label>(10)</label>
<mml:math id="M10">
<mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>&#x03D5;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mover>
<mml:mi>y</mml:mi>
<mml:mo>&#x2322;</mml:mo>
</mml:mover>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>k</mml:mi>
</mml:munder>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03A9;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula><mml:math id="INEQ9"><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x03B3;</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>T</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x2062;</mml:mo><mml:mi mathvariant="normal">&#x03BB;</mml:mi><mml:mo>&#x2062;</mml:mo><mml:msup><mml:mrow><mml:mo fence="true">||</mml:mo><mml:mi>w</mml:mi><mml:mo fence="true">||</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula></p>
</sec>
</sec>
<sec id="S3">
<title>Experiment Results</title>
<p>We implemented 10-cross validation experiments to test the performance of GXGCP. We divided our data into 10 groups. We used nine of 10 groups&#x2019; data to train the model and the data of the remaining one to test the model. After repeating this process 10 times, each group has been tested once. To show the accuracy of our model, we compared GXGCP with several other methods such as RWXGCP, GXGCP without PCA, GSVMCP, GANNCP, and GCNNCP. RWXGCP replaces the GCN part of GXGCP with random walk (RW). GSVMCP replaces the Xgboost part of GXGCP with support vector machine (SVM). GANNCP replaces the Xgboost part of GXGCP with artificial neural network (ANN). GCNNCP replaces the Xgboost part of GXGCP with convolutional neural network (CNN).</p>
<p>The area under the curve (AUC) and area under the precision recall (AUPR) curve of GXGCP are shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. The comparison results are listed in <xref ref-type="table" rid="T1">Table 1</xref>.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Receiver operating characteristic (ROC) curve of 10-cross validation experiments. The AUC and AUPR curve of GXGCP are 0.85 and 0.76, respectively. The comparison results are listed in <xref ref-type="table" rid="T1">Table 1</xref>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fcell-09-739715-g002.tif"/>
</fig>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Comparison results.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center">AUC</td>
<td valign="top" align="center">AUPR</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">GXGCP</td>
<td valign="top" align="center">0.85</td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td valign="top" align="left">RWXGCP</td>
<td valign="top" align="center">0.81</td>
<td valign="top" align="center">0.72</td>
</tr>
<tr>
<td valign="top" align="left">GXGCP without PCA</td>
<td valign="top" align="center">0.72</td>
<td valign="top" align="center">0.68</td>
</tr>
<tr>
<td valign="top" align="left">GSVMCP</td>
<td valign="top" align="center">0.74</td>
<td valign="top" align="center">0.65</td>
</tr>
<tr>
<td valign="top" align="left">GANNCP</td>
<td valign="top" align="center">0.76</td>
<td valign="top" align="center">0.71</td>
</tr>
<tr>
<td valign="top" align="left">GCNNCP</td>
<td valign="top" align="center">0.82</td>
<td valign="top" align="center">0.76</td>
</tr>
<tr>
<td valign="top" align="left">GDNNCP</td>
<td valign="top" align="center">0.80</td>
<td valign="top" align="center">0.74</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="T1">Table 1</xref>, GXGCP performed best among these five methods. These results show that GCN is more suitable for encoding network than RW, and Xgboost is more suitable for building model by sparse data than SVM and ANN.</p>
</sec>
<sec sec-type="conclusions" id="S4">
<title>Conclusion</title>
<p>Protein is the main executor of life activities. To decrypt the genome, you must first systematically understand the proteome. Identifying gastric cancer-related proteins can greatly help develop screening or testing tools for tumor detection, early diagnosis or differential diagnosis, prognostic analysis, efficacy evaluation, etc. Due to the high cost of biological experiments, we proposed GXGCP that fuses GCN, Xgboost, and PCA to identify gastric cancer-related proteins. To verify the accuracy of our method, we did 10-cross validation experiments. The results show that the AUC of GXGCP reached 0.85 and AUPR reached 0.76. To show the superiority of GXGCP, we compared it with several other methods, and GXGCP performed best. Overall, we propose a novel, efficient, and accurate method for large-scale identification of gastric cancer-related proteins, which would greatly benefit the study of the pathogenic mechanism and clinical research of gastric cancer.</p>
</sec>
<sec sec-type="data-availability" id="S5">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found in the article/supplementary material.</p>
</sec>
<sec id="S6">
<title>Ethics Statement</title>
<p>Ethical review and approval was not required for the study on human participants in accordance with the local legislation and institutional requirements. Written informed consent for participation was not required for this study in accordance with the national legislation and the institutional requirements.</p>
</sec>
<sec id="S7">
<title>Author Contributions</title>
<p>HZ and RX wrote this manuscript and did experiments. MD and YZ provided important ideas. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="S8">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Balluff</surname> <given-names>B.</given-names></name> <name><surname>Rauser</surname> <given-names>S.</given-names></name> <name><surname>Meding</surname> <given-names>S.</given-names></name> <name><surname>Elsner</surname> <given-names>M.</given-names></name> <name><surname>Sch&#x00F6;ne</surname> <given-names>C.</given-names></name> <name><surname>Feuchtinger</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>MALDI imaging identifies prognostic seven-protein signature of novel tissue markers in intestinal-type gastric cancer.</article-title> <source><italic>Am. J. Pathol.</italic></source> <volume>179</volume>, <fpage>2720</fpage>&#x2013;<lpage>2729</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajpath.2011.08.032</pub-id> <pub-id pub-id-type="pmid">22015459</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bray</surname> <given-names>F.</given-names></name> <name><surname>Ferlay</surname> <given-names>J.</given-names></name> <name><surname>Soerjomataram</surname> <given-names>I.</given-names></name> <name><surname>Siegel</surname> <given-names>R. L.</given-names></name> <name><surname>Torre</surname> <given-names>L. A.</given-names></name> <name><surname>Jemal</surname> <given-names>A.</given-names></name></person-group> (<year>2018</year>). <article-title>Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries.</article-title> <source><italic>CA Cancer J. Clin.</italic></source> <volume>68</volume> <fpage>394</fpage>&#x2013;<lpage>424</lpage>. <pub-id pub-id-type="doi">10.3322/caac.21492</pub-id> <pub-id pub-id-type="pmid">30207593</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>T.</given-names></name> <name><surname>Guestrin</surname> <given-names>C.</given-names></name></person-group> (<year>2016</year>). &#x201C;<article-title>Xgboost: a scalable tree boosting system</article-title>,&#x201D; in <source><italic>Proceedings of the 22nd ACM Sigkdd International Conference on Knowledge Discovery and Data Mining</italic></source>, (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>785</fpage>&#x2013;<lpage>794</lpage>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>L.</given-names></name> <name><surname>Han</surname> <given-names>X.</given-names></name> <name><surname>Zhu</surname> <given-names>Z.</given-names></name> <name><surname>Qi</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name></person-group> (<year>2021</year>). <article-title>Functional alterations caused by mutations reflect evolutionary trends of SARS-CoV-2.</article-title> <source><italic>Brief. Bioinform.</italic></source> <volume>22</volume> <fpage>1442</fpage>&#x2013;<lpage>1450</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbab042</pub-id> <pub-id pub-id-type="pmid">33580783</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Ju</surname> <given-names>P.</given-names></name> <name><surname>Peng</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2014</year>). <article-title>SemFunSim: a new method for measuring disease similarity by integrating semantic and gene functional association.</article-title> <source><italic>PLoS One</italic></source> <volume>9</volume>:<issue>e99415</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0099415</pub-id> <pub-id pub-id-type="pmid">24932637</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname> <given-names>L.</given-names></name> <name><surname>Zhao</surname> <given-names>H.</given-names></name> <name><surname>Wang</surname> <given-names>P.</given-names></name> <name><surname>Zhou</surname> <given-names>W.</given-names></name> <name><surname>Luo</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2019</year>). <article-title>Computational methods for identifying similar diseases.</article-title> <source><italic>Mol. Ther. Nucleic Acids</italic></source> <volume>18</volume> <fpage>590</fpage>&#x2013;<lpage>604</lpage>. <pub-id pub-id-type="doi">10.1016/j.omtn.2019.09.019</pub-id> <pub-id pub-id-type="pmid">31678735</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gullo</surname> <given-names>I.</given-names></name> <name><surname>Carneiro</surname> <given-names>F.</given-names></name> <name><surname>Oliveira</surname> <given-names>C.</given-names></name> <name><surname>Almeida</surname> <given-names>G. M.</given-names></name></person-group> (<year>2018</year>). <article-title>Heterogeneity in gastric cancer: from pure morphology to molecular classifications.</article-title> <source><italic>Pathobiology</italic></source> <volume>85</volume> <fpage>50</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1159/000473881</pub-id> <pub-id pub-id-type="pmid">28618420</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gygi</surname> <given-names>S. P.</given-names></name> <name><surname>Corthals</surname> <given-names>G. L.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Rochon</surname> <given-names>Y.</given-names></name> <name><surname>Aebersold</surname> <given-names>R.</given-names></name></person-group> (<year>2000</year>). <article-title>Evaluation of two-dimensional gel electrophoresis-based proteome analysis technology.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>97</volume> <fpage>9390</fpage>&#x2013;<lpage>9395</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.160270797</pub-id> <pub-id pub-id-type="pmid">10920198</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>P.</given-names></name> <name><surname>Yang</surname> <given-names>P.</given-names></name> <name><surname>Zhao</surname> <given-names>P.</given-names></name> <name><surname>Shang</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2019</year>). &#x201C;<article-title>GCN-MF: disease-gene association identification by graph convolutional networks and matrix factorization</article-title>,&#x201D; in <source><italic>Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &#x0026; Data Mining</italic></source>, (<publisher-loc>New York, NY</publisher-loc>: <publisher-name>Association for Computing Machinery</publisher-name>), <fpage>705</fpage>&#x2013;<lpage>713</lpage>.</citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jin</surname> <given-names>Z.</given-names></name> <name><surname>Jiang</surname> <given-names>W.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Biomarkers for gastric cancer: Progression in early diagnosis and prognosis.</article-title> <source><italic>Oncol. Lett.</italic></source> <volume>9</volume> <fpage>1502</fpage>&#x2013;<lpage>1508</lpage>. <pub-id pub-id-type="doi">10.3892/ol.2015.2959</pub-id> <pub-id pub-id-type="pmid">25788990</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>C.</given-names></name> <name><surname>Changlu</surname> <given-names>Q.</given-names></name> <name><surname>He</surname> <given-names>Z.</given-names></name> <name><surname>Tongze</surname> <given-names>F.</given-names></name> <name><surname>Xue</surname> <given-names>Z.</given-names></name></person-group> (<year>2019</year>). <article-title>gutMDisorder: a comprehensive database for dysbiosis of the gut microbiota in disorders and interventions.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>48</volume>:<issue>7603</issue>. <pub-id pub-id-type="doi">10.1093/nar/gkz843</pub-id> <pub-id pub-id-type="pmid">31584099</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mering</surname> <given-names>C. V.</given-names></name> <name><surname>Huynen</surname> <given-names>M.</given-names></name> <name><surname>Jaeggi</surname> <given-names>D.</given-names></name> <name><surname>Schmidt</surname> <given-names>S.</given-names></name> <name><surname>Bork</surname> <given-names>P.</given-names></name> <name><surname>Snel</surname> <given-names>B.</given-names></name></person-group> (<year>2003</year>). <article-title>STRING: a database of predicted functional associations between proteins.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>31</volume> <fpage>258</fpage>&#x2013;<lpage>261</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkg034</pub-id> <pub-id pub-id-type="pmid">12519996</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mo</surname> <given-names>F.</given-names></name> <name><surname>Luo</surname> <given-names>Y.</given-names></name> <name><surname>Fan</surname> <given-names>D. A.</given-names></name> <name><surname>Zeng</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>Y. N.</given-names></name> <name><surname>Luo</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2020</year>). <article-title>Integrated analysis of mRNA-seq and miRNA-seq to identify c-MYC, YAP1 and miR-3960 as major players in the anticancer effects of caffeic acid phenethyl ester in human small cell lung cancer cell line.</article-title> <source><italic>Curr. Gene Ther.</italic></source> <volume>20</volume> <fpage>15</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.2174/1566523220666200523165159</pub-id> <pub-id pub-id-type="pmid">32445454</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Fan</surname> <given-names>Y.</given-names></name> <name><surname>Xu</surname> <given-names>R.</given-names></name> <name><surname>Bai</surname> <given-names>Y.</given-names></name> <name><surname>Bai</surname> <given-names>L.</given-names></name></person-group> (<year>2018</year>). <article-title>Correlations of TNM staging and lymph node metastasis of gastric cancer with MRI features and VEGF expression.</article-title> <source><italic>Cancer Biomark.</italic></source> <volume>23</volume> <fpage>53</fpage>&#x2013;<lpage>59</lpage>. <pub-id pub-id-type="doi">10.3233/cbm-181287</pub-id> <pub-id pub-id-type="pmid">30010108</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Park</surname> <given-names>J. S.</given-names></name> <name><surname>Lee</surname> <given-names>N.</given-names></name> <name><surname>Beom</surname> <given-names>S. H.</given-names></name> <name><surname>Kim</surname> <given-names>H. S.</given-names></name> <name><surname>Lee</surname> <given-names>C.-K.</given-names></name> <name><surname>Rha</surname> <given-names>S. Y.</given-names></name><etal/></person-group> (<year>2018</year>). <article-title>The prognostic value of volume-based parameters using 18 F-FDG PET/CT in gastric cancer according to HER2 status.</article-title> <source><italic>Gastric Cancer</italic></source> <volume>21</volume> <fpage>213</fpage>&#x2013;<lpage>224</lpage>. <pub-id pub-id-type="doi">10.1007/s10120-017-0739-0</pub-id> <pub-id pub-id-type="pmid">28643145</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>J.</given-names></name> <name><surname>Zhao</surname> <given-names>T.</given-names></name></person-group> (<year>2020</year>). <article-title>Reduction in TOM1 expression exacerbates Alzheimer&#x2019;s disease.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>117</volume> <fpage>3915</fpage>&#x2013;<lpage>3916</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1917589117</pub-id> <pub-id pub-id-type="pmid">32047041</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ryu</surname> <given-names>J. W.</given-names></name> <name><surname>Kim</surname> <given-names>H. J.</given-names></name> <name><surname>Lee</surname> <given-names>Y. S.</given-names></name> <name><surname>Myong</surname> <given-names>N. H.</given-names></name> <name><surname>Hwang</surname> <given-names>C. H.</given-names></name> <name><surname>Lee</surname> <given-names>G. S.</given-names></name><etal/></person-group> (<year>2003</year>). <article-title>The proteomics approach to find biomarkers in gastric cancer.</article-title> <source><italic>J. Korean Med. Sci.</italic></source> <volume>18</volume>, <fpage>505</fpage>&#x2013;<lpage>509</lpage>. <pub-id pub-id-type="doi">10.3346/jkms.2003.18.4.505</pub-id> <pub-id pub-id-type="pmid">12923326</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sang</surname> <given-names>L.</given-names></name> <name><surname>Miller</surname> <given-names>J. J.</given-names></name> <name><surname>Corbit</surname> <given-names>K. C.</given-names></name> <name><surname>Giles</surname> <given-names>R. H.</given-names></name> <name><surname>Brauer</surname> <given-names>M. J.</given-names></name> <name><surname>Otto</surname> <given-names>E. A.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Mapping the NPHP-JBTS-MKS protein network reveals ciliopathy disease genes and pathways.</article-title> <source><italic>Cell</italic></source> <volume>145</volume> <fpage>513</fpage>&#x2013;<lpage>528</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2011.04.019</pub-id> <pub-id pub-id-type="pmid">21565611</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seyfried</surname> <given-names>N. T.</given-names></name> <name><surname>Dammer</surname> <given-names>E. B.</given-names></name> <name><surname>Swarup</surname> <given-names>V.</given-names></name> <name><surname>Nandakumar</surname> <given-names>D.</given-names></name> <name><surname>Duong</surname> <given-names>D. M.</given-names></name> <name><surname>Yin</surname> <given-names>L.</given-names></name><etal/></person-group> (<year>2017</year>). <article-title>A multi-network approach identifies protein-specific co-expression in asymptomatic and symptomatic Alzheimer&#x2019;s disease.</article-title> <source><italic>Cell Syst.</italic></source> <volume>4</volume> <fpage>60</fpage>&#x2013;<lpage>72.e4</lpage>. <pub-id pub-id-type="doi">10.1016/j.cels.2016.11.006</pub-id> <pub-id pub-id-type="pmid">27989508</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sz&#x00E1;sz</surname> <given-names>A. M.</given-names></name> <name><surname>L&#x00E1;nczky</surname> <given-names>A.</given-names></name> <name><surname>Nagy</surname> <given-names>&#x00C1;</given-names></name> <name><surname>F&#x00F6;rster</surname> <given-names>S.</given-names></name> <name><surname>Hark</surname> <given-names>K.</given-names></name> <name><surname>Green</surname> <given-names>J. E.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Cross-validation of survival associated biomarkers in gastric cancer using transcriptomic data of 1,065 patients.</article-title> <source><italic>Oncotarget</italic></source> <volume>7</volume> <fpage>49322</fpage>&#x2013;<lpage>49333</lpage>. <pub-id pub-id-type="doi">10.18632/oncotarget.10337</pub-id> <pub-id pub-id-type="pmid">27384994</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tannu</surname> <given-names>N. S.</given-names></name> <name><surname>Hemby</surname> <given-names>S. E.</given-names></name></person-group> (<year>2006</year>). <article-title>Two-dimensional fluorescence difference gel electrophoresis for comparative proteomics profiling.</article-title> <source><italic>Nat. Protoc.</italic></source> <volume>1</volume> <fpage>1732</fpage>&#x2013;<lpage>1742</lpage>. <pub-id pub-id-type="doi">10.1038/nprot.2006.256</pub-id> <pub-id pub-id-type="pmid">17487156</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tianyi</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Valsdottir</surname> <given-names>L. R.</given-names></name> <name><surname>Tianyi</surname> <given-names>Z.</given-names></name> <name><surname>Jiajie</surname> <given-names>P.</given-names></name></person-group> (<year>2020</year>). <article-title>Identifying drug&#x2013;target interactions based on graph convolutional network and deep neural network.</article-title> <source><italic>Brief. Bioinform.</italic></source> <volume>22</volume> <fpage>2141</fpage>&#x2013;<lpage>2150</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa044</pub-id> <pub-id pub-id-type="pmid">32367110</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tipping</surname> <given-names>M. E.</given-names></name> <name><surname>Bishop</surname> <given-names>C. M.</given-names></name></person-group> (<year>1999</year>). <article-title>Mixtures of probabilistic principal component analyzers.</article-title> <source><italic>Neural. Comput.</italic></source> <volume>11</volume>, <fpage>443</fpage>&#x2013;<lpage>482</lpage>. <pub-id pub-id-type="doi">10.1162/089976699300016728</pub-id> <pub-id pub-id-type="pmid">9950739</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Villanueva</surname> <given-names>M. T.</given-names></name></person-group> (<year>2011</year>). <article-title>Combination therapy: update on gastric cancer in East Asia.</article-title> <source><italic>Nat. Rev. Clin. Oncol.</italic></source> <volume>8</volume>:<issue>690</issue>. <pub-id pub-id-type="doi">10.1038/nrclinonc.2011.171</pub-id> <pub-id pub-id-type="pmid">22048625</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>T.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Cheng</surname> <given-names>L.</given-names></name></person-group> (<year>2020a</year>). <article-title>Deep-DRM: a computational method for identifying disease-related metabolites based on graph deep learning approaches.</article-title> <source><italic>Brief. Bioinform.</italic></source> <volume>22</volume>:<issue>bbaa212</issue>. <pub-id pub-id-type="doi">10.1093/bib/bbaa212</pub-id> <pub-id pub-id-type="pmid">33048110</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>T.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Peng</surname> <given-names>J.</given-names></name> <name><surname>Cheng</surname> <given-names>L.</given-names></name></person-group> (<year>2020b</year>). <article-title>DeepLGP: a novel deep learning method for prioritizing lncRNA target genes.</article-title> <source><italic>Bioinformatics</italic></source> <volume>36</volume> <fpage>4466</fpage>&#x2013;<lpage>4472</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa428</pub-id> <pub-id pub-id-type="pmid">32467970</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>T.</given-names></name> <name><surname>Hu</surname> <given-names>Y.</given-names></name> <name><surname>Zang</surname> <given-names>T.</given-names></name> <name><surname>Wang</surname> <given-names>Y.</given-names></name></person-group> (<year>2020c</year>). <article-title>Identifying protein biomarkers in blood for Alzheimer&#x2019;s disease.</article-title> <source><italic>Front. Cell Dev. Biol.</italic></source> <volume>8</volume>:<issue>472</issue>. <pub-id pub-id-type="doi">10.3389/fcell.2020.00472</pub-id> <pub-id pub-id-type="pmid">32626709</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Zeng</surname> <given-names>X.</given-names></name> <name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Zang</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2021a</year>). <article-title>Prediction and collection of protein&#x2013;metabolite interactions.</article-title> <source><italic>Brief. Bioinform.</italic></source> bbab014. <pub-id pub-id-type="doi">10.1093/bib/bbab014</pub-id> <pub-id pub-id-type="pmid">33554247</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhao</surname> <given-names>T.</given-names></name> <name><surname>Lyu</surname> <given-names>S.</given-names></name> <name><surname>Lu</surname> <given-names>G.</given-names></name> <name><surname>Juan</surname> <given-names>L.</given-names></name> <name><surname>Zeng</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2021b</year>). <article-title>SC2disease: a manually curated database of single-cell transcriptome for human diseases.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>49</volume> <fpage>D1413</fpage>&#x2013;<lpage>D1419</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa838</pub-id> <pub-id pub-id-type="pmid">33010177</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>M. J.</given-names></name> <name><surname>Hu</surname> <given-names>Z. Q.</given-names></name> <name><surname>Zhang</surname> <given-names>C. H.</given-names></name> <name><surname>Wu</surname> <given-names>L. Q.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Liang</surname> <given-names>D. S.</given-names></name></person-group> (<year>2020</year>). <article-title>Gene therapy for hemophilia A: where we stand.</article-title> <source><italic>Curr. Gene Ther.</italic></source> <volume>20</volume> <fpage>142</fpage>&#x2013;<lpage>151</lpage>. <pub-id pub-id-type="doi">10.2174/1566523220666200806110849</pub-id> <pub-id pub-id-type="pmid">32767930</pub-id></citation></ref>
</ref-list>
</back>
</article>