<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="methods-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2024.1489116</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Prediction of protein interactions between pine and pine wood nematode using deep learning and multi-dimensional feature fusion</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Wang</surname>
<given-names>Liuyan</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Rongguang</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Guan</surname>
<given-names>Xuemei</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1239040"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yan</surname>
<given-names>Shanchun</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/157837"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>College of Computer and Control Engineering, Northeast Forestry University</institution>, <addr-line>Harbin, Heilongjiang</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Key Laboratory of Sustainable Forest Ecosystem Management, School of Forestry, Northeast Forestry University</institution>, <addr-line>Harbin, Heilongjiang</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Lida Zhang, Shanghai Jiao Tong University, China</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Chao Wang, Chengdu University of Information Technology, China</p>
<p>Shanwen Sun, Northeast Forestry University, China</p>
<p>Ting Yun, Nanjing Forestry University, China</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Xuemei Guan, <email xlink:href="mailto:gxm_maomao1980@nefu.edu.cn">gxm_maomao1980@nefu.edu.cn</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>12</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>15</volume>
<elocation-id>1489116</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>08</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>11</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2024 Wang, Li, Guan and Yan</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Wang, Li, Guan and Yan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Pine Wilt Disease (PWD) is a devastating forest disease that has a serious impact on ecological balance ecological. Since the identification of plant-pathogen protein interactions (PPIs) is a critical step in understanding the pathogenic system of the pine wilt disease, this study proposes a Multi-feature Fusion Graph Attention Convolution (MFGAC-PPI) for predicting plant-pathogen PPIs based on deep learning. Compared with methods based on single-feature information, MFGAC-PPI obtains more 3D characterization information by utilizing AlphaFold and combining protein sequence features to extract multi-dimensional features via Transform with improved GCN. The performance of MFGAC-PPI was compared with the current representative methods of sequence-based, structure-based and hybrid characterization, demonstrating its superiority across all metrics. The experiments showed that learning multi-dimensional feature information effectively improved the ability of MFGAC-PPI in plant and pathogen PPI prediction tasks. Meanwhile, a pine wilt disease PPI network consisting of 2,688 interacting protein pairs was constructed based on MFGAC-PPI, which made it possible to systematically discover new disease resistance genes in pine trees and promoted the understanding of plant-pathogen interactions.</p>
</abstract>
<kwd-group>
<kwd>protein-protein interaction</kwd>
<kwd>pine wilt disease</kwd>
<kwd>deep learning</kwd>
<kwd>multi-dimensional feature</kwd>
<kwd>pine wood nematode (<italic>Bursaphelenchus xylophilus</italic>)</kwd>
</kwd-group>
<counts>
<fig-count count="6"/>
<table-count count="5"/>
<equation-count count="12"/>
<ref-count count="44"/>
<page-count count="13"/>
<word-count count="6858"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Plant Bioinformatics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Pine Wilt Disease (PWD) is a devastating and prevalent forest disease caused by the pine wood nematode (<italic>Bursaphelenchus xylophilus</italic>), which is known as the &#x201c;cancer&#x201d; of pine trees and the &#x201c;avian flu&#x201d; of pine forests. The disease can harm most pine (<italic>Pinus</italic>) plants, and it can destroy entire pine forests within 3 to 5 years from the initial infection, the pine forest resources, natural landscape and ecological environment caused serious damage (<xref ref-type="bibr" rid="B37">Xu et&#xa0;al., 2023</xref>). Therefore, exploring the pathogenic mechanisms of pine wilt disease and achieving effective early prevention have become top priorities in global forest protection.</p>
<p>Protein-protein interactions (PPI) between plants and pathogens are fundamental to understanding infection mechanisms and host response strategies. The first line of defense for plant disease resistance is to recognize pathogen-associated molecular patterns (PAMPs) through cell surface receptors (PRRs), which activates pattern-triggered immunity (PTI) (<xref ref-type="bibr" rid="B40">Yuan et&#xa0;al., 2021</xref>). At this stage, to destroy the host immune system, pathogens will secrete effector proteins that interact directly or indirectly with plant proteins, interfering with the plant&#x2019;s PTI response. To counteract pathogen virulence, plants initiate a second line of defense that specifically recognizes effectors through intracellular resistance proteins (R), thereby activating effector-triggered immunity (ETI) (<xref ref-type="bibr" rid="B27">Naveed et&#xa0;al., 2020</xref>). In summary, the interactions between host resistance proteins and pathogen-effector proteins play a pivotal role in plant-pathogen molecular recognition (<xref ref-type="bibr" rid="B7">Cardoso et&#xa0;al., 2024</xref>). The interaction between the pine wood nematode and pine tree serves as a quintessential model of plant-pathogen relationships, encompassing both the processes of pathogen infection and destruction of the host, as well as host perception and defense against invasion. Therefore, investigating the protein-protein interaction networks between the pine wood nematode and its host pine trees is crucial for elucidating the pathogenic mechanisms of pine wood nematode disease. This understanding is of great significance for achieving early pest control, maintaining forest health and ecosystem balance.</p>
<p>Traditional methods for protein-protein interactions identification, including yeast two-hybrid screen(Y2H) (<xref ref-type="bibr" rid="B32">Uetz et&#xa0;al., 2000</xref>), affinity purification-mass spectrometry (AP-MS) (<xref ref-type="bibr" rid="B11">Gavin et&#xa0;al., 2002</xref>), and co-immunoprecipitation (Co-IP) (<xref ref-type="bibr" rid="B4">Bennett et&#xa0;al., 2010</xref>), were initially widely used in human-virus PPI studies (<xref ref-type="bibr" rid="B44">Zhou et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B18">Kim et&#xa0;al., 2023</xref>). These techniques laid the foundation for phytopathology network studies, resulting in a series of plant-pathogen PPI databases such as HPIDB (<xref ref-type="bibr" rid="B2">Ammari et&#xa0;al., 2016</xref>), PHI-base (<xref ref-type="bibr" rid="B36">Winnenburg et&#xa0;al., 2007</xref>), and UVPID (<xref ref-type="bibr" rid="B15">Jiehua et&#xa0;al., 2019</xref>). However, these methods are time-consuming and costly, which make the study of disease resistance mechanisms in non-model plants generally lack of holistic nature, and experimentally validated PPIs between Pinus sylvestris and Pinus sylvestris nematodes are even fewer (<xref ref-type="bibr" rid="B25">Meng et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B23">Liu et&#xa0;al., 2021</xref>). Consequently, there is an urgent need to develop a fast and accurate plant-pathogen PPI prediction method to elucidate the pathogenic system of pine wilt disease.</p>
<p>Computational prediction methods for PPIs play an increasingly important role benefiting from the rapid development of computational biology. Methods based on machine learning (ML) have been widely applied in early computational modeling studies of cross-species PPI (<xref ref-type="bibr" rid="B31">Tang et&#xa0;al., 2023</xref>). A range of ML methods, have demonstrated effectiveness in PPI prediction tasks for species like human hepatitis C virus and Arabidopsis (<xref ref-type="bibr" rid="B35">Wang et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B1">Ahmed et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B19">Lei et&#xa0;al., 2023</xref>). With the increasing availability of high-throughput sequencing data, traditional ML methods are becoming inadequate for handling vast amounts of data, and deep learning (DL) has attracted wide attention in the field of bioinformatics due to its powerful model expression ability (<xref ref-type="bibr" rid="B41">Zhang et&#xa0;al., 2024</xref>). Protein sequences serve as the primary data source for PPI prediction, and many models leverage sequence information to conduct predictive research, such as DNN-PPI (<xref ref-type="bibr" rid="B20">Li et&#xa0;al., 2018</xref>), DeepFE-PPI (<xref ref-type="bibr" rid="B38">Yao et&#xa0;al., 2019</xref>), PIPR (<xref ref-type="bibr" rid="B9">Chen et&#xa0;al., 2019</xref>), and so on. A series of results have also been achieved in plant-pathogen PPI prediction research work, for example, Zheng (<xref ref-type="bibr" rid="B43">Zheng et&#xa0;al., 2023</xref>) fused protein sequence, structural domain and gene ontology (GO) information to construct a deep learning framework based on the combination of word2vec and RCNN to predict protein-protein interactions in <italic>Arabidopsis thaliana</italic>, and validated its ability to cross-species prediction in different datasets. Pan (<xref ref-type="bibr" rid="B28">Pan et&#xa0;al., 2022</xref>) fuses protein sequence information with behavioral information to predict interactions between different plant proteins using DNN and obtains more than 92% accuracy in a variety of datasets. Li (<xref ref-type="bibr" rid="B21">Li et&#xa0;al., 2022</xref>) proposes a plant-pathogen prediction model by combining position-specific scoring matrices (PSSMs) and evaluates the effectiveness of the model in <italic>Arabidopsis thaliana</italic>, <italic>Zea mays</italic>, and <italic>Oryza sativa</italic> datasets.</p>
<p>Recently, the general interest of researchers in predicting PPI based on structural information has been driven by the rapid growth of three-dimensional (3D) structural data of proteins, especially by the vacated introduction of AIphaFold (<xref ref-type="bibr" rid="B16">Jumper et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B34">Varadi et&#xa0;al., 2022</xref>), a series of structural and graphical neural network-based prediction models have emerged, such as DensePPI (<xref ref-type="bibr" rid="B12">Halsana et&#xa0;al., 2023</xref>), GraphPPIS (<xref ref-type="bibr" rid="B39">Yuan et&#xa0;al., 2022</xref>), TAGPPI (<xref ref-type="bibr" rid="B30">Song et&#xa0;al., 2022</xref>), and Struct2Graph (<xref ref-type="bibr" rid="B3">Baranwal et&#xa0;al., 2022</xref>). Zheng (<xref ref-type="bibr" rid="B42">Zheng et&#xa0;al., 2021</xref>) proposed a computational framework based on structural and homology modeling, which not only predicted the PPI networks of rice and the rice blast fungus, generating an interactions network with 2,018 protein pairs, but also analyzed the network to make systematic discovery of plant disease resistance genes possible. Improvement of PPI prediction in combination with AIphaFold may be a solution to this problem by making predictions based on primary structure, which in turn results in more protein structural information that can be used on a genome-wide scale (<xref ref-type="bibr" rid="B13">Homma et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B6">Bryant et&#xa0;al., 2022</xref>).</p>
<p>In summary, the structural information of proteins is more conserved relative to sequences during evolution and can be obtained with higher accuracy, but the sequence evolution information of proteins is crucial for predicting functionally relevant interactions. Models relying solely on structural information might overlook such vital biological data, limiting their applicability and generalization capability (<xref ref-type="bibr" rid="B17">Kang et&#xa0;al., 2023</xref>; <xref ref-type="bibr" rid="B33">Vajdi et&#xa0;al., 2020</xref>). Therefore, an approach that fuses sequence and structural features by integrating a deep learning framework can reveal the details of protein interactions, which can be beneficial for generalization to species-wide prediction efforts and elucidation of genome-to-phenomenon (<xref ref-type="bibr" rid="B29">Sledzieski et&#xa0;al., 2021</xref>).</p>
<p>In this study, Graph Convolutional Networks (GCN) and attention mechanisms focused on critical interaction nodes, combined with AlphaFold, were used to predict the PPIs between pine wood nematodes and their host pine trees. Compared to the sole use of amino acid sequences or protein structure data, the proposed method took a multi-level feature fusion approach to acquire more comprehensive protein representation information, thereby reducing evolutionary differences in cross-species PPIs and improving predictive accuracy. Moreover, by converting structural data into graph data, it was represented in the form of graph theory to identify new interaction residues in the plant-pathogen system in an unsupervised manner. In addition, PPIs for pine and pine wood nematode proteins were constructed using a multidimensional feature fusion method, providing valuable insights for systematically understanding plantpathogen interactions and the pathogenic mechanisms of Pine wilt disease.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Materials and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Construction of datasets</title>
<p>High-quality training data is crucial for deep learning models, but the availability of known PPI data between plants and pathogens is very limited. Therefore, training models on known protein interaction datasets to predict PPIs in new host-pathogen systems becomes particularly important. In this study, to improve the accuracy of the model while preserving the biological significance to the greatest extent, protein-related data of pine nematode and host pine as well as PPI data verified by biological experiments, were selected to build a dataset suitable for training a pine nematode-pine PPIs prediction model.</p>
<p>First, the raw protein sequence data for both the pine wood nematode and pine tree, which served as the basis for subsequent construction of protein 3D structures and extraction of sequence features, was obtained. The protein sequence data for the pine wood nematode was gained from the NBIC database (Taxonomy ID: 6326), totaling 53,412 sequences. The protein sequence data for the pine tree was obtained from UNIPROT(<ext-link ext-link-type="uri" xlink:href="https://www.uniprot.org">https://www.uniprot.org</ext-link>), totaling 200,806 sequences. The genome annotation project database for pine trees was TreeGenes (<ext-link ext-link-type="uri" xlink:href="https://www.treegenesdb.org">https://www.treegenesdb.org</ext-link>). To minimize potential errors from raw protein sequence data, sequences containing short protein sequences (e.g., lengths less than 50 bp) and homologous sequences were removed.</p>
<p>The experimentally validated protein-protein interaction data between the pine wood nematode and its host pine tree was gained from the PHI-base (<ext-link ext-link-type="uri" xlink:href="https://www.phi-base.org">https://www.phi-base.org</ext-link>) database.</p>
<p>PPI data for both the pine wood nematode and pine tree were queried from the STRING database. Specifically, there were 6,231 interaction pairs for the pine wood nematode and 2,009 interaction pairs for the pine tree. All interaction data selected here were physical interactions, excluding weak and transient interactions. The interaction data between the two were analyzed and compared, and the overlapping data was regarded as interspecific interaction. Finally, combining data obtained from PHI-base, 8,259 protein-protein interaction pairs were constructed using 4,792 proteins, including 5,258 positive reference datasets. It is worth noting that protein interactions can take various forms, only direct physical interactions were considered as positive data. Non-interacting proteins cannot be directly obtained from PPI databases, so proteins with no interactions in the PPI network were marked as negative samples not recorded in the PPI dataset.</p>
<p>AlphaFold can obtain the representation information of 3D structure through protein sequences. Therefore, AlphaFold DB would be used to view the structural information of protein interactions generated based on the above rules, and structure prediction would be made for protein sequences lacking structural information to generate PDB files to establish a protein structure database for PPI prediction of pine wood nematode disease.</p>
<p>According to the structural graph of protein network interactions, an interaction between two proteins was marked as 1, otherwise as 0. There were 8,259 positive and negative samples in total, which were divided into training, testing sets in an 8:2. The validation set was used to determine the optimal parameters for the model, and the trained model with the determined optimal parameters was used for training. To improve model performance, different proportions of positive and negative samples were set, and 1/2 of the positive (negative) samples from the training set were mixed with the negative (positive) samples from the training set to enhance the model&#x2019;s generalization capability under imbalanced sample conditions.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Network overall framework</title>
<p>An end-to-end deep learning framework MFGA-PPI was introduced for identifying plant-pathogen PPIs. Here, we take the sequence and structure information of the two proteins as input and define it as a binary classification problem, with the final output being a set of 0 or 1 predictions as to whether they interact or not. The overall architecture of the MFGAC-PPI model is shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>. This model architecture consisted of three parts: the feature extraction module, the feature aggregation module, and the prediction module. In the first part, to better represent protein structures, both tertiary and primary structure information and design feature extraction modules were used for each. The second part used a linear interpolation method to effectively combine the two feature vectors, achieving a multidimensional protein representation. The third part inputed the resulting protein pairs into an attention network, calculated attention scores, and used them for PPI prediction.</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Overall structure of the model.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g001.tif"/>
</fig>
<sec id="s2_2_1">
<label>2.2.1</label>
<title>Construction of graph data</title>
<p>After encoding the 3D structural features, the spatial information of protein pairs needed to be converted into corresponding protein graphs, and inputed into the graph neural network for processing.</p>
<p>The protein feature coding process was divided into two parts: extraction of tertiary structure features and extraction of primary structure features.</p>
<p>As shown in <xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>, the protein feature encoding process was divided into two parts, i.e., extracting tertiary structure features and extracting primary structure features. Firstly, the main chain of the protein was selected for feature extraction. Due to the different atoms composing the 20 amino acids, each atom had distinct characteristics in different amino acids.</p>
<p>According to the atoms that made up the amino acid residues, four atomic features, namely van der Waals radius, electronic charge, B-factor and atomic mass, were extracted from the PDB file of the tertiary structure, and are used to represent <inline-formula>
<mml:math display="inline" id="im1">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mtext>...</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mn>4</mml:mn>
<mml:mo>;</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mtext>...</mml:mtext>
<mml:mo>,</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the <italic>i</italic>th feature of the <italic>j</italic>th atom in the residue. Then, the average of each atomic feature within the residue is calculated and represented by <inline-formula>
<mml:math display="inline" id="im3">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mtext>&#x2009;</mml:mtext>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>i</italic> represents the <italic>i</italic>th atomic feature of the residue, and <inline-formula>
<mml:math display="inline" id="im4">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> indicates the average feature value of all atoms in the residue. Finally, a four-dimensional atomic feature for each residue was obtained. The calculation formula is as follows:</p>
<disp-formula id="eq1">
<label>(1)</label>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:msub>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>N</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>N</mml:mi>
</mml:munderover>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</disp-formula>
<p>After the atomic features were obtained, the geometric structure of the protein was analyzed to measure the distances between residues and assess the possibility of interactions.</p>
<p>There are various methods to construct geometric shapes based on atomic spatial coordinates. Borgwardt (<xref ref-type="bibr" rid="B5">Borgwardt et&#xa0;al., 2005</xref>) described spatial contacts between atoms through van der Waals forces or hydrogen bonds. Based on the idea of graph theory, Cha (<xref ref-type="bibr" rid="B8">Cha et&#xa0;al., 2022</xref>)took the atomic positions after averaging as the spatial position coordinates of amino acids and regarded them as the vertices of the graph.</p>
<p>Given that graph theory-based methods can reveal the topological structure of protein networks, the concept of graph theory was firstly adopted by aggregating amino acids based on group numbers and calculating the average distance between atoms within the group to obtain the coordinates of each residue. Subsequently, using the residue coordinates, the pairwise distances between residues were calculated using the Euclidean distance. Generally, if the distance between two amino acids in space was less than 8 &#xc5;, they were considered to be in contact. Therefore, a spatial proximity relationship was determined if the distance between residues was less than 8.0 &#xc5; (<xref ref-type="bibr" rid="B26">Mou et&#xa0;al., 2023</xref>). The protein graph of length <italic>l</italic> is a square matrix <inline-formula>
<mml:math display="inline" id="im5">
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mo>=</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mo>}</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> of order <italic>l</italic>, where</p>
<disp-formula id="eq2">
<label>(2)</label>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mtable columnalign="left">
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mn>1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>f</mml:mi>
<mml:mtext>&#x2009;</mml:mtext>
<mml:msub>
<mml:mi>&#x3b1;</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>b</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&lt;</mml:mo>
<mml:mn>8.0</mml:mn>
<mml:mstyle mathvariant="bold" mathsize="normal">
<mml:mi>&#xc5;</mml:mi>
</mml:mstyle>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr columnalign="left">
<mml:mtd columnalign="left">
<mml:mn>0</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>h</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>w</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Finally, we construct an <italic>n</italic> &#xd7; <italic>n</italic> residue matrix was constructed and the amino acid residues to unique integer identifiers were mapped, being added to the atomic features to generate the node feature matrix <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>v</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>4</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>n</italic> is the number of residues and 4 represents the four-dimensional node vector extracted. The constructed protein graph object can be denoted as <inline-formula>
<mml:math display="inline" id="im8">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>=</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>V</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>E</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>A</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>V</italic> is the set of vertices, <italic>E</italic> is the set of edges between them, vertices represent amino acid residues, edges represent spatial proximity relationships between residues, and <italic>A</italic> is the adjacency matrix, <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>Protein sequence features were extracted based on the main chain, which typically consists of three atoms: nitrogen (N), alpha-carbon (C<italic>&#x3b1;</italic>), and carbonyl carbon (C). The alpha-carbon (C<italic>&#x3b1;</italic>) atom is commonly used to mark the position of amino acids (<xref ref-type="bibr" rid="B22">Lin et&#xa0;al., 2022</xref>). The ESM-2 model was employed to extract features from protein sequences. ESM2 is a language model based on the Transformer architecture, which maps protein sequences to representations in a high-dimensional space (<xref ref-type="bibr" rid="B24">Liu and Shen, 2023</xref>). First, the coordinates of the C<italic>&#x3b1;</italic> in the protein structure information were used as reference points for amino acids and traversed each main chain and residue to obtain the amino acid sequence. Then, we use the ESM2 model was used to encode the protein sequence, mapping each protein sequence to a 1280-dimensional vector and obtaining embeddings at the protein level. Finally, the sequence embeddings were added to the graph embeddings computed by the GCN block, which were then normalized to the final output embeddings to obtain the sequence feature vector <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
</sec>
<sec id="s2_2_2">
<label>2.2.2</label>
<title>Improved graph convolutional network processes protein graph</title>
<p>GCN is a convolutional neural network for processing graph data, where nodes represent residues and edges represent relationships between residues. Here the features of different nodes in the protein graph data are aggregated using the improved GCN module, which continuously learns and updates the node features.</p>
<p>Node features first enter a 1D convolution layer to extract local features, followed by processing through a bidirectional gated recurrent unit (GRU) to capture global dependencies of the residue node features and output a feature sequence. Finally, an average pooling layer compressed the GRU output to generate fixed-size feature representations. For the stability of convolution operations, the adjacency matrix <italic>A</italic> is normalized. The feature was propagated to the neighbor node through the adjacency matrix <italic>A</italic>1 and <italic>A</italic>2, which were the adjacency matrices of two proteins respectively, and the updated feature of this node containing the neighbor node was saved. The main calculation formula of GCN module is as follows:</p>
<disp-formula id="eq3">
<label>(3)</label>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
<mml:mo>&#xa0;</mml:mo>
<mml:mover accent="true">
<mml:mi>A</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:msup>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mn>2</mml:mn>
</mml:mfrac>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>A</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mi>A</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, <italic>A</italic> is the adjacency matrix and <italic>I</italic> is the identity matrix. Similarly, <inline-formula>
<mml:math display="inline" id="im12">
<mml:mrow>
<mml:mover accent="true">
<mml:mi>D</mml:mi>
<mml:mo>&#x2dc;</mml:mo>
</mml:mover>
<mml:mo>=</mml:mo>
<mml:mi>D</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, where <italic>D</italic> is the degree matrix of <italic>A</italic>, and <italic>I</italic> is the identity matrix. <inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> represents the updated features of the <italic>l</italic>-th layer residue nodes. When <italic>l</italic> = 0, it indicates that the node has not been updated, thus <inline-formula>
<mml:math display="inline" id="im14">
<mml:mrow>
<mml:msup>
<mml:mi>H</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>v</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im15">
<mml:mrow>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>v</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the initial residue node features. <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:msup>
<mml:mi>W</mml:mi>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> is the trainable weight matrix of the <italic>l</italic>-th layer, and <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is the nonlinear activation function. The feature of the last updated residue node was inputted into <inline-formula>
<mml:math display="inline" id="im18">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mi>g</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and a total of 3 residue node updates were carried out in this experiment.</p>
<p>Therefore, for a pair of proteins A and B, we can extract richer structural feature vectors can be extracted as <inline-formula>
<mml:math display="inline" id="im19">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im20">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im21">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula>
<mml:math display="inline" id="im22">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:msup>
<mml:mi>&#x211d;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula>. Finally, the resulting structural feature vectors were combined with the ESM output <inline-formula>
<mml:math display="inline" id="im23">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im24">
<mml:mrow>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> to obtain <inline-formula>
<mml:math display="inline" id="im25">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:math>
</inline-formula>, where <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mtext>&#x3bb;</mml:mtext>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>2</mml:mn>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mtext>&#x3bb;</mml:mtext>
<mml:msub>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula>, with <inline-formula>
<mml:math display="inline" id="im29">
<mml:mtext>&#x3bb;</mml:mtext>
</mml:math>
</inline-formula> being an adjustable parameter.</p>
</sec>
<sec id="s2_2_3">
<label>2.2.3</label>
<title>Attention network for PPI prediction</title>
<p>The scaled dot-product attention mechanism was employed to compute the attention between F1 and F2. A scaling factor was introduced before the dot-product calculation to balance the magnitude of the results, evaluating which residues play critical roles in the interaction.</p>
<p>When calculating the attention of <italic>F</italic>1 on <italic>F</italic>2, the scaled dot-product attention primarily accepted three parameters: Query (<italic>F</italic>1), Key (<italic>F</italic>2), and Value (<italic>F</italic>2). First, the attention score(AS) was computed by taking the dot product of the query and key matrices. Secondly, to control the range of attention scores and prevent gradient explosion or vanishing, the key vector dimension <inline-formula>
<mml:math display="inline" id="im30">
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> was used to scale it and the scaled attention score (SAS) was obtained. Finally, the attention weight matrix was multiplied by the value matrix to compute the weighted sum. The formula of the attention network is as follows:</p>
<disp-formula id="eq4">
<label>(4)</label>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>A</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:msup>
<mml:mn>2</mml:mn>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq5">
<label>(5)</label>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>S</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:msup>
<mml:mn>2</mml:mn>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq6">
<label>(6)</label>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:msup>
<mml:mn>2</mml:mn>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>,</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:msup>
<mml:mn>1</mml:mn>
<mml:mi>T</mml:mi>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:msub>
<mml:mi>d</mml:mi>
<mml:mi>k</mml:mi>
</mml:msub>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>The final attention scores <inline-formula>
<mml:math display="inline" id="im31">
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im32">
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> were concatenated to obtain <inline-formula>
<mml:math display="inline" id="im33">
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>=</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, which served as the output of the attention mechanism.</p>
</sec>
<sec id="s2_2_4">
<label>2.2.4</label>
<title>FNN layer</title>
<p>A fully connected layer was used to predict the model output. The vector s obtained from the attention layer was used as the input to the fully connected layer, which produced the vector <inline-formula>
<mml:math display="inline" id="im34">
<mml:mrow>
<mml:mi>z</mml:mi>
<mml:mo>=</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>. Finally, a softmax activation function was applied to yield the binary classification result.</p>
<disp-formula id="eq7">
<label>(7)</label>
<mml:math display="block" id="M7">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:munder accentunder="true">
<mml:mo>&#xa0;</mml:mo>
<mml:mo>_</mml:mo>
</mml:munder>
<mml:mi>p</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>s</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>z</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
<mml:mo>+</mml:mo>
<mml:msup>
<mml:mi>e</mml:mi>
<mml:mrow>
<mml:msub>
<mml:mi>z</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Model training and hyperparameter setting</title>
<p>In this study, Intel(R) Xeon(R) Gold 6354 CPU @ 3.00GHz&#xd7;2, NVIDIA GeForce RTX 3090&#xd7;4 graphics processor, 256.0GB memory, 19TB MR9364-8i storage were used, and operating system was Ubuntu 18.04.4 LTS.</p>
<p>In the process of model training, 5-fold cross-validation was adopted. For each fold training set, 1000 samples were randomly selected for training. For each selected sample, the characteristics and labels of the sample were obtained. In this experiment, the Adam optimizer was used to optimize the model. The initial value of the learning rate was 1 &#xd7; 10<sup>&#x2212;3</sup>, and the learning rate attenuated to half of the original value after every 10 rounds of training to prevent the model from jumping out of the optimal solution. In this model, three layers of GCN neural networks are used, and the embedding dimension of GCN in each layer was 20, which reduces the computational overhead on the premise of ensuring sufficient feature information capture. The feature dimension of the protein sequence captured by ESM was 1280 dimensions. Binary cross-entropy is used as the loss function of the model in this paper, and the formula is as follows:</p>
<disp-formula id="eq8">
<label>(8)</label>
<mml:math display="block" id="M8">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>s</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>y</mml:mi>
<mml:mtext>&#xa0;log</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mtext>&#xa0;log</mml:mtext>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where, <italic>y</italic> is the true label and <italic>p</italic> is the probability that the model predicts 1. <italic>y</italic> log(<italic>p</italic>) represents the loss when the true label is 1 and (1 &#x2212; <italic>y</italic>) log(1 &#x2212; <italic>p</italic>) represents the loss when the true label is 0.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Evaluation metrics</title>
<p>In the experiment, eight common indicators were used to evaluate the model: Precision, Recall, Accuracy, Specificity, F1-score, Matthews Correlation Coefficient (MCC), Area Under the P-R Curve(AUPRC) and Area Under the ROC Curve (AUROC). Their relevant definitions are as follows:</p>
<disp-formula id="eq9">
<label>(9)</label>
<mml:math display="block" id="M9">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq10">
<label>(10)</label>
<mml:math display="block" id="M10">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>u</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq11">
<label>(11)</label>
<mml:math display="block" id="M11">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>p</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>t</mml:mi>
<mml:mi>y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
<mml:mi>F</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mi>r</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>o</mml:mi>
<mml:mi>n</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>e</mml:mi>
<mml:mi>c</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>l</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="eq12">
<label>(12)</label>
<mml:math display="block" id="M12">
<mml:mrow>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>C</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>P</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mi>N</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>F</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mrow>
</mml:math>
</disp-formula>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results</title>
<sec id="s3_1">
<label>3.1</label>
<title>Ablation experiment</title>
<p>The basic framework of the MFGAC-PPI model mainly consisted of four submodules: ESM, improved GCN, Scaled dot-product attention, and FNN. According to the aforementioned hyperparameter settings, each submodule was treated as a different variable and the &#x201c;control variable method&#x201d; was used to investigate the influence and contribution of each submodule to the proposed model. Therefore, 5-fold cross-validation was used to conduct ablation studies on these four structures, six indicators were used for evaluation, and the maximum value of the results of each ablation was selected, as shown in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. It can be seen that, when any of the sub-modules was ablated, the overall prediction performance of the model was degraded, indicating that each submodule played a role, and the structural design was reasonable without structural redundancy. Similarly, it was found that, the model performance decreases the most among the evaluation metrics when the improved GCN module was ablated, especially specificity and MCC, which dropped by 1.73% and 2.19%, respectively. In contrast, when the FNN module was ablated, the model performance changed the least across all metrics, with precision only dropping by 0.01%, and the largest change being in MCC, which decreased by 1.39%.</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>The performance results for different modules.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Module ablation</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Precision</th>
<th valign="middle" align="center">Recall</th>
<th valign="middle" align="center">Specificity</th>
<th valign="middle" align="center">F1</th>
<th valign="middle" align="center">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">Attention</td>
<td valign="bottom" align="center">0.9779</td>
<td valign="bottom" align="center">0.9915</td>
<td valign="bottom" align="center">0.9796</td>
<td valign="bottom" align="center">0.9770</td>
<td valign="bottom" align="center">0.9816</td>
<td valign="bottom" align="center">0.9400</td>
</tr>
<tr>
<td valign="top" align="center">ESM</td>
<td valign="top" align="center">0.9751</td>
<td valign="top" align="center">0.9887</td>
<td valign="top" align="center">0.9756</td>
<td valign="top" align="center">0.9770</td>
<td valign="top" align="center">0.9802</td>
<td valign="top" align="center">0.9431</td>
</tr>
<tr>
<td valign="top" align="center">GCN</td>
<td valign="top" align="center">0.9723</td>
<td valign="top" align="center">0.9836</td>
<td valign="top" align="center">0.9741</td>
<td valign="top" align="center">0.9655</td>
<td valign="top" align="center">0.9761</td>
<td valign="top" align="center">0.9340</td>
</tr>
<tr>
<td valign="top" align="center">FNN</td>
<td valign="top" align="center">0.9779</td>
<td valign="top" align="center">0.9915</td>
<td valign="top" align="center">0.9798</td>
<td valign="top" align="center">0.9809</td>
<td valign="top" align="center">0.9638</td>
<td valign="top" align="center">0.9420</td>
</tr>
<tr>
<td valign="top" align="center">All</td>
<td valign="top" align="center">
<bold>0.9812</bold>
</td>
<td valign="top" align="center">
<bold>0.9916</bold>
</td>
<td valign="top" align="center">
<bold>0.9893</bold>
</td>
<td valign="top" align="center">
<bold>0.9828</bold>
</td>
<td valign="top" align="center">
<bold>0.9904</bold>
</td>
<td valign="top" align="center">
<bold>0.9559</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>*Best performance is shown in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>From the results, it was evident that the improved GCN module contributed the most to the overall model and significantly affected its prediction performance. In contrast, the FNN had the least impact on the predictive ability of the model.</p>
<p>To assess the importance of each sub-module more intuitively and comprehensively, the results of the five-fold cross-validation in each fold of the six evaluation metrics are presented here using box plots, as shown in <xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>. It can be visually observed from the figure that when the improved GCN was ablated, the model showed the lowest average values and the lowest outliers across all metrics. The model without FNN showed the most stable performance in the five-fold cross-validation, with the highest average values across all metrics. However, in precision and specificity, it had nearly the same values as the model without the attention module, but the model without attention had lower dispersion points. In summary, each sub-module had an important contribution and influenced the model&#x2019;s effectiveness to some extent. Among them, the improved GCN contributed the most to the model&#x2019;s overall performance, followed by the ESM module, both of which reflected the importance of feature fusion to some extent. Then there was attention and FNN, with FNN having the smallest impact on the overall model.</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>The 5-fold cross-validation of performance results.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g002.tif"/>
</fig>
<p>The model constructed in this study differed from the GAT module by introducing a Scaled dot-product attention to capture subtle differences within proteins and their interaction interfaced while retaining more interpretable biological characteristics, which helped to identify amino acid residues with important functions, and had better robustness when dealing with proteins of different lengths. To validate that the introduction of the scaling factor improved the predictive ability of the model, as well as to explore the impact of different attention mechanisms on the model, three new models were added to the original structure: self-attention replacing scaled dot-product attention, mutual-attention replacing scaled dot-product attention, and multi-head-attention replacing scaled dot-product attention. Keeping other sub-modules and hyperparameters unchanged, these models were subjected to ablation experiments using 5-fold cross-validation, and evaluated using three comprehensive metrics: specificity, F1, and MCC for evaluation. <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> shows the maximum values in the cross-validation.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>The performance results of different attention modules.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Attention module</th>
<th valign="middle" align="center">Specificity</th>
<th valign="middle" align="center">F1</th>
<th valign="middle" align="center">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">Mutual-Attention</td>
<td valign="bottom" align="center">0.9743</td>
<td valign="bottom" align="center">0.9795</td>
<td valign="bottom" align="center">0.9370</td>
</tr>
<tr>
<td valign="top" align="center">Self-Attention</td>
<td valign="top" align="center">0.9798</td>
<td valign="top" align="center">0.9842</td>
<td valign="top" align="center">0.9545</td>
</tr>
<tr>
<td valign="top" align="center">Multihead-Attention</td>
<td valign="top" align="center">0.9655</td>
<td valign="top" align="center">0.9808</td>
<td valign="top" align="center">0.9418</td>
</tr>
<tr>
<td valign="top" align="center">All</td>
<td valign="top" align="center">
<bold>0.9828</bold>
</td>
<td valign="top" align="center">
<bold>0.9847</bold>
</td>
<td valign="top" align="center">
<bold>0.9559</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>*Best performance is shown in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>To clearly illustrate the variation of each fold value in cross-validation, radar charts of Specificity, F1, and MCC on the validation dataset were plotted as shown in <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>. It can be seen that the model using mutual-attention had the smallest area in the metrics, while the MFGAC-PPI model using the scaled dot-product attention mechanism had the largest area in the three comprehensive metrics, indicating a significant reduction in false positive (FP) samples and an increase in true positive (TP) samples. Moreover, it was apparent from <xref ref-type="fig" rid="f3">
<bold>Figures&#xa0;3A, C</bold>
</xref> that the Specificity indicator&#x2019;s area for the multi-head-attention model was much larger than that for the self-attention model, but the area in MCC was the opposite.</p>
<fig id="f3" position="float">
<label>Figure&#xa0;3</label>
<caption>
<p>The 5-fold cross-validation radar map. <bold>(A)</bold> specificity 5-fold cross-validation radar map. <bold>(B)</bold> F1 5-fold cross-validation radar map. <bold>(C)</bold> MCC 5-fold cross-validation radar map.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g003.tif"/>
</fig>
<p>This may be due to the different focus of the two attention mechanisms when processing input data. Multi-head-attention captured different feature subspaces through multiple heads, performing better on negative sample features, while self-attention uses a single attention head to compute attention globally, focusing on optimizing overall performance, thus performing better in F1 and MCC metrics.</p>
<p>Combining <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref> and <xref ref-type="fig" rid="f3">
<bold>Figure&#xa0;3</bold>
</xref>, it was found that the applicability of the attention modules, from highest to lowest, was scaled dot-product attention &gt; self-attention &gt; multihead-attention &gt; mutual-attention.</p>
<p>Additionally, different &#x3bb; values were tested to evaluate the contribution of sequence and structural features to the model during feature fusion. During model training, the &#x3bb; value for feature fusion in the constructed MFGAC-PPI model was adjusted while keeping other parameters constant, and the results are shown in <xref ref-type="table" rid="T3">
<bold>Table&#xa0;3</bold>
</xref>. It can be seen that when the &#x3bb; values were set to 0.5 and 0.7, the performance of various evaluation metrics was the optimal, especially when &#x3bb;=0.7, F1 and MCC reached the best.</p>
<table-wrap id="T3" position="float">
<label>Table&#xa0;3</label>
<caption>
<p>The performance results for different &#x3bb; values.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">&#x3bb; values</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Specificity</th>
<th valign="middle" align="center">F1</th>
<th valign="middle" align="center">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">0</td>
<td valign="bottom" align="center">0.9735</td>
<td valign="bottom" align="center">0.9565</td>
<td valign="bottom" align="center">0.9789</td>
<td valign="bottom" align="center">0.9439</td>
</tr>
<tr>
<td valign="top" align="center">0.1</td>
<td valign="top" align="center">0.9831</td>
<td valign="top" align="center">0.9743</td>
<td valign="top" align="center">0.9797</td>
<td valign="top" align="center">0.9454</td>
</tr>
<tr>
<td valign="top" align="center">0.3</td>
<td valign="top" align="center">0.9763</td>
<td valign="top" align="center">0.9655</td>
<td valign="top" align="center">0.9761</td>
<td valign="top" align="center">0.9522</td>
</tr>
<tr>
<td valign="top" align="center">0.5</td>
<td valign="top" align="center">
<bold>0.9812</bold>
</td>
<td valign="top" align="center">
<bold>0.9828</bold>
</td>
<td valign="top" align="center">0.9833</td>
<td valign="top" align="center">0.9528</td>
</tr>
<tr>
<td valign="top" align="center">0.7</td>
<td valign="top" align="center">0.9807</td>
<td valign="top" align="center">0.9733</td>
<td valign="top" align="center">
<bold>0.9904</bold>
</td>
<td valign="top" align="center">
<bold>0.9559</bold>
</td>
</tr>
<tr>
<td valign="top" align="center">0.9</td>
<td valign="top" align="center">0.9708</td>
<td valign="top" align="center">0.9742</td>
<td valign="top" align="center">0.9818</td>
<td valign="top" align="center">0.9437</td>
</tr>
<tr>
<td valign="top" align="center">1</td>
<td valign="top" align="center">0.9704</td>
<td valign="top" align="center">0.9731</td>
<td valign="top" align="center">0.9833</td>
<td valign="top" align="center">0.9419</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>*Best performance is shown in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Comparison with other competitive methods</title>
<p>The constructed MFGAC-PPI model were compared with five classic PPI prediction models: DeepPPI (<xref ref-type="bibr" rid="B10">Du et&#xa0;al., 2017</xref>), PIPR (<xref ref-type="bibr" rid="B9">Chen et&#xa0;al., 2019</xref>), Struct2Graph (<xref ref-type="bibr" rid="B3">Baranwal et&#xa0;al., 2022</xref>), AFTGAN (<xref ref-type="bibr" rid="B17">Kang et&#xa0;al., 2023</xref>), and TAGPPI (<xref ref-type="bibr" rid="B30">Song et&#xa0;al., 2022</xref>), and the results were shown in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>. The proposed model exhibited the best performance across accuracy, precision, recall, specificity, and F1, although MCC was slightly lower than that of Struct2Graph. TAGPPI&#x2019;s performance was close to the best across all metrics, particularly specificity, which was only 0.0017 lower than MFGAC-PPI. The sequence-based models PIPR and DeepPPI showed relatively low performance in MCC but maintained a relatively balanced performance across other metrics. Meanwhile, AFTGAN showed poor performance on the constructed dataset, significantly lagging behind other methods in key metrics such as F1 and recall, which were 0.1588 and 0.1555 lower than MFGAC-PPI, respectively, indicating a notable deficiency in recognizing positive samples.</p>
<table-wrap id="T4" position="float">
<label>Table&#xa0;4</label>
<caption>
<p>The performance evaluation results of different methods.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Methods</th>
<th valign="middle" align="center">Accuracy</th>
<th valign="middle" align="center">Precision</th>
<th valign="middle" align="center">Recall</th>
<th valign="middle" align="center">Specificity</th>
<th valign="middle" align="center">F1</th>
<th valign="middle" align="center">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">DeepPPI</td>
<td valign="bottom" align="center">0.9435</td>
<td valign="bottom" align="center">0.9503</td>
<td valign="bottom" align="center">0.9527</td>
<td valign="bottom" align="center">0.9702</td>
<td valign="bottom" align="center">0.9515</td>
<td valign="bottom" align="center">0.9436</td>
</tr>
<tr>
<td valign="top" align="center">PIPR</td>
<td valign="top" align="center">0.9609</td>
<td valign="top" align="center">0.9617</td>
<td valign="top" align="center">0.9756</td>
<td valign="top" align="center">0.9613</td>
<td valign="top" align="center">0.9609</td>
<td valign="top" align="center">0.9316</td>
</tr>
<tr>
<td valign="top" align="center">Struct2-Graph</td>
<td valign="top" align="center">0.9796</td>
<td valign="top" align="center">0.9830</td>
<td valign="top" align="center">0.9725</td>
<td valign="top" align="center">0.9543</td>
<td valign="top" align="center">
<bold>0.9777</bold>
</td>
<td valign="top" align="center">0.9725</td>
</tr>
<tr>
<td valign="top" align="center">AFTGAN</td>
<td valign="top" align="center">0.8437</td>
<td valign="top" align="center">0.8295</td>
<td valign="top" align="center">0.8338</td>
<td valign="top" align="center">0.9511</td>
<td valign="top" align="center">0.8316</td>
<td valign="top" align="center">0.9582</td>
</tr>
<tr>
<td valign="top" align="center">TAGPPI</td>
<td valign="top" align="center">0.9781</td>
<td valign="top" align="center">0.9710</td>
<td valign="top" align="center">0.9726</td>
<td valign="top" align="center">0.9811</td>
<td valign="top" align="center">0.9718</td>
<td valign="top" align="center">0.9525</td>
</tr>
<tr>
<td valign="top" align="center">MFGAC-PPI</td>
<td valign="top" align="center">
<bold>0.9812</bold>
</td>
<td valign="top" align="center">
<bold>0.9916</bold>
</td>
<td valign="top" align="center">
<bold>0.9893</bold>
</td>
<td valign="top" align="center">
<bold>0.9828</bold>
</td>
<td valign="top" align="center">0.9904</td>
<td valign="top" align="center">
<bold>0.9559</bold>
</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>*Best performance is shown in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>To further assess the robustness and generalization capability of the model, the proposed model was compared with five other competitive algorithms using the original dataset and an independent dataset, ara data (<xref ref-type="bibr" rid="B43">Zheng et&#xa0;al., 2023</xref>). The AUROC and P-R curves for these models on both datasets were calculated, as shown in <xref ref-type="fig" rid="f4">
<bold>Figures&#xa0;4</bold>
</xref>, <xref ref-type="fig" rid="f5">
<bold>5</bold>
</xref>.</p>
<fig id="f4" position="float">
<label>Figure&#xa0;4</label>
<caption>
<p>AUROC results compared with competing methods. <bold>(A)</bold> ROC curve verified using the dataset built in this article. <bold>(B)</bold> ROC curve verified using the ara data dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g004.tif"/>
</fig>
<fig id="f5" position="float">
<label>Figure&#xa0;5</label>
<caption>
<p>P-R results compared with competing methods. <bold>(A)</bold> P-R curve verified using the dataset built in this article. <bold>(B)</bold> P-R curve verified using the ara data dataset.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g005.tif"/>
</fig>
<p>As shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4A</bold>
</xref>, in the constructed original dataset, the model&#x2019;s AUC area was the largest, reaching 0.95, which was 0.03, 0.09, 0.13, 0.17, and 0.19 higher than TAGPPI, Struct2Graph, PIPR, DeepPPI, and AFTGAN, respectively. Additionally, the analysis of the P-R curve in <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5A</bold>
</xref> indicated that MFGAC-PPI and Struct2Graph performed the best. The results were consistent with the performance of various metrics in <xref ref-type="table" rid="T4">
<bold>Table&#xa0;4</bold>
</xref>, where the top three models in overall performance on the constructed dataset were MFGAC-PPI, TAGPPI, and Struct2Graph, respectively. As shown in <xref ref-type="fig" rid="f4">
<bold>Figure&#xa0;4B</bold>
</xref>, when training on the independent dataset, the AUC of MFGAC-PPI dropped by 0.03, whereas TAGPPI increased by 0.03, making TAGPPI to be the best performance in AUROC. The AUROC curves of the other models remained relatively stable, similar to the results on the original dataset, although Struct2Graph shows a slight decline. In addition, the analysis of <xref ref-type="fig" rid="f5">
<bold>Figure&#xa0;5B</bold>
</xref> showed that the performance of MFGAC-PPI on the P-R curve in the ara data dataset declined slightly, while Struct2Graph exhibited the best performance.</p>
<p>In general, MFGAC-PPI demonstrated the best performance in predictive ability and comprehensive metrics, but its performance declined when tested in different data sets, indicating that its generalization ability had room for improvement.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Verification of unbalanced data</title>
<p>In actual biological systems, most proteins interact with specific proteins rather than random combinations. Therefore, only a small fraction of protein pair combinations has true interactions, especially when PPI-related data is scarce, and the imbalance between positive and negative samples in protein interactions can reach a ratio of 1:100 or more. To verify the superiority of the proposed model in the case of unbalanced samples, the MFGAC-PPI model was trained using the constructed dataset with balanced (1:1) to unbalanced (1:3, 1:5, 1:10, 1:20, 1:30) data, using precision, recall, and MCC as evaluation metrics to illustrate the advantages of the proposed model in imbalanced datasets.</p>
<p>The results are shown in <xref ref-type="table" rid="T5">
<bold>Table&#xa0;5</bold>
</xref>, the MFGAC-PPI model showed better performance in various indexes in unbalanced data sets. Although the performance of each index decreased with the increase of the degree of imbalance, the precision and MCC remained above 91%. At the same time, recall remained relatively stable, indicating that the model can correctly identify positive samples even under highly imbalanced conditions. Overall, the MFGA-PPI model was robust, but its overall predictive performance decreased with increasing unevenness, especially affecting precision and MCC.</p>
<table-wrap id="T5" position="float">
<label>Table&#xa0;5</label>
<caption>
<p>The performance evaluation results of unbalanced dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">P:N ratio</th>
<th valign="middle" align="center">Precision</th>
<th valign="middle" align="center">Recall</th>
<th valign="middle" align="center">MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="bottom" align="center">1:1</td>
<td valign="bottom" align="center">
<bold>0.9916</bold>
</td>
<td valign="bottom" align="center">
<bold>0.9893</bold>
</td>
<td valign="bottom" align="center">
<bold>0.9559</bold>
</td>
</tr>
<tr>
<td valign="bottom" align="center">1:3</td>
<td valign="top" align="center">0.9798</td>
<td valign="top" align="center">0.9873</td>
<td valign="top" align="center">0.9496</td>
</tr>
<tr>
<td valign="bottom" align="center">1:5</td>
<td valign="top" align="center">0.9655</td>
<td valign="top" align="center">0.9502</td>
<td valign="top" align="center">0.9460</td>
</tr>
<tr>
<td valign="bottom" align="center">1:10</td>
<td valign="top" align="center">0.9258</td>
<td valign="top" align="center">0.9455</td>
<td valign="top" align="center">0.9319</td>
</tr>
<tr>
<td valign="bottom" align="center">1:20</td>
<td valign="top" align="center">0.9131</td>
<td valign="top" align="center">0.9421</td>
<td valign="top" align="center">0.9201</td>
</tr>
<tr>
<td valign="bottom" align="center">1:30</td>
<td valign="top" align="center">0.9124</td>
<td valign="top" align="center">0.9322</td>
<td valign="top" align="center">0.9157</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>*Best performance is shown in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Generated protein interaction network in pine wood nematode disease</title>
<p>The optimized plant-pathogen PPI prediction model was used to predict 20,765 sets of pine nematode and pine tree protein interaction data, and the full prediction results are shown in <xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary Table&#xa0;1</bold>
</xref>. In particular, protein pairs with interaction scores greater than 0.5 were judged to have
interactions, and a total of 2,688 pine-pine wood nematode protein-protein interaction networks were generated, which contained 46 pine wood nematode proteins and 354 pine proteins, with the results shown in <xref ref-type="supplementary-material" rid="SM2">
<bold>Supplementary Table&#xa0;2</bold>
</xref>. Among the predicted pine-pine wood nematode PPIs results, 16 predicted PPIs were validated by previous biological experiments. Meanwhile, the comparison of different models revealed that about 19% or so of the PPIs could be derived from sequence-based prediction models, and there existed about 21% of the interaction pairs that could be predicted by structure-based methods, and there was a high degree of overlap between these two parts of the PPIs. These results indicated that a multi-dimensional feature fusion approach can effectively uncover new PPI data and was successfully applied in the study of the pine wilt disease system.</p>
<p>A topological analysis of the predicted PPI network revealed that pine-pine wood nematode PPIs
exhibited scale-free properties similar to other biological networks. Notably, pine nematode proteins had more interaction links than pine proteins, with one pine nematode protein able to interact with an average of four pine proteins and at least 20 pine wood nematode proteins interact with over 10 pine proteins each. Here the nodes with higher degrees were analyzed for centrality, and the results are shown in <xref ref-type="supplementary-material" rid="SF1">
<bold>Supplementary Figure&#xa0;1</bold>
</xref>, and a PPI network diagram was drawn as shown in <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6A</bold>
</xref>. From <xref ref-type="supplementary-material" rid="SF1">
<bold>Supplementary Figure&#xa0;1</bold>
</xref> and <xref ref-type="fig" rid="f6">
<bold>Figure&#xa0;6A</bold>
</xref>, it can be seen that the proteins with the most links are the effector protein <italic>BxSap1</italic> of pine nematode as well as the autophagy gene <italic>BxATG16</italic>, which played crucial roles in the virulence of the pine wood nematode. This was followed by <italic>A0A1l7SCF8 BURXY</italic>, <italic>Bx tlp 1</italic> of pine wood nematode and <italic>P41649.2</italic> protein of pine. This result suggested that potentially pathogen-associated proteins were more involved than resistance-associated proteins in the pine nematode system. Surface representations of the 3D structures of the interacting proteins are shown in <xref ref-type="fig" rid="f6">
<bold>Figures&#xa0;6B, C</bold>
</xref>, demonstrating that MFGAC-PPI can efficiently predict regions of plant-pathogen protein interaction residues, illustrated by the surface interaction between effector proteins <italic>SapB3</italic> (blue) and <italic>P41649.2</italic> (green), with the interaction region in red. By revealing 3D structural analysis and protein surface interaction regions, it helped understand molecular communication between pine trees and pathogens, explaining how pathogens evaded or suppressed the host immune response and effectively invaded host cells.</p>
<fig id="f6" position="float">
<label>Figure&#xa0;6</label>
<caption>
<p>Schematic diagram of protein-protein interaction networks. <bold>(A)</bold> protein-protein interaction network diagram of pine wood nematode disease system. <bold>(B)</bold> is the schematic diagram of the interaction between effector protein SapB3 and P41649.2, and <bold>(C)</bold> is the schematic diagram after 180 rotation.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-15-1489116-g006.tif"/>
</fig>
</sec>
</sec>
<sec id="s4" sec-type="conclusions">
<label>4</label>
<title>Conclusion</title>
<p>
<italic>Bursaphelenchus xylophilus</italic> is one of the most devastating forest pathogens worldwide. Predicting and analyzing plant-pathogen protein interactions is a crucial step in understanding the molecular mechanisms of plant diseases. In this study, we propose MFGAC-PPI, an improved graph attention convolutional network-based deep learning method for plant-pathogen PPI prediction. It was leveraged multi-level feature fusion to provide a comprehensive research perspective and accurately predict protein-protein interactions. By utilizing AlphaFold to obtain more 3D structural information of pathogen proteins and extracting features from both amino acid sequences and structural information using the Transformer structure and GCN, the prediction accuracy was enhanced. Additionally, the scaled dot-product attention mechanism identified important interacting residues in an unsupervised manner, facilitating downstream analysis. Experimental results indicated that MFGAC-PPI achieved high accuracy on two datasets, with an AUC exceeding 92%, and performed well on imbalanced datasets. It outperformed current state-of-the-art prediction methods, making it suitable for plant-pathogen interaction prediction tasks.</p>
<p>Through the optimized plant-pathogen interaction prediction model, we generated a pine wood nematode disease PPIs comprising 2,688 interacting protein pairs involving 36 <italic>Pinus</italic> proteins and 356 <italic>B. xylophilus</italic> proteins. Notably, <italic>B. xylophilus</italic> proteins exhibited more interaction relationships and partners compared to <italic>Pinus</italic> proteins. This involved pathogen-related proteins in plant-pathogen interactions, potentially due to co-evolutionary arms race dynamics. The predicted PPI networks successfully identified interactions such as <italic>BxSap1</italic> effector with <italic>Pinus</italic> PR protein (<italic>PtPR-1b</italic>), previously validated through experimental methods (<xref ref-type="bibr" rid="B14">Hu et&#xa0;al., 2019</xref>). The results demonstrate that the MFGAC-PPI model&#x2019;s successful application in the pine wilt disease system provides a comprehensive PPI network, aiding in the identification of resistance genes and advancing our understanding of plant-pathogen interaction mechanisms.</p>
<p>This study revealed that embedding protein sequence information into protein structural representations can extract more effective biological information, improving the accuracy of PPI prediction tasks. The performance metrics of this approach surpassed those of single-dimension protein feature learning methods. In the future, it is expected that new strategies for protein representation learning will be applied to other prediction tasks, such as protein function prediction. By combining these strategies with the protein structure prediction techniques, more high-precision structural data can be obtained, thereby expanding the model applicability and enhancing its generalization capabilities.</p>
</sec>
</body>
<back>
<sec id="s5" sec-type="data-availability">
<title>Data availability statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="SM1">
<bold>Supplementary Material</bold>
</xref>. Further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6" sec-type="author-contributions">
<title>Author contributions</title>
<p>LW: Writing &#x2013; original draft. RL: Writing &#x2013; review &amp; editing. XG: Writing &#x2013; review &amp; editing. SY: Writing &#x2013; review &amp; editing.</p>
</sec>
<sec id="s7" sec-type="funding-information">
<title>Funding</title>
<p>The author(s) declare that financial support was received for the research, authorship, and/or publication of this article. This work is supported in part by funds from the Open Grant for Key Laboratory of Sustainable Forest Ecosystem Management (Northeast Forestry University), Ministry of Education (KFJJ2023YB03), Manufacturing Innovation Talent Project of Harbin Science and Technology Bureau (CXRC20221110393) and National Natural Science Foundation of China (32171691).</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>The authors thank the anonymous reviewers for their valuable suggestions.</p>
</ack>
<sec id="s8" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The reviewer SS declared a shared affiliation with the authors to the handling editor at the time of review.</p>
</sec>
<sec id="s9" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s10" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2024.1489116/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpls.2024.1489116/full#supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Image1.tiff" id="SF1" mimetype="image/tiff"/>
<supplementary-material xlink:href="Table1.xlsx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
<supplementary-material xlink:href="Table2.xlsx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ahmed</surname> <given-names>F. F.</given-names>
</name>
<name>
<surname>Khatun</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Mosharaf</surname> <given-names>M. P.</given-names>
</name>
<name>
<surname>Mollah</surname> <given-names>M. N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Prediction of protein-protein interactions in arabidopsis thaliana using partial training samples in a machine learning framework</article-title>. <source>Curr. Bioinf.</source> <volume>16</volume>, <fpage>865</fpage>&#x2013;<lpage>879</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2174/1574893616666210204145254</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ammari</surname> <given-names>M. G.</given-names>
</name>
<name>
<surname>Gresham</surname> <given-names>C. R.</given-names>
</name>
<name>
<surname>McCarthy</surname> <given-names>F. M.</given-names>
</name>
<name>
<surname>Nanduri</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Hpidb 2.0: a curated database for host&#x2013;pathogen interactions</article-title>. <source>Database</source> <volume>2016</volume>, <fpage>baw103</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/database/baw103</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baranwal</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Magner</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Saldinger</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Turali-Emre</surname> <given-names>E. S.</given-names>
</name>
<name>
<surname>Elvati</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Kozarekar</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Struct2graph: a graph attention network for structure based predictions of protein&#x2013;protein interactions</article-title>. <source>BMC Bioinf.</source> <volume>23</volume>, <fpage>370</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12859-022-04910-9</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bennett</surname> <given-names>E. J.</given-names>
</name>
<name>
<surname>Rush</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Gygi</surname> <given-names>S. P.</given-names>
</name>
<name>
<surname>Harper</surname> <given-names>J. W.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Dynamics of cullin-ring ubiquitin ligase network revealed by systematic quantitative proteomics</article-title>. <source>Cell</source> <volume>143</volume>, <fpage>951</fpage>&#x2013;<lpage>965</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cell.2010.11.017</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Borgwardt</surname> <given-names>K. M.</given-names>
</name>
<name>
<surname>Ong</surname> <given-names>C. S.</given-names>
</name>
<name>
<surname>Sch&#xf6;nauer</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Vishwanathan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Smola</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Kriegel</surname> <given-names>H.-P.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Protein function prediction via graph kernels</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>i47</fpage>&#x2013;<lpage>i56</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/bti1007</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bryant</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Pozzati</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Elofsson</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Improved prediction of protein-protein interactions using alphafold2</article-title>. <source>Nat. Commun.</source> <volume>13</volume>(<issue>1</issue>), <fpage>1265</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41467-022-28865-w</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cardoso</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Manadas</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Abrantes</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Robertson</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Arcos</surname> <given-names>S. C.</given-names>
</name>
<name>
<surname>Troya</surname> <given-names>M. T.</given-names>
</name>
<etal/>
</person-group>. (<year>2024</year>). <article-title>Pine wilt disease: what do we know from proteomics</article-title>? <source>BMC Plant Biol.</source> <volume>24</volume>, <fpage>98</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12870-024-04771-9</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cha</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Emre</surname> <given-names>E. S. T.</given-names>
</name>
<name>
<surname>Xiao</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Kim</surname> <given-names>J.-Y.</given-names>
</name>
<name>
<surname>Bogdan</surname> <given-names>P.</given-names>
</name>
<name>
<surname>VanEpps</surname> <given-names>J. S.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Unifying structural descriptors for biological and bioinspired nanoscale complexes</article-title>. <source>Nat. Comput. Sci.</source> <volume>2</volume>, <fpage>243</fpage>&#x2013;<lpage>252</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s43588-022-00229-w</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ju</surname> <given-names>C. J.-T.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Chang</surname> <given-names>K.-W.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Multifaceted protein&#x2013;protein interaction prediction based on siamese residual rcnn</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>i305</fpage>&#x2013;<lpage>i314</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btz328</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Du</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Hu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yan</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Deepppi: boosting prediction of protein&#x2013;protein interactions with deep neural networks</article-title>. <source>J. Chem. Inf. modeling</source> <volume>57</volume>, <fpage>1499</fpage>&#x2013;<lpage>1510</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1021/acs.jcim.7b00028</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gavin</surname> <given-names>A.-C.</given-names>
</name>
<name>
<surname>B&#xf6;sche</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Krause</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Grandi</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Marzioch</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Bauer</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2002</year>). <article-title>Functional organization of the yeast proteome by systematic analysis of protein complexes</article-title>. <source>Nature</source> <volume>415</volume>, <fpage>141</fpage>&#x2013;<lpage>147</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/415141a</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Halsana</surname> <given-names>A. A.</given-names>
</name>
<name>
<surname>Chakroborty</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Halder</surname> <given-names>A. K.</given-names>
</name>
<name>
<surname>Basu</surname> <given-names>S.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Denseppi: A novel image-based deep learning method for prediction of protein&#x2013;protein interactions</article-title>. <source>IEEE Trans. NanoBioscience</source> <volume>22</volume>, <fpage>904</fpage>&#x2013;<lpage>911</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1109/TNB.2023.3251192</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Homma</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>van der Hoorn</surname> <given-names>R. A.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Alphafold-multimer predicts cross-kingdom interactions at the plant-pathogen interface</article-title>. <source>Nat. Commun.</source> <volume>14</volume>, <fpage>6040</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41467-023-41721-9</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hu</surname> <given-names>L.-J.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>X.-Q.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>H.-Y.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Y.-C.</given-names>
</name>
<name>
<surname>Ye</surname> <given-names>J.-R.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>An effector, bxsapb1, induces cell death and contributes to virulence in the pine wood nematode bursaphelenchus xylophilus</article-title>. <source>Mol. Plant-Microbe Interact.</source> <volume>32</volume>, <fpage>452</fpage>&#x2013;<lpage>463</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1094/MPMI-10-18-0275-R</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiehua</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Shuai</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Yizhen</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Shiwen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Yanjun</surname> <given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Ustilaginoidea virens: A fungus infects rice flower and threats world rice production</article-title>. <source>Rice Sci.</source> <volume>26</volume>, <fpage>199</fpage>&#x2013;<lpage>206</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.rsci.2018.10.007</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jumper</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Evans</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Pritzel</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Green</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Figurnov</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ronneberger</surname> <given-names>O.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Highly accurate protein structure prediction with alphafold</article-title>. <source>nature</source> <volume>596</volume>, <fpage>583</fpage>&#x2013;<lpage>589</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41586-021-03819-2</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Elofsson</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Aftgan: prediction of multi-type ppi based on attention free transformer and graph attention network</article-title>. <source>Bioinformatics</source> <volume>39</volume>, <fpage>btad052</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btad052</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>D.-K.</given-names>
</name>
<name>
<surname>Weller</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lin</surname> <given-names>C.-W.</given-names>
</name>
<name>
<surname>Sheykhkarimli</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Knapp</surname> <given-names>J. J.</given-names>
</name>
<name>
<surname>Dugied</surname> <given-names>G.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>A proteome-scale map of the sars-cov-2&#x2013;human contactome</article-title>. <source>Nat. Biotechnol.</source> <volume>41</volume>, <fpage>140</fpage>&#x2013;<lpage>149</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41587-022-01475-z</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lei</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>He</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Arapathogen2. 0: An improved prediction of plant&#x2013;pathogen protein&#x2013;protein interactions empowered by the natural language processing technique</article-title>. <source>J. Proteome Res.</source> <volume>23</volume>, <fpage>494</fpage>&#x2013;<lpage>499</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1021/acs.jproteome.3c00364</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Gong</surname> <given-names>X.-J.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Deep neural network based predictions of protein interactions using primary sequences</article-title>. <source>Molecules</source> <volume>23</volume>(<issue>8</issue>), <elocation-id>1923</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/molecules23081923</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>L.-P.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Cheng</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Cpiela: Computational prediction of plant protein&#x2013;protein interactions by ensemble learning approach from protein sequences and evolutionary information</article-title>. <source>Front. Genet.</source> <volume>13</volume>, <elocation-id>857839</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2022.857839</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Akin</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Rao</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Hie</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Lu</surname> <given-names>W.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Language models of protein sequences at the scale of evolution enable accurate structure prediction</article-title>. <source>BioRxiv</source> <volume>2022</volume>, <fpage>500902</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1101/2022.07.20.500902</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Two terpene synthases in resistant pinus massoniana contribute to defence against bursaphelenchus xylophilus</article-title>. <source>Plant Cell Environ.</source> <volume>44</volume>, <fpage>257</fpage>&#x2013;<lpage>274</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/pce.13873</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Shen</surname> <given-names>H.-B.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Foldexplorer: Fast and accurate protein structure search with sequenceenhanced graph embedding</article-title>. <source>arXiv preprint arXiv:2311.18219</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.21203/rs.3.rs-3831396/v1</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meng</surname> <given-names>F.-L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Y.-X.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.-Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Expression analysis of thaumatin-like proteins from bursaphelenchus xylophilus and pinus massoniana</article-title>. <source>Physiol. Mol. Plant Pathol.</source> <volume>100</volume>, <fpage>178</fpage>&#x2013;<lpage>184</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.pmpp.2017.10.002</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Pan</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Shi</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>A transformer-based ensemble framework for the prediction of protein&#x2013;protein interaction sites</article-title>. <source>Research</source> <volume>6</volume>, <elocation-id>0240</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.34133/research.0240</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Naveed</surname> <given-names>Z. A.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Mubeen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Ali</surname> <given-names>G. S.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>The pti to eti continuum in phytophthora-plant interactions</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>, <elocation-id>593905</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2020.593905</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pan</surname> <given-names>J.</given-names>
</name>
<name>
<surname>You</surname> <given-names>Z.-H.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>L.-P.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>W.-Z.</given-names>
</name>
<name>
<surname>Guo</surname> <given-names>J.-X.</given-names>
</name>
<name>
<surname>Yu</surname> <given-names>C.-Q.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Dwppi: a deep learning approach for predicting protein&#x2013;protein interactions in plants based on multi-source information with a large-scale biological network</article-title>. <source>Front. Bioengineering Biotechnol.</source> <volume>10</volume>, <elocation-id>807522</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fbioe.2022.807522</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sledzieski</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Cowen</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Berger</surname> <given-names>B.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>D-script translates genome to phenome with sequence-based, structure-aware, genome-scale predictions of protein-protein interactions</article-title>. <source>Cell Syst.</source> <volume>12</volume>, <fpage>969</fpage>&#x2013;<lpage>982</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cels.2021.08.010</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Niu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>X.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Learning spatial structures of proteins improves protein&#x2013;protein interaction prediction</article-title>. <source>Briefings Bioinf.</source> <volume>23</volume>, <fpage>bbab558</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbab558</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tang</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Peng</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Yin</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Machine learning on protein&#x2013;protein interaction prediction: models, challenges and trends</article-title>. <source>Briefings Bioinf.</source> <volume>24</volume>, <fpage>bbad076</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbad076</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Uetz</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Giot</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Cagney</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Mansfield</surname> <given-names>T. A.</given-names>
</name>
<name>
<surname>Judson</surname> <given-names>R. S.</given-names>
</name>
<name>
<surname>Knight</surname> <given-names>J. R.</given-names>
</name>
<etal/>
</person-group>. (<year>2000</year>). <article-title>A comprehensive analysis of protein&#x2013;protein interactions in saccharomyces cerevisiae</article-title>. <source>Nature</source> <volume>403</volume>, <fpage>623</fpage>&#x2013;<lpage>627</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/35001009</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vajdi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Zarringhalam</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Haspel</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Patch-dca: improved protein interface prediction by utilizing structural information and clustering dca scores</article-title>. <source>Bioinformatics</source> <volume>36</volume>, <fpage>1460</fpage>&#x2013;<lpage>1467</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btz791</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Varadi</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Anyango</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Deshpande</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Nair</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Natassia</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Yordanova</surname> <given-names>G.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Alphafold protein structure database: massively expanding the structural coverage of protein-sequence space with high-accuracy models</article-title>. <source>Nucleic Acids Res.</source> <volume>50</volume>, <fpage>D439</fpage>&#x2013;<lpage>D444</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkab1061</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zou</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Machine learning for phytopathology: from the molecular scale towards the network scale</article-title>. <source>Briefings Bioinf.</source> <volume>22</volume>, <fpage>bbab037</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bib/bbab037</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Winnenburg</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Urban</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Beacham</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Baldwin</surname> <given-names>T. K.</given-names>
</name>
<name>
<surname>Holland</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Lindeberg</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2007</year>). <article-title>Phi-base update: additions to the pathogen&#x2013;host interaction database</article-title>. <source>Nucleic Acids Res.</source> <volume>36</volume>, <fpage>D572</fpage>&#x2013;<lpage>D576</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkm858</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xu</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ren</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Luo</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Pine wilt disease in northeast and northwest China: A comprehensive risk review</article-title>. <source>Forests</source> <volume>14</volume>, <fpage>174</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/f14020174</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Du</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Diao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>An integration of deep learning with feature embedding for protein&#x2013;protein interaction prediction</article-title>. <source>PeerJ</source> <volume>7</volume>, <fpage>e7126</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.7717/peerj.7126</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Structure-aware protein&#x2013;protein interaction site prediction using deep graph convolutional network</article-title>. <source>Bioinformatics</source> <volume>38</volume>, <fpage>125</fpage>&#x2013;<lpage>132</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btab643</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Ngou</surname> <given-names>B. P. M.</given-names>
</name>
<name>
<surname>Ding</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Xin</surname> <given-names>X.-F.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Pti-eti crosstalk: an integrative view of plant immunity</article-title>. <source>Curr. Opin. Plant Biol.</source> <volume>62</volume>, <fpage>102030</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.pbi.2021.102030</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Durham</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Cong</surname> <given-names>Q.</given-names>
</name>
</person-group> (<year>2024</year>). <article-title>Revolutionizing protein&#x2013;protein interaction prediction with deep learning</article-title>. <source>Curr. Opin. Struct. Biol.</source> <volume>85</volume>, <fpage>102775</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.sbi.2024.102775</pub-id>
</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Sun</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Predicting protein&#x2013;protein interactions between rice and blast fungus using structure-based approaches</article-title>. <source>Front. Plant Sci.</source> <volume>12</volume>, <elocation-id>690124</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2021.690124</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zheng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Yang</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Wuchty</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
</person-group> (<year>2023</year>). <article-title>Deep learning-assisted prediction of protein&#x2013;protein interactions in arabidopsis thaliana</article-title>. <source>Plant J.</source> <volume>114</volume>, <fpage>984</fpage>&#x2013;<lpage>994</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/tpj.v114.4</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Gupta</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Paramo</surname> <given-names>M. I.</given-names>
</name>
<name>
<surname>Hou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Mao</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>A comprehensive sars-cov-2&#x2013; human protein&#x2013;protein interactome reveals covid-19 pathobiology and potential host therapeutic targets</article-title>. <source>Nat. Biotechnol.</source> <volume>41</volume>, <fpage>128</fpage>&#x2013;<lpage>139</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41587-022-01474-0</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>