<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fgene.2021.745228</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Prediction of Protein&#x2013;Protein Interactions in <italic>Arabidopsis</italic>, Maize, and Rice by Combining Deep Neural Network With Discrete Hilbert Transform</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Pan</surname> <given-names>Jie</given-names></name>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1415285/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Li</surname> <given-names>Li-Ping</given-names></name>
<xref ref-type="corresp" rid="c002"><sup>&#x002A;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>You</surname> <given-names>Zhu-Hong</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/406779/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yu</surname> <given-names>Chang-Qing</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Ren</surname> <given-names>Zhong-Hao</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/1449677/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Guan</surname> <given-names>Yong-Jian</given-names></name>
</contrib>
</contrib-group>
<aff><institution>School of Information Engineering, Xijing University</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Robert Friedman, Retired, Columbia, SC, United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Loris Nanni, University of Padua, Italy; Huang Yu-an, Shenzhen University, China; Zhi-An Huang, City University of Hong Kong, Hong Kong, SAR China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Jie Pan, <email>jiepan960930@gmail.com</email></corresp>
<corresp id="c002">Li-Ping Li, <email>cs2bioinformatics@gmail.com</email></corresp>
<fn fn-type="other" id="fn004"><p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>20</day>
<month>09</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>745228</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>08</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2021 Pan, Li, You, Yu, Ren and Guan.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Pan, Li, You, Yu, Ren and Guan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Protein&#x2013;protein interactions (PPIs) in plants play an essential role in the regulation of biological processes. However, traditional experimental methods are expensive, time-consuming, and need sophisticated technical equipment. These drawbacks motivated the development of novel computational approaches to predict PPIs in plants. In this article, a new deep learning framework, which combined the discrete Hilbert transform (DHT) with deep neural networks (DNN), was presented to predict PPIs in plants. To be more specific, plant protein sequences were first transformed as a position-specific scoring matrix (PSSM). Then, DHT was employed to capture features from the PSSM. To improve the prediction accuracy, we used the singular value decomposition algorithm to decrease noise and reduce the dimensions of the feature descriptors. Finally, these feature vectors were fed into DNN for training and predicting. When performing our method on three plant PPI datasets <italic>Arabidopsis thaliana</italic>, maize, and rice, we achieved good predictive performance with average area under receiver operating characteristic curve values of 0.8369, 0.9466, and 0.9440, respectively. To fully verify the predictive ability of our method, we compared it with different feature descriptors and machine learning classifiers. Moreover, to further demonstrate the generality of our approach, we also test it on the yeast and human PPI dataset. Experimental results anticipated that our method is an efficient and promising computational model for predicting potential plant&#x2013;protein interacted pairs.</p>
</abstract>
<kwd-group>
<kwd>deep neural networks</kwd>
<kwd>discrete hilbert transform</kwd>
<kwd>plant</kwd>
<kwd>protein&#x2013;protein interactions</kwd>
<kwd>position-specific scoring matrix</kwd>
</kwd-group>
<counts>
<fig-count count="9"/>
<table-count count="8"/>
<equation-count count="13"/>
<ref-count count="56"/>
<page-count count="11"/>
<word-count count="8139"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="S1">
<title>Introduction</title>
<p>Identification of protein&#x2013;protein interactions (PPIs) in plants is essential for exploring the mechanisms underlying of biological processes, such as organ formation, homeostasis control (<xref ref-type="bibr" rid="B7">Canovas et al., 2004</xref>), plant defense (<xref ref-type="bibr" rid="B54">Zhang et al., 2010</xref>), signal transduction (<xref ref-type="bibr" rid="B28">Khan and Kihara, 2016</xref>), and stress response (<xref ref-type="bibr" rid="B5">Bracha-Drori et al., 2004</xref>). Although numerous high-throughput techniques have been developed to identify PPIs of model species, such as affinity purification mass spectrometry (<xref ref-type="bibr" rid="B14">Fukao, 2012</xref>; <xref ref-type="bibr" rid="B3">Armean et al., 2013</xref>) and yeast two-hybrid (<xref ref-type="bibr" rid="B8">Causier and Davies, 2002</xref>; <xref ref-type="bibr" rid="B13">Fang et al., 2002</xref>), these approaches are cumbersome, costly, particularly time consuming, and always suffer from high false positive rate. To overcome these problems, there is an urgent need to develop sequence-based computational methods that can accurately predict potential PPIs while analyzing the functions of plant genes.</p>
<p>In recent years, many studies have been introduced for detecting PPIs. These methods can be broadly classified into several categories: protein structure&#x2013;based method (<xref ref-type="bibr" rid="B20">Hayashi et al., 2018</xref>), genomic information&#x2013;based method (<xref ref-type="bibr" rid="B51">Zahiri et al., 2014</xref>), evolutionary relationship&#x2013;based approach (<xref ref-type="bibr" rid="B47">Xu et al., 2011</xref>), and protein sequence&#x2013;based method (<xref ref-type="bibr" rid="B40">Richoux et al., 2019</xref>). In fact, the first three methods have better prediction performance. However, these methods typically require the structural details of proteins such as 3D structural and protein homology information. If this prior knowledge is not available, then the method will not perform as expected. Theoretically, amino acid sequence contains all the necessary information to detect PPIs. In addition, with the improvement of sequencing technology, more and more plant genome sequences are available. Hence, it is meaningful to develop computational methods to predict potential PPIs from sequence information.</p>
<p>To date, some new approaches have been proposed to predict PPIs using the feature descriptors of protein sequence, such as the composition-transition-distribution descriptor (<xref ref-type="bibr" rid="B48">Yang et al., 2010</xref>), auto-covariance descriptor (<xref ref-type="bibr" rid="B17">Guo et al., 2008</xref>), Zernike moments descriptor (<xref ref-type="bibr" rid="B46">Wang et al., 2017</xref>), and local descriptor (<xref ref-type="bibr" rid="B11">Davies et al., 2008</xref>). These descriptors summarize specific aspects of amino acid sequence, including frequencies of local patterns, physicochemical properties, and positional distribution of protein sequence. However, the coverage of these feature descriptors is still limited. Recently, many deep learning techniques also have been applied on PPI-based prediction. For example, <xref ref-type="bibr" rid="B12">Du et al. (2017)</xref> presented an approach called DeepPPI, which adopted deep neural networks (DNN) to extract high-level features from raw input features of protein sequence to identify PPIs. <xref ref-type="bibr" rid="B52">Zeng et al. (2020)</xref> were inspired by the deep learning algorithm and proposed a framework called DeepPPISP, which extracts local and global features from amino acid sequences and employs DNN to predict PPIs. <xref ref-type="bibr" rid="B44">Sun et al. (2017)</xref> employed stacked autoencoder (SAE), which is a deep learning algorithm to predict PPIs from human protein sequence. <xref ref-type="bibr" rid="B19">Hashemifar et al. (2018)</xref> developed a novel sequence-based approach called DPPI that used Siamese-like convolutional neural networks (CNN) combined with data augmentation and random projection to improve PPI prediction. <xref ref-type="bibr" rid="B41">Sledzieski et al. (2021)</xref> proposed a novel model named D-SCRIPT, which indicated that employing a deep learning language modeling of protein sequence data is effective for PPI prediction. <xref ref-type="bibr" rid="B9">Chen et al. (2019)</xref> put forward an end-to-end framework that combined contextualized information and local features with a deep residual recurrent CNN in the Siamese architecture to predict PPIs only using protein sequence information. <xref ref-type="bibr" rid="B49">Yi et al. (2018)</xref> proposed the RPI-SAN model using a deep learning stacked autoencoder network to extract features from RNA and amino acid sequences. Finally, they fed these features to the RF model for training and predicting. Despite these advances in previous studies, there is still a need to improve the accuracy and efficiency of the PPI prediction models.</p>
<p>In this article, we combined DNN with discrete Hilbert transform (DHT) and singular value decomposition (SVD) to predict PPIs in plants. More specifically, for each plant primary sequence, position-specific score matrix (PSSM) was constructed, and then DHT was applied to gather important information from the protein PSSM. Subsequently, SVD algorithm was adopted to reduce feature dimension and noise interference and finally generated a 600-dimensional feature vector. Lastly, a deep neural network was applied to make predictions between target plant proteins. When the proposed method was applied on the <italic>Arabidopsis thaliana</italic>, maize (<italic>Zea mays</italic>), and rice (<italic>Oryza sativa</italic>) PPI datasets, it yielded promising results of average AUC (area under ROC curve) values of 0.8369, 0.9466, and 0.9440. When compared with some different feature selection methods and state-of-the-art machine learning classifiers, our method obtained better results. In addition, to achieve more convincing evidence, we also applied our method to the yeast and human PPI dataset. These combined results suggest that the proposed approach is effective and trustworthy for predicting potential PPIs in plants.</p>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec id="S2.SS1">
<title>Data Collection and Construction of the Benchmarking Set</title>
<p>To validate the robustness and effectiveness of the proposed model, we performed it on three plant PPI datasets, <italic>A. thaliana</italic>, <italic>Z. mays</italic>, and <italic>O. sativa</italic>. The <italic>A. thaliana</italic> dataset was collected from TAIR<sup><xref ref-type="fn" rid="footnote1">1</xref></sup> (<xref ref-type="bibr" rid="B39">Rhee et al., 2003</xref>), IntAct<sup><xref ref-type="fn" rid="footnote2">2</xref></sup> (<xref ref-type="bibr" rid="B27">Kerrien et al., 2012</xref>), and BioGRID<sup><xref ref-type="fn" rid="footnote3">3</xref></sup> (<xref ref-type="bibr" rid="B42">Stark et al., 2006</xref>). After removing the redundancy, the final <italic>A. thaliana</italic>&#x2013;positive dataset comprised 28,110 PPI pairs containing 7,437 <italic>A. thaliana</italic> proteins. These protein-interacted pairs constructed the primary <italic>A. thaliana</italic> PPI network. For the construction of the negative dataset, we employed a bipartite to formulate a network of plant PPIs, where the nodes represent the plant proteins and the links denote the interactions between them. Here, we use <italic>A. thaliana</italic> as an example. The whole associations between the 7,437 proteins are 55,308,969 (7,437 &#x00D7; 7,437) in the corresponding bipartite. However, only 28,110 PPIs had been demonstrated to have the interactions. Thus, the possible number of negative pairs is 55,280,859 (55,308,969&#x2013;28,110), which is significantly more than the positive samples. To handle this binary classification problem, we randomly collected 28,110 non-interacting pairs as the negative dataset. In theoretical terms, the negative samples may contain a small number of positive samples; however, given the size of the whole non-interaction pairs, the probability of this situation is very small. In this way, the whole <italic>A. thaliana</italic> dataset consists of 56,220 protein pairs.</p>
<p>Maize and rice are the main cash crops in the world. The maize (<italic>Z. mays</italic>) dataset contains 14,800 positive pairs, which was downloaded from PPIM<sup><xref ref-type="fn" rid="footnote4">4</xref></sup> (<xref ref-type="bibr" rid="B55">Zhu et al., 2016</xref>) and agriGO<sup><xref ref-type="fn" rid="footnote5">5</xref></sup> (<xref ref-type="bibr" rid="B45">Tian et al., 2017</xref>). Similarly, we assumed that the proteins in different subcellular work compartments have no interactions and finally achieved 14,800 non-interacting protein pairs. The rice (<italic>O. sativa</italic>) dataset consisted of 9,600 protein pairs, 4,800 positive pairs, and 4,800 negative pairs collected from the PRIN database<sup><xref ref-type="fn" rid="footnote6">6</xref></sup> (<xref ref-type="bibr" rid="B16">Gu et al., 2011</xref>).</p>
</sec>
<sec id="S2.SS2">
<title>Representation of the Plant Amino Acid Sequence</title>
<p>To mine highly efficient features for training the models, each protein pair is encoded as 800-dimensional feature vector by PSSM (<xref ref-type="bibr" rid="B15">Gribskov et al., 1987</xref>). PSSM has been successfully employed in various fields of biological research including the prediction of PPI site, subcellular localization, and DNA-binding protein identification. In this section, we applied PSI-BLAST (<xref ref-type="bibr" rid="B2">Altschul and Koonin, 1998</xref>) tool to represent protein sequence as a <italic>U</italic> &#x00D7; 20 matrix, where <italic>Q</italic> = {<italic>&#x03B7;</italic><sub><italic>a</italic>,<italic>b</italic></sub>:<italic>a</italic> = 1&#x22EF;<italic>U</italic> <italic>a</italic><italic>n</italic><italic>d</italic>b = 1&#x22EF;20}, and it can obtain the information of plant sequential evolution. PSSM can be defined as</p>
<disp-formula id="S2.Ex1"><label>(1)</label><mml:math id="M1" display="block">
<mml:mrow>
<mml:mi>Q</mml:mi>
<mml:mo>=</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mtable displaystyle="true" rowspacing="0pt">
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo rspace="8.1pt">,</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo rspace="8.1pt">,</mml:mo>
<mml:mrow>
<mml:mpadded width="+5.6pt">
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo rspace="8.1pt">,</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo rspace="8.1pt">,</mml:mo>
<mml:mrow>
<mml:mpadded width="+5.6pt">
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mrow>
<mml:mpadded>
<mml:mi mathvariant="normal">&#x22EE;</mml:mi>
</mml:mpadded>
<mml:mo mathvariant="italic" separator="true">&#x2003;&#x2003;&#x2006;</mml:mo>
<mml:mi mathvariant="normal">&#x22EE;</mml:mi>
<mml:mo mathvariant="italic" separator="true">&#x2003;</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo mathvariant="italic" separator="true">&#x2003;</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="left">
<mml:mrow>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mi>U</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mo>&#x2062;</mml:mo>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo rspace="8.1pt">,</mml:mo>
<mml:mrow>
<mml:mpadded width="+5.6pt">
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>&#x03B7;</mml:mi>
<mml:mrow>
<mml:mi>U</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>20</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03B7;<sub><italic>a,b</italic></sub> represents probability that the <italic>a-</italic>th mutate to <italic>b-</italic>th amino acid during the evolutionary process. In the experiment, plant protein sequences were adopted as seeds to search and align homogenous sequences from SwissProt database by PSI-BLAST tool. The tool will be used to recognize members of gene family and evolutionary relationships between plant protein sequences. It is also able to generate a 20-dimensional vector to denote the probabilities of conservation against mutations to the 20 amino acids. The number of iterations is set to 3 and the <italic>E</italic>-value is cut off at 0.001 to achieve homologous sequences. The PSI-BLAST tool and SwissProt database can be accessed online<sup><xref ref-type="fn" rid="footnote7">7</xref></sup>.</p>
</sec>
<sec id="S2.SS3">
<title>Discrete Hilbert Transform</title>
<p>In this section, we introduce discrete Hilbert transform (DHT; <xref ref-type="bibr" rid="B10">Cizek, 1970</xref>) to extract feature descriptors from the PSSM to make the prediction more convenient and accurate. DHT is used as a tool for signal analysis in the time and frequency domains. Before describing the 2-dimensional DHT, the 1-D DHT (<xref ref-type="bibr" rid="B37">Ponomareva et al., 2018</xref>) is used in the spatial and frequency domain and has been previously described (<xref ref-type="bibr" rid="B43">Stark, 1971</xref>; <xref ref-type="bibr" rid="B4">Bracewell and Bracewell, 1986</xref>; <xref ref-type="bibr" rid="B56">Zhu et al., 1990</xref>; <xref ref-type="bibr" rid="B36">Onodera et al., 2005</xref>).</p>
<p>To better extract the feature descriptors, we used the 2-D DHT for constructing the local energy of PSSM. In this work, we applied the 2-D DHT, which is defined by <xref ref-type="bibr" rid="B38">Read and Treitel (1973)</xref> in the frequency domain. Our Matlab code is shown as follows:</p>
<p>function x = hilbert2(xr,m,n)</p>
<p>%HILBERT2 Discrete-time 2D analytic signal via Hilbert transform.</p>
<p>% X = HILBERT2(Xr) computes the 2D discrete-time analytic signal</p>
<p>% X = Xr + i<sup>&#x2217;</sup>Xi such that Xi is the Hilbert transform of real image Xr.</p>
<p>% If the input Xr is complex, then only the real part is used: Xr = real(Xr).</p>
<p>% HILBERT2(Xr,M,N) computes the MxN-point Hilbert transform. Xr is padded</p>
<p>% zeros if it has less than MxN points, and truncated if it has more.</p>
<p>if nargin &#x003C; 2, n = []; end</p>
<p>if &#x223C;isreal (xr)</p>
<p>&#x00A0;&#x00A0;&#x00A0;warning (&#x2019;HILBERT2 ignores imaginary part of input.&#x2019;)</p>
<p>&#x00A0;&#x00A0;&#x00A0;xr = real (xr);</p>
<p>end</p>
<p>if isempty (n)</p>
<p>&#x00A0;&#x00A0;&#x00A0;[m, n] = size (xr);</p>
<p>end</p>
<p>if <italic>m</italic> &#x003C; 2 | | <italic>n</italic> &#x003C; 2,</p>
<p>&#x00A0;&#x00A0;&#x00A0;x = Hilbert (xr); % 1D analytic signal</p>
<p>&#x00A0;&#x00A0;&#x00A0;return;</p>
<p>end;</p>
<p>In this work, PSI-BLAST encoded each protein sequence as a <italic>U</italic>&#x00D7;20 matrix. Due to the different lengths of protein sequences, the size of each matrix constructed by PSSM is also different. To handle this problem, we transformed the variably sized PSSM into a 20&#x00D7;20 matrix, and the 2-D DHT is applied to extract feature vectors from the PSSM profile. In this way, each plant protein sequence will be converted into a 400-dimensional vector by 2-D DHT. As a non-linear filtering technique, SVD has been widely applied in noise reduction of vibration signals. This is because the signals after noise reduction have a small phase-shift and there is no time delay effect. To improve the prediction accuracy and reduce the dimensionality of the input feature matrix, we applied SVD (<xref ref-type="bibr" rid="B31">Klema and Laub, 1980</xref>) algorithm to reduce size of feature vectors from 400 to 300. At the same time, the lower dimensions could reduce the complexity of the model and increase the generalization error of the classifier. Finally, each protein pair will be represented as a 600-dimensional DHT descriptor.</p>
</sec>
<sec id="S2.SS4">
<title>Deep Neural Networks</title>
<p>Considering the larger numbers of hidden layers that can be used for training networks, artificial neural networks consist of two or more hidden layers that are often referred as DNN as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. The depth of a neural network relates to the quantity of hidden layers, and the largest number of neurons determines the width of DNN (<xref ref-type="bibr" rid="B22">Hinton et al., 2006</xref>; <xref ref-type="bibr" rid="B23">Hinton and Salakhutdinov, 2006</xref>).</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The construction of deep neural networks.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g001.tif"/>
</fig>
<p>In terms of structure, DNN is composed of many plain modules, which appear as a multilayer stack. The data are first received by the input layer, and then converted through a non-linear way across many hidden layers. Before calculating the final output, the average gradient is first computed and the corresponding weights are adjusted. Neurons of a hidden layer or input layer are associated with the neurons of the existing layer. Each neuron will compute a weighted sum of its input and perform a non-linear activation function to capture its outputs. The non-linear activation functions usually include sigmoid, rectified linear unit (ReLU), and hyperbolic tangent. In this work, we used the sigmoid and ReLU. We constructed a DNN-based model using the TensorFlow platform shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. This model consists of two hidden layers with 48 neurons each. The DHT feature descriptors are employed as the inputs for the DNN model. After that, these features were set into the hidden layers for training and predicting PPIs. Adam algorithm (<xref ref-type="bibr" rid="B30">Kingma and Ba, 2014</xref>), which is an adaptive learning rate approach, was adopted in our methods to accelerate the training process. At the same time, to avoid overfitting, the dropout technique was also applied to our model (<xref ref-type="bibr" rid="B29">Khan et al., 2019</xref>). We also used the cross-entropy loss and ReLU activation function to speed our training and achieve better predictive performance (<xref ref-type="bibr" rid="B21">Hinton et al., 2015</xref>). The loss can be calculated by the following formulas:</p>
<disp-formula id="S2.Ex2"><label>(2)</label><mml:math id="M2" display="block">
<mml:mrow>
<mml:msubsup>
<mml:mo mathvariant="italic">R</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo rspace="8.1pt">;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex3"><mml:math id="M3" display="block">
<mml:mrow>
<mml:msubsup>
<mml:mo mathvariant="italic">R</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mi>m</mml:mi>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mpadded width="+1.7pt">
<mml:mi>i</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo rspace="8.1pt">;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex4"><label>(3)</label><mml:math id="M4" display="block">
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>4</mml:mn>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>;</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex5"><label>(4)</label><mml:math id="M5" display="block">
<mml:mrow>
<mml:msubsup>
<mml:mo mathvariant="italic">R</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mn>3</mml:mn>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:msubsup>
<mml:mo>&#x2295;</mml:mo>
<mml:msubsup>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex6"><label>(5)</label><mml:math id="M6" display="block">
<mml:mrow>
<mml:msubsup>
<mml:mo mathvariant="italic">R</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mn>3</mml:mn>
</mml:msubsup>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>T</mml:mi>
<mml:mmultiscripts>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mpadded width="+1.7pt">
<mml:mi>i</mml:mi>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:none/>
<mml:mprescripts/>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:none/>
</mml:mmultiscripts>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>=</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi mathvariant="normal">&#x22EF;</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex7"><mml:math id="M7" display="block">
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mo>=</mml:mo>
<mml:mo>-</mml:mo>
<mml:mfrac>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:mfrac>
<mml:munderover>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>n</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex8"><label>(6)</label><mml:math id="M8" display="block">
<mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>-</mml:mo>
<mml:msub>
<mml:mi>y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>-</mml:mo>
<mml:msub>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mi>T</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mi>R</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>b</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>In Eqs. 2&#x2013;6, <italic>n</italic> describes the amount of protein pairs that need to be trained, <italic>m</italic> denotes the individual network, <italic>h</italic><sub><italic>1</italic></sub> represents the depth of two individual networks, and <italic>h</italic><sub><italic>2</italic></sub> denotes the depth of the fused network. The activation function of ReLU and output layer with sigmoid is &#x03C3;<sub><italic>1</italic></sub> and &#x03C3;<sub><italic>2</italic></sub>, respectively; &#x2295; is the concatenation operator. <italic>R</italic> represents the output of hidden layer and <italic>y</italic> is the corresponding desired output. <italic>T</italic> and <italic>b</italic> indicate the weight matrix and bias vectors.</p>
</sec>
</sec>
<sec sec-type="results" id="S3">
<title>Results</title>
<sec id="S3.SS1">
<title>Evaluation Criteria</title>
<p>To prevent overfitting and validate the robustness of our method, five-fold cross-validation (CV) scheme is performed on our method. Specifically, the entire plant&#x2019;s PPI dataset will be randomly split into five equal parts; four of them will be employed for training and the remaining one was used for testing. The training and testing data will not overlap with each other to prevent overfitting. The final validation results were the mean value obtained by the five-fold CV scheme. The predictive performance of the proposed approach is verified by five different measurements, including accuracy (Acc), precision (PR), sensitivity (Sens), specificity (Spec), and MCC. They can be represented by</p>
<disp-formula id="S2.Ex9"><label>(7)</label><mml:math id="M9" display="block">
<mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex10"><label>(8)</label><mml:math id="M10" display="block">
<mml:mrow>
<mml:mrow>
<mml:mi>P</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>R</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex11"><label>(9)</label><mml:math id="M11" display="block">
<mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex12"><label>(10)</label><mml:math id="M12" display="block">
<mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>p</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<disp-formula id="S2.Ex13"><label>(11)</label><mml:math id="M13" display="block">
<mml:mrow>
<mml:mrow>
<mml:mi>M</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>-</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mi>F</mml:mi>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x00D7;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>N</mml:mi>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi>F</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>P</mml:mi>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where TP, FP, TN, and FN are associated with the number of true positive, false negative, true negative, and false negative, respectively. In addition, receiver operating characteristic (ROC) curves (<xref ref-type="bibr" rid="B18">Hand, 2009</xref>) were plotted for better accessing the predictive performance of the proposed model. Furthermore, AUC (area under ROC curve) <xref ref-type="bibr" rid="B24">Huang and Ling (2005)</xref> values were also used as an evaluation criterion.</p>
</sec>
<sec id="S3.SS2">
<title>Predictive Performance of Our Model on Three Plant Datasets</title>
<p>We validated the predictive performance of the proposed model on three plant PPI datasets by five-fold CV scheme, including <italic>A. thaliana</italic>, <italic>Z. mays</italic>, and <italic>O. sativa</italic>. It can be observed from <xref ref-type="table" rid="T1">Table 1</xref> that the average accuracy (Acc), precision (PR), sensitivity (Sens), specificity (Spec), and Matthews correlation coefficient (MCC) and AUC values obtained on the <italic>A. thaliana</italic> dataset are 71.48%, 66.64%, 86.09%, 56.88%, 44.94%, and 0.8369, respectively. Their SDs are 0.69, 0.89, 1.08, 2.21, 1.14, and 0.36%, respectively. <xref ref-type="table" rid="T2">Table 2</xref> lists the prediction results obtained on the <italic>Z. mays</italic> dataset, from which we can see the average Acc of 85.41%, PR of 81.54%, Sens of 91.67%, Spec of 79.17%, MCC of 71.43%, and AUC of 0.9466, respectively. Their SDs are 1.18, 2.38, 1.18, 3.24, 2.00, and 0.26%, respectively. On the <italic>O. sativa</italic> dataset, shown in <xref ref-type="table" rid="T3">Table 3</xref>, our model performs at an Acc of 82.60%, PR of 75.79%, Sens of 95.89%, Spec of 69.31%, MCC of 67.65%, and AUC of 0.9442, with SDs of 1.79, 2.43, 0.91, 3.53, 2.98, and 0.58%, respectively. <xref ref-type="fig" rid="F2">Figures 2</xref>&#x2013;<xref ref-type="fig" rid="F4">4</xref> illustrate the ROC curves yielded on <italic>A. thaliana</italic>, <italic>Z. mays</italic>, and <italic>O. sativa</italic> datasets. In the figure of ROC curves, <italic>x</italic>-axis is the false positive rate and <italic>y</italic>-axis represents the true positive rate.</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>Five-fold CV results performed on the <italic>A. thaliana</italic> dataset by the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">71.54</td>
<td valign="top" align="center">66.45</td>
<td valign="top" align="center">87.08</td>
<td valign="top" align="center">55.98</td>
<td valign="top" align="center">45.31</td>
<td valign="top" align="center">0.8415</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">72.05</td>
<td valign="top" align="center">67.73</td>
<td valign="top" align="center">84.64</td>
<td valign="top" align="center">59.36</td>
<td valign="top" align="center">45.49</td>
<td valign="top" align="center">0.8340</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">72.25</td>
<td valign="top" align="center">67.30</td>
<td valign="top" align="center">85.69</td>
<td valign="top" align="center">59.03</td>
<td valign="top" align="center">46.35</td>
<td valign="top" align="center">0.8378</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">70.87</td>
<td valign="top" align="center">66.28</td>
<td valign="top" align="center">85.80</td>
<td valign="top" align="center">55.73</td>
<td valign="top" align="center">43.59</td>
<td valign="top" align="center">0.8325</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">70.71</td>
<td valign="top" align="center">65.46</td>
<td valign="top" align="center">87.25</td>
<td valign="top" align="center">54.30</td>
<td valign="top" align="center">43.98</td>
<td valign="top" align="center">0.8386</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>71.48 &#x00B1; 0.69</bold></td>
<td valign="top" align="center"><bold>66.64 &#x00B1; 0.89</bold></td>
<td valign="top" align="center"><bold>86.09 &#x00B1; 1.08</bold></td>
<td valign="top" align="center"><bold>56.88 &#x00B1; 2.21</bold></td>
<td valign="top" align="center"><bold>44.94 &#x00B1; 1.14</bold></td>
<td valign="top" align="center"><bold>0.8369 &#x00B1; 0.0036</bold></td>
</tr>
</tbody>
</table></table-wrap>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>Five-fold CV results performed on the <italic>Zea mays</italic> dataset by the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">84.63</td>
<td valign="top" align="center">80.07</td>
<td valign="top" align="center">91.80</td>
<td valign="top" align="center">77.59</td>
<td valign="top" align="center">70.04</td>
<td valign="top" align="center">0.9471</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">84.36</td>
<td valign="top" align="center">78.90</td>
<td valign="top" align="center">93.40</td>
<td valign="top" align="center">75.50</td>
<td valign="top" align="center">69.95</td>
<td valign="top" align="center">0.9479</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">85.84</td>
<td valign="top" align="center">83.41</td>
<td valign="top" align="center">90.28</td>
<td valign="top" align="center">81.19</td>
<td valign="top" align="center">71.87</td>
<td valign="top" align="center">0.9421</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">84.94</td>
<td valign="top" align="center">80.73</td>
<td valign="top" align="center">91.95</td>
<td valign="top" align="center">77.89</td>
<td valign="top" align="center">70.56</td>
<td valign="top" align="center">0.9474</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">87.26</td>
<td valign="top" align="center">84.59</td>
<td valign="top" align="center">90.91</td>
<td valign="top" align="center">83.67</td>
<td valign="top" align="center">74.74</td>
<td valign="top" align="center">0.9485</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>85.41 &#x00B1; 1.18</bold></td>
<td valign="top" align="center"><bold>81.54 &#x00B1; 2.38</bold></td>
<td valign="top" align="center"><bold>91.67 &#x00B1; 1.18</bold></td>
<td valign="top" align="center"><bold>79.17 &#x00B1; 3.24</bold></td>
<td valign="top" align="center"><bold>71.43 &#x00B1; 2.00</bold></td>
<td valign="top" align="center"><bold>0.9466 &#x00B1; 0.0026</bold></td>
</tr>
</tbody>
</table></table-wrap>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>Five-fold CV results performed on the <italic>Oryza sativa</italic> dataset by the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">80.21</td>
<td valign="top" align="center">72.29</td>
<td valign="top" align="center">96.03</td>
<td valign="top" align="center">65.28</td>
<td valign="top" align="center">64.03</td>
<td valign="top" align="center">0.9419</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">82.60</td>
<td valign="top" align="center">75.00</td>
<td valign="top" align="center">96.24</td>
<td valign="top" align="center">69.74</td>
<td valign="top" align="center">68.04</td>
<td valign="top" align="center">0.9490</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">85.05</td>
<td valign="top" align="center">78.77</td>
<td valign="top" align="center">96.73</td>
<td valign="top" align="center">72.93</td>
<td valign="top" align="center">71.95</td>
<td valign="top" align="center">0.9503</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">83.33</td>
<td valign="top" align="center">77.17</td>
<td valign="top" align="center">94.33</td>
<td valign="top" align="center">72.49</td>
<td valign="top" align="center">68.40</td>
<td valign="top" align="center">0.9360</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">81.82</td>
<td valign="top" align="center">75.71</td>
<td valign="top" align="center">96.12</td>
<td valign="top" align="center">66.12</td>
<td valign="top" align="center">65.84</td>
<td valign="top" align="center">0.9437</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>82.60 &#x00B1; 1.79</bold></td>
<td valign="top" align="center"><bold>75.79 &#x00B1; 2.43</bold></td>
<td valign="top" align="center"><bold>95.89 &#x00B1; 0.91</bold></td>
<td valign="top" align="center"><bold>69.31 &#x00B1; 3.53</bold></td>
<td valign="top" align="center"><bold>67.65 &#x00B1; 2.98</bold></td>
<td valign="top" align="center"><bold>0.9440 &#x00B1; 0.0058</bold></td>
</tr>
</tbody>
</table></table-wrap>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>The ROC curves of our approach on the <italic>A. thaliana</italic> dataset under five-fold CV.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>The ROC curves of our approach on the <italic>Zea mays</italic> dataset under five-fold CV.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>The ROC curves of our approach on the <italic>Oryza sativa</italic> dataset under five-fold CV.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g004.tif"/>
</fig>
<p>Based on the experimental results, it can be indicated that the proposed model is effective for identifying PPIs in plants. We attributed this better prediction performance to the powerful DHT&#x2013;SVD descriptors and the excellent DNN classifier. The PSSM not only encodes the sequence into matrix but also obtains the sufficient prior information of plant proteins. In addition, the application of DHT extracted robust feature descriptors from PSSM, and then, SVD algorithm was employed to reduce the noise and decrease the dimension of feature matrix that can better improve the prediction performance. As a popular deep learning classifier, DNN shows the powerful ability for training and predicting, which makes us more convinced that our method can be a useful tool for plant PPI prediction.</p>
</sec>
<sec id="S3.SS3">
<title>Comparison With Random Forest and K-Nearest Neighbor Classifier</title>
<p>There are many machine learning classifiers that have been applied to predict PPIs. K-nearest neighbor (KNN) (<xref ref-type="bibr" rid="B26">Keller et al., 1985</xref>) and random forest (RF) (<xref ref-type="bibr" rid="B6">Breiman, 2001</xref>) are the most widely used algorithms. The KNN algorithm is one of the simplest classification approaches and it has been widely applied to detect PPIs (<xref ref-type="bibr" rid="B33">Li et al., 2009</xref>). RF is a decision tree&#x2013;based ensemble learning method, and it is known for its powerful ability of classification (<xref ref-type="bibr" rid="B32">Li et al., 2012</xref>). To further verify the predictive ability of DNN classifier, we compared it with the KNN and RF model by the five-fold CV scheme and adopted the same DHT feature descriptors. The results list in <xref ref-type="table" rid="T4">Table 4</xref> illustrates that our method achieved higher AUC values across the <italic>A. thaliana</italic>, <italic>Z. mays</italic>, and <italic>O. sativa</italic> datasets. It can be observed that the average AUC values of the DNN classifier are 0.1023, 0.1215, and 0.1354 higher than those of KNN classifier. Similarly, when compared with the RF classifier, the AUC value of our model improved 0.0036, 0.013, and 0.0241, respectively. From the comparison results shown in <xref ref-type="fig" rid="F5">Figure 5</xref>, we considered that the combination of DNN classifier and DHT descriptors can significantly improve the performance in plant PPI prediction.</p>
<table-wrap position="float" id="T4">
<label>TABLE 4</label>
<caption><p>Five-fold CV results yielded by KNN and RF classifier on the three plant PPI datasets.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Dataset</td>
<td valign="top" align="center">Classifier</td>
<td valign="top" align="center">AUC</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>A. thaliana</italic></td>
<td valign="top" align="center">KNN</td>
<td valign="top" align="center">0.7346 &#x00B1; 0.22</td>
<td valign="top" align="center">71.12 &#x00B1; 0.44</td>
<td valign="top" align="center">79.00 &#x00B1; 0.54</td>
<td valign="top" align="center">67.92 &#x00B1; 0.43</td>
<td valign="top" align="center">60.77 &#x00B1; 0.22</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">RF</td>
<td valign="top" align="center">0.8333 &#x00B1; 0.77</td>
<td valign="top" align="center">82.63 &#x00B1; 0.94</td>
<td valign="top" align="center">68.31 &#x00B1; 1.23</td>
<td valign="top" align="center">85.63 &#x00B1; 0.88</td>
<td valign="top" align="center">64.01 &#x00B1; 0.69</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Our method</td>
<td valign="top" align="center">0.8369 &#x00B1; 0.36</td>
<td valign="top" align="center">66.64 &#x00B1; 0.89</td>
<td valign="top" align="center">86.09 &#x00B1; 1.08</td>
<td valign="top" align="center">56.88 &#x00B1; 2.21</td>
<td valign="top" align="center">44.94 &#x00B1; 1.14</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Zea mays</italic></td>
<td valign="top" align="center">KNN</td>
<td valign="top" align="center">0.8251 &#x00B1; 0.42</td>
<td valign="top" align="center">78.38 &#x00B1; 0.77</td>
<td valign="top" align="center">89.77 &#x00B1; 0.48</td>
<td valign="top" align="center">75.25 &#x00B1; 0.77</td>
<td valign="top" align="center">70.83 &#x00B1; 0.57</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">RF</td>
<td valign="top" align="center">0.9336 &#x00B1; 0.40</td>
<td valign="top" align="center">96.98 &#x00B1; 0.28</td>
<td valign="top" align="center">89.52 &#x00B1; 0.48</td>
<td valign="top" align="center">97.21 &#x00B1; 0.34</td>
<td valign="top" align="center">87.57 &#x00B1; 0.49</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Our method</td>
<td valign="top" align="center">0.9466 &#x00B1; 0.26</td>
<td valign="top" align="center">81.54 &#x00B1; 2.38</td>
<td valign="top" align="center">91.67 &#x00B1; 1.18</td>
<td valign="top" align="center">79.17 &#x00B1; 3.24</td>
<td valign="top" align="center">71.43 &#x00B1; 2.00</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Oryza sativa</italic></td>
<td valign="top" align="center">KNN</td>
<td valign="top" align="center">0.8086 &#x00B1; 0.89</td>
<td valign="top" align="center">76.41 &#x00B1; 1.55</td>
<td valign="top" align="center">89.28 &#x00B1; 0.78</td>
<td valign="top" align="center">72.44 &#x00B1; 1.58</td>
<td valign="top" align="center">68.59 &#x00B1; 1.17</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">RF</td>
<td valign="top" align="center">0.9199 &#x00B1; 0.58</td>
<td valign="top" align="center">87.30 &#x00B1; 1.35</td>
<td valign="top" align="center">88.00 &#x00B1; 1.34</td>
<td valign="top" align="center">87.22 &#x00B1; 1.16</td>
<td valign="top" align="center">78.26 &#x00B1; 1.28</td>
</tr>
<tr>
<td/>
<td valign="top" align="center">Our method</td>
<td valign="top" align="center">0.9440 &#x00B1; 0.58</td>
<td valign="top" align="center">75.79 &#x00B1; 2.43</td>
<td valign="top" align="center">95.89 &#x00B1; 0.91</td>
<td valign="top" align="center">69.31 &#x00B1; 3.53</td>
<td valign="top" align="center">67.65 &#x00B1; 2.98</td>
</tr>
</tbody>
</table></table-wrap>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Comparison results of AUC values obtained by deep neural network (DNN), K-nearest neighbor (KNN), and random forest (RF) classifiers on the three plant PPI datasets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g005.tif"/>
</fig>
</sec>
<sec id="S3.SS4">
<title>Comparison of Position-Specific Scoring Matrix With Different Protein Representation Methods</title>
<p>To evaluate the performance of PSSM, we compared it with the substitution matrix representation (SMR), which was proposed by <xref ref-type="bibr" rid="B50">Yu et al. (2012)</xref> to represent protein sequence. In this section, we employed the BLOSUM62 matrix to encode the <italic>A. thaliana</italic> protein sequence as a 20 &#x00D7; 20 matrix. Then, the DHT algorithm was applied to extract feature descriptors from SMR matrix and SVD was also adopted to reduce the feature dimensions. By this way, we can generate a 600-dimensional SMR&#x2013;DHT descriptor for each protein pair. The five-fold CV results of SMR&#x2013;DHT descriptors combined with DNN classifier on the <italic>A. thaliana</italic> dataset are summarized in <xref ref-type="table" rid="T5">Table 5</xref>. It can be observed that the PSSM-based method performs significantly better than the SMR-based method. For example, the accuracy and AUC gaps between PSSM and SMR-based method are 4.38 and 4.94%, respectively. The higher predictive accuracy and lower SDs further indicated that our method performs better than the SMR-based approach (<xref ref-type="fig" rid="F6">Figure 6</xref>).</p>
<table-wrap position="float" id="T5">
<label>TABLE 5</label>
<caption><p>Comparison of PSSM with SMR-based method on the <italic>A. thaliana</italic> dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">71.54</td>
<td valign="top" align="center">71.26</td>
<td valign="top" align="center">72.27</td>
<td valign="top" align="center">70.81</td>
<td valign="top" align="center">43.08</td>
<td valign="top" align="center">78.72</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">61.05</td>
<td valign="top" align="center">57.03</td>
<td valign="top" align="center">90.82</td>
<td valign="top" align="center">31.04</td>
<td valign="top" align="center">27.29</td>
<td valign="top" align="center">79.00</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">58.44</td>
<td valign="top" align="center">54.79</td>
<td valign="top" align="center">92.68</td>
<td valign="top" align="center">24.74</td>
<td valign="top" align="center">23.70</td>
<td valign="top" align="center">78.43</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">72.47</td>
<td valign="top" align="center">74.73</td>
<td valign="top" align="center">68.51</td>
<td valign="top" align="center">76.50</td>
<td valign="top" align="center">45.14</td>
<td valign="top" align="center">78.66</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">72.02</td>
<td valign="top" align="center">71.34</td>
<td valign="top" align="center">73.25</td>
<td valign="top" align="center">70.80</td>
<td valign="top" align="center">44.06</td>
<td valign="top" align="center">78.94</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>67.10 &#x00B1; 6.79</bold></td>
<td valign="top" align="center"><bold>65.83 &#x00B1; 9.2</bold></td>
<td valign="top" align="center"><bold>79.51 &#x00B1; 11.34</bold></td>
<td valign="top" align="center"><bold>54.78 &#x00B1; 24.76</bold></td>
<td valign="top" align="center"><bold>36.65 &#x00B1; 10.29</bold></td>
<td valign="top" align="center"><bold>0.7875 &#x00B1; 0.0023</bold></td>
</tr>
<tr>
<td valign="top" align="left"><bold>Our method</bold></td>
<td valign="top" align="center"><bold>71.48 &#x00B1; 0.69</bold></td>
<td valign="top" align="center"><bold>66.64 &#x00B1; 0.89</bold></td>
<td valign="top" align="center"><bold>86.09 &#x00B1; 1.08</bold></td>
<td valign="top" align="center"><bold>56.88 &#x00B1; 2.21</bold></td>
<td valign="top" align="center"><bold>44.94 &#x00B1; 1.14</bold></td>
<td valign="top" align="center"><bold>0.8369 &#x00B1; 0.0036</bold></td>
</tr>
</tbody>
</table></table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>ROC curves obtained from SMR-based method on the <italic>A. thaliana</italic> dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g006.tif"/>
</fig>
</sec>
<sec id="S3.SS5">
<title>Comparison With Different Feature Extraction Methods</title>
<p>To illustrate the effectivity of our feature extraction approach, we compared DHT with some popular correlative methods, including discrete cosine transform (DCT) (<xref ref-type="bibr" rid="B1">Ahmed et al., 1974</xref>), fast Fourier transform (FFT) (<xref ref-type="bibr" rid="B35">Nussbaumer, 1981</xref>), discrete wavelet transform (DWT) (<xref ref-type="bibr" rid="B34">Nanni et al., 2012</xref>), and auto-covariance (AC) (<xref ref-type="bibr" rid="B53">Zeng et al., 2009</xref>). As shown in <xref ref-type="table" rid="T6">Table 6</xref> and <xref ref-type="fig" rid="F7">Figure 7</xref>, on the <italic>O. sativa</italic> dataset, our method obtained a high prediction accuracy of 82.60%. The prediction accuracy values of other methods are 80.95, 75.31, 81.54, and 66.63%, respectively. Our method performs better than the other four methods. Especially compared with the AC-based method, our approach improved the Acc, Spec, MCC, and AUC by over 15%, and PR and Sens by over 7%, respectively. Although the Sens value of our method is not the highest, it still obtains an excellent value of 95.89%. The Acc, PR, Sens, Spec, MCC, and AUC values obtained from our model are 1.06, 0.69, 1.08, 1.05, 2.15, and 1.31% higher than the values of the DWT-based method. These comparison results further indicated the superiority of the proposed method.</p>
<table-wrap position="float" id="T6">
<label>TABLE 6</label>
<caption><p>Performance comparison of the DHT with different feature extraction methods on <italic>Oryza sativa</italic> dataset.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Descriptors</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">DCT+DNN</td>
<td valign="top" align="center">80.95 &#x00B1; 1.10</td>
<td valign="top" align="center">73.70 &#x00B1; 1.41</td>
<td valign="top" align="center"><bold>96.12 &#x00B1; 1.15</bold></td>
<td valign="top" align="center">65.64 &#x00B1; 2.40</td>
<td valign="top" align="center">64.99 &#x00B1; 1.97</td>
<td valign="top" align="center">0.9360 &#x00B1; 0.0017</td>
</tr>
<tr>
<td valign="top" align="left">FFT+DNN</td>
<td valign="top" align="center">75.31 &#x00B1; 1.37</td>
<td valign="top" align="center">68.61 &#x00B1; 1.03</td>
<td valign="top" align="center">93.34 &#x00B1; 1.59</td>
<td valign="top" align="center">57.23 &#x00B1; 2.90</td>
<td valign="top" align="center">54.26 &#x00B1; 2.81</td>
<td valign="top" align="center">0.8760 &#x00B1; 0.0096</td>
</tr>
<tr>
<td valign="top" align="left">DWT+DNN</td>
<td valign="top" align="center">81.54 &#x00B1; 3.05</td>
<td valign="top" align="center">75.10 &#x00B1; 3.84</td>
<td valign="top" align="center">94.81 &#x00B1; 0.65</td>
<td valign="top" align="center">68.26 &#x00B1; 6.61</td>
<td valign="top" align="center">65.50 &#x00B1; 4.99</td>
<td valign="top" align="center">0.9309 &#x00B1; 0.0052</td>
</tr>
<tr>
<td valign="top" align="left">AC+DNN</td>
<td valign="top" align="center">66.63 &#x00B1; 4.48</td>
<td valign="top" align="center">62.02 &#x00B1; 4.91</td>
<td valign="top" align="center">88.42 &#x00B1; 4.77</td>
<td valign="top" align="center">45.02 &#x00B1; 12.49</td>
<td valign="top" align="center">37.39 &#x00B1; 5.39</td>
<td valign="top" align="center">0.7931 &#x00B1; 0.0126</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Our method</bold></td>
<td valign="top" align="center"><bold>82.60 &#x00B1; 1.79</bold></td>
<td valign="top" align="center"><bold>75.79 &#x00B1; 2.43</bold></td>
<td valign="top" align="center">95.89 &#x00B1; 0.91</td>
<td valign="top" align="center"><bold>69.31 &#x00B1; 3.53</bold></td>
<td valign="top" align="center"><bold>67.65 &#x00B1; 2.98</bold></td>
<td valign="top" align="center"><bold>0.9440 &#x00B1; 0.0058</bold></td>
</tr>
</tbody>
</table></table-wrap>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Five-fold CV results obtained by DNN classifier with different feature descriptors on the <italic>Oryza sativa</italic> dataset. <bold>(A)</bold> is the ROC curves obtained by DCT descriptors; <bold>(B)</bold> is the ROC curves obtained by FFT descriptors; <bold>(C)</bold> is the ROC curves obtained by DWT; <bold>(D)</bold> is the ROC curves obtained by AC.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g007.tif"/>
</fig>
</sec>
<sec id="S3.SS6">
<title>Predictive Ability on Yeast and Human Dataset</title>
<p>To further validate the potential of the presented method, we performed it on the yeast and human PPI dataset, which was introduced by <xref ref-type="bibr" rid="B17">Guo et al. (2008)</xref> and <xref ref-type="bibr" rid="B25">Huang et al. (2015)</xref>. The predictive results of the two datasets are listed in <xref ref-type="table" rid="T7">Tables 7</xref>, <xref ref-type="table" rid="T8">8</xref>, and the corresponding ROC curves are shown in <xref ref-type="fig" rid="F8">Figures 8</xref>, <xref ref-type="fig" rid="F9">9</xref>. When performing on the yeast dataset, it achieved average Acc, PR, Sens, Spec, MCC, and AUC value of 79.54%, 73.46%, 92.63%, 66.47%, 61.27%, and 0.9203, with SDs of 1.43, 2.11, 1.18, 3.60, 2.16, and 0.46%, respectively. From <xref ref-type="table" rid="T8">Table 8</xref>, it can be observed that the proposed model yielded great results on the human dataset, an average Acc of 82.76%, PR of 75.79%, Sens of 94.18%, Spec of 72.30%, MCC of 67.72%, and AUC of 0.9473, with SDs of 1.68, 2.79, 1.64, 4.74, 2.37, and 0.26%, respectively. From these results, we can observe that the powerful DNN-based classifier combined with the DHT feature descriptor is accurate and robust for exploring cross-species predictions of PPIs.</p>
<table-wrap position="float" id="T7">
<label>TABLE 7</label>
<caption><p>Five-fold CV results performed on the yeast dataset by the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">77.20</td>
<td valign="top" align="center">70.38</td>
<td valign="top" align="center">93.33</td>
<td valign="top" align="center">61.31</td>
<td valign="top" align="center">57.60</td>
<td valign="top" align="center">0.9176</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">79.88</td>
<td valign="top" align="center">73.51</td>
<td valign="top" align="center">91.84</td>
<td valign="top" align="center">68.50</td>
<td valign="top" align="center">61.82</td>
<td valign="top" align="center">0.9241</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">79.44</td>
<td valign="top" align="center">73.17</td>
<td valign="top" align="center">93.73</td>
<td valign="top" align="center">64.80</td>
<td valign="top" align="center">61.27</td>
<td valign="top" align="center">0.9181</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">80.20</td>
<td valign="top" align="center">73.97</td>
<td valign="top" align="center">93.31</td>
<td valign="top" align="center">67.03</td>
<td valign="top" align="center">62.56</td>
<td valign="top" align="center">0.9263</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">81.00</td>
<td valign="top" align="center">76.27</td>
<td valign="top" align="center">90.95</td>
<td valign="top" align="center">70.70</td>
<td valign="top" align="center">63.09</td>
<td valign="top" align="center">0.9158</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>79.54 &#x00B1; 1.43</bold></td>
<td valign="top" align="center"><bold>73.46 &#x00B1; 2.11</bold></td>
<td valign="top" align="center"><bold>92.63 &#x00B1; 1.18</bold></td>
<td valign="top" align="center"><bold>66.47 &#x00B1; 3.60</bold></td>
<td valign="top" align="center"><bold>61.27 &#x00B1; 2.16</bold></td>
<td valign="top" align="center"><bold>0.9203 &#x00B1; 0.0046</bold></td>
</tr>
</tbody>
</table></table-wrap>
<table-wrap position="float" id="T8">
<label>TABLE 8</label>
<caption><p>Five-fold CV results performed on the human dataset by the proposed model.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Testing set</td>
<td valign="top" align="center">Acc (%)</td>
<td valign="top" align="center">PR (%)</td>
<td valign="top" align="center">Sens (%)</td>
<td valign="top" align="center">Spec (%)</td>
<td valign="top" align="center">MCC (%)</td>
<td valign="top" align="center">AUC</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">1</td>
<td valign="top" align="center">82.41</td>
<td valign="top" align="center">74.69</td>
<td valign="top" align="center">94.07</td>
<td valign="top" align="center">72.28</td>
<td valign="top" align="center">67.18</td>
<td valign="top" align="center">0.9487</td>
</tr>
<tr>
<td valign="top" align="left">2</td>
<td valign="top" align="center">82.05</td>
<td valign="top" align="center">74.09</td>
<td valign="top" align="center">95.19</td>
<td valign="top" align="center">70.34</td>
<td valign="top" align="center">66.92</td>
<td valign="top" align="center">0.9484</td>
</tr>
<tr>
<td valign="top" align="left">3</td>
<td valign="top" align="center">83.76</td>
<td valign="top" align="center">78.61</td>
<td valign="top" align="center">92.33</td>
<td valign="top" align="center">75.36</td>
<td valign="top" align="center">68.60</td>
<td valign="top" align="center">0.9428</td>
</tr>
<tr>
<td valign="top" align="left">4</td>
<td valign="top" align="center">84.99</td>
<td valign="top" align="center">78.87</td>
<td valign="top" align="center">92.96</td>
<td valign="top" align="center">77.92</td>
<td valign="top" align="center">71.17</td>
<td valign="top" align="center">0.9492</td>
</tr>
<tr>
<td valign="top" align="left">5</td>
<td valign="top" align="center">80.59</td>
<td valign="top" align="center">72.70</td>
<td valign="top" align="center">96.36</td>
<td valign="top" align="center">65.59</td>
<td valign="top" align="center">64.75</td>
<td valign="top" align="center">0.9481</td>
</tr>
<tr>
<td valign="top" align="left"><bold>Average</bold></td>
<td valign="top" align="center"><bold>82.76 &#x00B1; 1.68</bold></td>
<td valign="top" align="center"><bold>75.79 &#x00B1; 2.79</bold></td>
<td valign="top" align="center"><bold>94.18 &#x00B1; 1.64</bold></td>
<td valign="top" align="center"><bold>72.30 &#x00B1; 4.74</bold></td>
<td valign="top" align="center"><bold>67.72 &#x00B1; 2.37</bold></td>
<td valign="top" align="center"><bold>0.9473 &#x00B1; 0.0026</bold></td>
</tr>
</tbody>
</table></table-wrap>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption><p>ROC curves performed by the proposed model on yeast dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g008.tif"/>
</fig>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption><p>ROC curves performed by the proposed model on human dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fgene-12-745228-g009.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="S4">
<title>Discussion</title>
<p>In this article, we proposed a deep learning framework to predict PPIs in plants only using the information of amino acid sequence. This approach is based on DNN combined with DHT descriptors and PSSM. More specifically, we first used the PSSM to represent plant protein sequences, and then extracted feature vectors from these matrices by DHT. To improve the prediction accuracy and reduce the computational complexity, the SVD algorithm was adopted to reduce the feature dimensions. Lastly, these feature descriptors were sent to the DNN classifier for training and predicting. To verify the performance of the proposed approach, we performed it on <italic>A. thaliana</italic>, <italic>Z. mays</italic>, and <italic>O. sativa</italic> datasets. To evaluate the power of the DNN-based classifier, we compared it with the KNN and RF classifier using the same DHT descriptors. In addition, we also compared the DHT with some different feature descriptors. To further indicate the generality of our model, we also applied it to the yeast and human datasets. The experimental results indicated that our model performs significantly well in predicting PPIs in plants. In further work, we will continue to design more effective computational models for better analyzing biomolecular interactions in plants.</p>
</sec>
<sec sec-type="data-availability" id="S5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="http://arabidopsis.org/">http://arabidopsis.org/</ext-link>; <ext-link ext-link-type="uri" xlink:href="http://www.ebi.ac.uk/intact">http://www.ebi.ac.uk/intact</ext-link>; <ext-link ext-link-type="uri" xlink:href="http://www.thebiogrid.org/">http://www.thebiogrid.org/</ext-link>; <ext-link ext-link-type="uri" xlink:href="http://comp-sysbio.org/ppim">http://comp-sysbio.org/ppim</ext-link>; <ext-link ext-link-type="uri" xlink:href="http://bis.zju.edu.cn/prin/">http://bis.zju.edu.cn/prin/</ext-link>.</p>
</sec>
<sec id="S6">
<title>Author Contributions</title>
<p>JP, L-PL, and Z-HY: conceptualization, methodology, software, validation, formal analysis, investigation, resources, and data curation. C-QY and Z-HR: writing &#x2013; original draft preparation, writing, review, editing, visualization, and supervision. Y-JG: project administration. Z-HY: funding acquisition. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The reviewer Z-AH declared past co-authorships with one of the author Z-HY and the reviewer HY-a declared past co-authorships with two of the authors Z-HY and C-QY to the handling Editor.</p>
</sec>
<sec sec-type="disclaimer" id="S7">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec sec-type="funding-information" id="S8">
<title>Funding</title>
<p>This research was funded by the National Natural Science Foundation of China, grant numbers 62002297 and 61722212.</p>
</sec>
<ack>
<p>Our deepest gratitude goes to the editor Robert Friedman and three reviewers for their careful work and thoughtful suggestions that have helped improve this manuscript substantially.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ahmed</surname> <given-names>N.</given-names></name> <name><surname>Natarajan</surname> <given-names>T.</given-names></name> <name><surname>Rao</surname> <given-names>K. R.</given-names></name></person-group> (<year>1974</year>). <article-title>Discrete cosine transform.</article-title> <source><italic>IEEE Trans. Comput.</italic></source> <volume>100</volume> <fpage>90</fpage>&#x2013;<lpage>93</lpage>. <pub-id pub-id-type="doi">10.1109/T-C.1974.223784</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Koonin</surname> <given-names>E. V.</given-names></name></person-group> (<year>1998</year>). <article-title>Iterated profile searches with PSI-BLAST&#x2014;a tool for discovery in protein databases.</article-title> <source><italic>Trends Biochem. Sci.</italic></source> <volume>23</volume> <fpage>444</fpage>&#x2013;<lpage>447</lpage>. <pub-id pub-id-type="doi">10.1016/s0968-0004(98)01298-5</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Armean</surname> <given-names>I. M.</given-names></name> <name><surname>Lilley</surname> <given-names>K. S.</given-names></name> <name><surname>Trotter</surname> <given-names>M. W.</given-names></name></person-group> (<year>2013</year>). <article-title>Popular computational methods to assess multiprotein complexes derived from label-free affinity purification and mass spectrometry (AP-MS) experiments.</article-title> <source><italic>Mol. Cell. Proteomics</italic></source> <volume>12</volume> <fpage>1</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1074/mcp.r112.019554</pub-id> <pub-id pub-id-type="pmid">23071097</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bracewell</surname> <given-names>R. N.</given-names></name> <name><surname>Bracewell</surname> <given-names>R. N.</given-names></name></person-group> (<year>1986</year>). <source><italic>The Fourier Transform And Its Applications.</italic></source> <publisher-loc>New York</publisher-loc>: <publisher-name>McGraw-Hill</publisher-name>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bracha-Drori</surname> <given-names>K.</given-names></name> <name><surname>Shichrur</surname> <given-names>K.</given-names></name> <name><surname>Katz</surname> <given-names>A.</given-names></name> <name><surname>Oliva</surname> <given-names>M.</given-names></name> <name><surname>Angelovici</surname> <given-names>R.</given-names></name> <name><surname>Yalovsky</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2004</year>). <article-title>Detection of protein&#x2013;protein interactions in plants using bimolecular fluorescence complementation.</article-title> <source><italic>Plant J.</italic></source> <volume>40</volume> <fpage>419</fpage>&#x2013;<lpage>427</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-313X.2004.02206.x</pub-id> <pub-id pub-id-type="pmid">15469499</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Breiman</surname> <given-names>L.</given-names></name></person-group> (<year>2001</year>). <article-title>Random forests.</article-title> <source><italic>Mach. Learn.</italic></source> <volume>45</volume> <fpage>5</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1023/A:1010933404324</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Canovas</surname> <given-names>F. M.</given-names></name> <name><surname>Dumas-Gaudot</surname> <given-names>E.</given-names></name> <name><surname>Recorbet</surname> <given-names>G.</given-names></name> <name><surname>Jorrin</surname> <given-names>J.</given-names></name> <name><surname>Mock</surname> <given-names>H. P.</given-names></name> <name><surname>Rossignol</surname> <given-names>M.</given-names></name></person-group> (<year>2004</year>). <article-title>Plant proteome analysis.</article-title> <source><italic>Proteomics</italic></source> <volume>4</volume> <fpage>285</fpage>&#x2013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1002/pmic.200300602</pub-id> <pub-id pub-id-type="pmid">14760698</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Causier</surname> <given-names>B.</given-names></name> <name><surname>Davies</surname> <given-names>B.</given-names></name></person-group> (<year>2002</year>). <article-title>Analysing protein-protein interactions with the yeast two-hybrid system.</article-title> <source><italic>Plant Mol. Biol.</italic></source> <volume>50</volume> <fpage>855</fpage>&#x2013;<lpage>870</lpage>. <pub-id pub-id-type="doi">10.1023/A:1021214007897</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>M.</given-names></name> <name><surname>Ju</surname> <given-names>C. J.-T.</given-names></name> <name><surname>Zhou</surname> <given-names>G.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Zhang</surname> <given-names>T.</given-names></name> <name><surname>Chang</surname> <given-names>K.-W.</given-names></name><etal/></person-group> (<year>2019</year>). <article-title>Multifaceted protein&#x2013;protein interaction prediction based on siamese residual rcnn.</article-title> <source><italic>Bioinformatics</italic></source> <volume>35</volume> <fpage>i305</fpage>&#x2013;<lpage>i314</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz328</pub-id> <pub-id pub-id-type="pmid">31510705</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cizek</surname> <given-names>V.</given-names></name></person-group> (<year>1970</year>). <article-title>Discrete hilbert transform.</article-title> <source><italic>IEEE Tran. Audio Electroacoustics</italic></source> <volume>18</volume> <fpage>340</fpage>&#x2013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1109/TAU.1970.1162139</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Davies</surname> <given-names>M. N.</given-names></name> <name><surname>Secker</surname> <given-names>A.</given-names></name> <name><surname>Freitas</surname> <given-names>A. A.</given-names></name> <name><surname>Clark</surname> <given-names>E.</given-names></name> <name><surname>Timmis</surname> <given-names>J.</given-names></name> <name><surname>Flower</surname> <given-names>D. R.</given-names></name></person-group> (<year>2008</year>). <article-title>Optimizing amino acid groupings for GPCR classification.</article-title> <source><italic>Bioinformatics</italic></source> <volume>24</volume> <fpage>1980</fpage>&#x2013;<lpage>1986</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btn382</pub-id> <pub-id pub-id-type="pmid">18676973</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Du</surname> <given-names>X.</given-names></name> <name><surname>Sun</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>C.</given-names></name> <name><surname>Yao</surname> <given-names>Y.</given-names></name> <name><surname>Yan</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name></person-group> (<year>2017</year>). <article-title>DeepPPI: boosting prediction of protein&#x2013;protein interactions with deep neural networks.</article-title> <source><italic>J. Chem. Inform. Model.</italic></source> <volume>57</volume> <fpage>1499</fpage>&#x2013;<lpage>1510</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jcim.7b00028</pub-id> <pub-id pub-id-type="pmid">28514151</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fang</surname> <given-names>Y.</given-names></name> <name><surname>Macool</surname> <given-names>D.</given-names></name> <name><surname>Xue</surname> <given-names>Z.</given-names></name> <name><surname>Heppard</surname> <given-names>E.</given-names></name> <name><surname>Hainey</surname> <given-names>C.</given-names></name> <name><surname>Tingey</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2002</year>). <article-title>Development of a high-throughput yeast two-hybrid screening system to study protein-protein interactions in plants.</article-title> <source><italic>Mol. Genet. Genomics</italic></source> <volume>267</volume> <fpage>142</fpage>&#x2013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1007/s00438-002-0656-7</pub-id> <pub-id pub-id-type="pmid">11976957</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fukao</surname> <given-names>Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Protein&#x2013;protein interactions in plants.</article-title> <source><italic>Plant Cell Physiol.</italic></source> <volume>53</volume> <fpage>617</fpage>&#x2013;<lpage>625</lpage>. <pub-id pub-id-type="doi">10.1093/pcp/pcs026</pub-id> <pub-id pub-id-type="pmid">22383626</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gribskov</surname> <given-names>M.</given-names></name> <name><surname>Mclachlan</surname> <given-names>A. D.</given-names></name> <name><surname>Eisenberg</surname> <given-names>D.</given-names></name></person-group> (<year>1987</year>). <article-title>Profile analysis: detection of distantly related proteins.</article-title> <source><italic>Proc. Natl. Acad. Sci. U. S. A.</italic></source> <volume>84</volume> <fpage>4355</fpage>&#x2013;<lpage>4358</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.84.13.4355</pub-id> <pub-id pub-id-type="pmid">3474607</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>H.</given-names></name> <name><surname>Zhu</surname> <given-names>P.</given-names></name> <name><surname>Jiao</surname> <given-names>Y.</given-names></name> <name><surname>Meng</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>PRIN: a predicted rice interactome network.</article-title> <source><italic>BMC Bioinformatics</italic></source> <volume>12</volume>:<fpage>161</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-12-161</pub-id> <pub-id pub-id-type="pmid">21575196</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>Y.</given-names></name> <name><surname>Yu</surname> <given-names>L.</given-names></name> <name><surname>Wen</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Using support vector machine combined with auto covariance to predict protein&#x2013;protein interactions from protein sequences.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>36</volume> <fpage>3025</fpage>&#x2013;<lpage>3030</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkn159</pub-id> <pub-id pub-id-type="pmid">18390576</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hand</surname> <given-names>D. J.</given-names></name></person-group> (<year>2009</year>). <article-title>Measuring classifier performance: a coherent alternative to the area under the ROC curve.</article-title> <source><italic>Mach. Learn.</italic></source> <volume>77</volume> <fpage>103</fpage>&#x2013;<lpage>123</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-009-5119-5</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hashemifar</surname> <given-names>S.</given-names></name> <name><surname>Neyshabur</surname> <given-names>B.</given-names></name> <name><surname>Khan</surname> <given-names>A. A.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name></person-group> (<year>2018</year>). <article-title>Predicting protein&#x2013;protein interactions through sequence-based deep learning.</article-title> <source><italic>Bioinformatics</italic></source> <volume>34</volume> <fpage>i802</fpage>&#x2013;<lpage>i810</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty573</pub-id> <pub-id pub-id-type="pmid">30423091</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hayashi</surname> <given-names>T.</given-names></name> <name><surname>Matsuzaki</surname> <given-names>Y.</given-names></name> <name><surname>Yanagisawa</surname> <given-names>K.</given-names></name> <name><surname>Ohue</surname> <given-names>M.</given-names></name> <name><surname>Akiyama</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>MEGADOCK-Web: an integrated database of high-throughput structure-based protein-protein interaction predictions.</article-title> <source><italic>BMC Bioinformatics</italic></source> <volume>19</volume>:<fpage>62</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-018-2073-x</pub-id> <pub-id pub-id-type="pmid">29745830</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G.</given-names></name> <name><surname>Vinyals</surname> <given-names>O.</given-names></name> <name><surname>Dean</surname> <given-names>J.</given-names></name></person-group> (<year>2015</year>). <article-title>Distilling the knowledge in a neural network.</article-title> <source><italic>arXiv</italic> [Preprint]</source>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1503.02531">arXiv:1503.02531</ext-link></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Osindero</surname> <given-names>S.</given-names></name> <name><surname>Teh</surname> <given-names>Y.-W.</given-names></name></person-group> (<year>2006</year>). <article-title>A fast learning algorithm for deep belief nets.</article-title> <source><italic>Neural Comput.</italic></source> <volume>18</volume> <fpage>1527</fpage>&#x2013;<lpage>1554</lpage>. <pub-id pub-id-type="doi">10.1162/neco.2006.18.7.1527</pub-id> <pub-id pub-id-type="pmid">16764513</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hinton</surname> <given-names>G. E.</given-names></name> <name><surname>Salakhutdinov</surname> <given-names>R. R.</given-names></name></person-group> (<year>2006</year>). <article-title>Reducing the dimensionality of data with neural networks.</article-title> <source><italic>Science</italic></source> <volume>313</volume> <fpage>504</fpage>&#x2013;<lpage>507</lpage>. <pub-id pub-id-type="doi">10.1126/science.1127647</pub-id> <pub-id pub-id-type="pmid">16873662</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>J.</given-names></name> <name><surname>Ling</surname> <given-names>C. X.</given-names></name></person-group> (<year>2005</year>). <article-title>Using AUC and accuracy in evaluating learning algorithms.</article-title> <source><italic>IEEE Trans. Knowl. Data Eng.</italic></source> <volume>17</volume> <fpage>299</fpage>&#x2013;<lpage>310</lpage>. <pub-id pub-id-type="doi">10.1109/tkde.2005.50</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>Y.-A.</given-names></name> <name><surname>You</surname> <given-names>Z.-H.</given-names></name> <name><surname>Gao</surname> <given-names>X.</given-names></name> <name><surname>Wong</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name></person-group> (<year>2015</year>). <article-title>Using weighted sparse representation model combined with discrete cosine transformation to predict protein-protein interactions from protein sequence.</article-title> <source><italic>BioMed Res. Int.</italic></source> <volume>2015</volume> <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <pub-id pub-id-type="doi">10.1155/2015/902198</pub-id> <pub-id pub-id-type="pmid">26634213</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Keller</surname> <given-names>J. M.</given-names></name> <name><surname>Gray</surname> <given-names>M. R.</given-names></name> <name><surname>Givens</surname> <given-names>J. A.</given-names></name></person-group> (<year>1985</year>). <article-title>A fuzzy k-nearest neighbor algorithm.</article-title> <source><italic>IEEE Trans. Syst. Man Cybern.</italic></source> <volume>15</volume> <fpage>580</fpage>&#x2013;<lpage>585</lpage>. <pub-id pub-id-type="doi">10.1109/TSMC.1985.6313426</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kerrien</surname> <given-names>S.</given-names></name> <name><surname>Aranda</surname> <given-names>B.</given-names></name> <name><surname>Breuza</surname> <given-names>L.</given-names></name> <name><surname>Bridge</surname> <given-names>A.</given-names></name> <name><surname>Broackes-Carter</surname> <given-names>F.</given-names></name> <name><surname>Chen</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>The IntAct molecular interaction database in 2012.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>40</volume> <fpage>D841</fpage>&#x2013;<lpage>D846</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr1088</pub-id> <pub-id pub-id-type="pmid">22121220</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khan</surname> <given-names>I. K.</given-names></name> <name><surname>Kihara</surname> <given-names>D.</given-names></name></person-group> (<year>2016</year>). <article-title>Genome-scale prediction of moonlighting proteins using diverse protein association information.</article-title> <source><italic>Bioinformatics</italic></source> <volume>32</volume> <fpage>2281</fpage>&#x2013;<lpage>2288</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btw166</pub-id> <pub-id pub-id-type="pmid">27153604</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khan</surname> <given-names>S. H.</given-names></name> <name><surname>Hayat</surname> <given-names>M.</given-names></name> <name><surname>Porikli</surname> <given-names>F.</given-names></name></person-group> (<year>2019</year>). <article-title>Regularization of deep neural networks with spectral dropout.</article-title> <source><italic>Neural Netw.</italic></source> <volume>110</volume> <fpage>82</fpage>&#x2013;<lpage>90</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2018.09.009</pub-id> <pub-id pub-id-type="pmid">30504041</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kingma</surname> <given-names>D. P.</given-names></name> <name><surname>Ba</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>Adam: a method for stochastic optimization.</article-title> <source><italic>arXiv</italic> [Preprint]</source>. <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1412.6980">arXiv:1412.6980</ext-link></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Klema</surname> <given-names>V.</given-names></name> <name><surname>Laub</surname> <given-names>A.</given-names></name></person-group> (<year>1980</year>). <article-title>The singular value decomposition: Its computation and some applications.</article-title> <source><italic>IEEE Trans. Automat. Contr.</italic></source> <volume>25</volume> <fpage>164</fpage>&#x2013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1109/tac.1980.1102314</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B.-Q.</given-names></name> <name><surname>Feng</surname> <given-names>K.-Y.</given-names></name> <name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>T.</given-names></name> <name><surname>Cai</surname> <given-names>Y.-D.</given-names></name></person-group> (<year>2012</year>). <article-title>Prediction of protein-protein interaction sites by random forest algorithm with mRMR and IFS.</article-title> <source><italic>PLoS One</italic></source> <volume>7</volume>:<fpage>e43927</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0043927</pub-id> <pub-id pub-id-type="pmid">22937126</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>L.</given-names></name> <name><surname>Jing</surname> <given-names>L.</given-names></name> <name><surname>Huang</surname> <given-names>D.</given-names></name></person-group> (<year>2009</year>). &#x201C;<article-title>Protein-protein interaction extraction from biomedical literatures based on modified SVM-KNN</article-title>,&#x201D; in <source><italic>2009 International Conference on Natural Language Processing and Knowledge Engineering</italic></source>, (<publisher-loc>Dalian, China</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1109/NLPKE.2009.5313735</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nanni</surname> <given-names>L.</given-names></name> <name><surname>Brahnam</surname> <given-names>S.</given-names></name> <name><surname>Lumini</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>Wavelet images and Chou&#x2019;s pseudo amino acid composition for protein classification.</article-title> <source><italic>Amino Acids</italic></source> <volume>43</volume> <fpage>657</fpage>&#x2013;<lpage>665</lpage>. <pub-id pub-id-type="doi">10.1007/s00726-011-1114-9</pub-id> <pub-id pub-id-type="pmid">21993538</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nussbaumer</surname> <given-names>H. J.</given-names></name></person-group> (<year>1981</year>). <article-title>&#x201C;The fast Fourier transform,&#x201D;</article-title> in <source><italic>Fast Fourier Transform and Convolution Algorithms</italic></source> (<publisher-loc>Berlin</publisher-loc>: <publisher-name>Springer</publisher-name>) <fpage>80</fpage>&#x2013;<lpage>111</lpage>.</citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Onodera</surname> <given-names>R.</given-names></name> <name><surname>Watanabe</surname> <given-names>H.</given-names></name> <name><surname>Ishii</surname> <given-names>Y.</given-names></name></person-group> (<year>2005</year>). <article-title>Interferometric phase-measurement using a one-dimensional discrete Hilbert transform.</article-title> <source><italic>Opt. Rev.</italic></source> <volume>12</volume> <fpage>29</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1007/s10043-005-0029-7</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ponomareva</surname> <given-names>O.</given-names></name> <name><surname>Ponomarev</surname> <given-names>A.</given-names></name> <name><surname>Ponomarev</surname> <given-names>V.</given-names></name></person-group> (<year>2018</year>). &#x201C;<article-title>Evolution of forward and inverse discrete fourier transform</article-title>,&#x201D; in <source><italic>2018 IEEE East-West Design &#x0026; Test Symposium (EWDTS)</italic></source>, (<publisher-loc>Kazan, Russia</publisher-loc>: <publisher-name>IEEE</publisher-name>), <fpage>1</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1109/EWDTS.2018.8524820</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Read</surname> <given-names>R. R.</given-names></name> <name><surname>Treitel</surname> <given-names>S.</given-names></name></person-group> (<year>1973</year>). <article-title>The stabilization of two-dimensional recursive filters via the discrete Hilbert transform.</article-title> <source><italic>IEEE Trans. Geosci. Electron.</italic></source> <volume>11</volume> <fpage>153</fpage>&#x2013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1109/tge.1973.294304</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rhee</surname> <given-names>S. Y.</given-names></name> <name><surname>Beavis</surname> <given-names>W.</given-names></name> <name><surname>Berardini</surname> <given-names>T. Z.</given-names></name> <name><surname>Chen</surname> <given-names>G.</given-names></name> <name><surname>Dixon</surname> <given-names>D.</given-names></name> <name><surname>Doyle</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2003</year>). <article-title>The Arabidopsis Information Resource (TAIR): a model organism database providing a centralized, curated gateway to Arabidopsis biology, research materials and community.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>31</volume> <fpage>224</fpage>&#x2013;<lpage>228</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkg076</pub-id> <pub-id pub-id-type="pmid">12519987</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Richoux</surname> <given-names>F.</given-names></name> <name><surname>Servantie</surname> <given-names>C.</given-names></name> <name><surname>Bor&#x00E8;s</surname> <given-names>C.</given-names></name> <name><surname>T&#x00E9;letch&#x00E9;a</surname> <given-names>S.</given-names></name></person-group> (<year>2019</year>). <article-title>Comparing two deep learning sequence-based models for protein-protein interaction prediction.</article-title> <source><italic>arXiv</italic> [Preprint].</source> <ext-link ext-link-type="uri" xlink:href="https://arxiv.org/abs/1901.06268">arXiv:1901.06268</ext-link></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sledzieski</surname> <given-names>S.</given-names></name> <name><surname>Singh</surname> <given-names>R.</given-names></name> <name><surname>Cowen</surname> <given-names>L.</given-names></name> <name><surname>Berger</surname> <given-names>B.</given-names></name></person-group> (<year>2021</year>). <article-title>Sequence-based prediction of protein-protein interactions: a structure-aware interpretable deep learning model.</article-title> <source><italic>bioRxiv</italic> [preprint].</source> <pub-id pub-id-type="doi">10.1101/2021.01.22.427866</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stark</surname> <given-names>C.</given-names></name> <name><surname>Breitkreutz</surname> <given-names>B.-J.</given-names></name> <name><surname>Reguly</surname> <given-names>T.</given-names></name> <name><surname>Boucher</surname> <given-names>L.</given-names></name> <name><surname>Breitkreutz</surname> <given-names>A.</given-names></name> <name><surname>Tyers</surname> <given-names>M.</given-names></name></person-group> (<year>2006</year>). <article-title>BioGRID: a general repository for interaction datasets.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>34</volume> <fpage>D535</fpage>&#x2013;<lpage>D539</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkj109</pub-id> <pub-id pub-id-type="pmid">16381927</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stark</surname> <given-names>H.</given-names></name></person-group> (<year>1971</year>). <article-title>An extension of the Hilbert transform product theorem.</article-title> <source><italic>Proc. IEEE</italic></source> <volume>59</volume> <fpage>1359</fpage>&#x2013;<lpage>1360</lpage>. <pub-id pub-id-type="doi">10.1109/proc.1971.8420</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>T.</given-names></name> <name><surname>Zhou</surname> <given-names>B.</given-names></name> <name><surname>Lai</surname> <given-names>L.</given-names></name> <name><surname>Pei</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Sequence-based prediction of protein protein interaction using a deep-learning algorithm.</article-title> <source><italic>BMC Bioinformatics</italic></source> <volume>18</volume>:<fpage>277</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1700-2</pub-id> <pub-id pub-id-type="pmid">28545462</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tian</surname> <given-names>T.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Yan</surname> <given-names>H.</given-names></name> <name><surname>You</surname> <given-names>Q.</given-names></name> <name><surname>Yi</surname> <given-names>X.</given-names></name> <name><surname>Du</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2017</year>). <article-title>agriGO v2. 0: a GO analysis toolkit for the agricultural community, 2017 update.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>45</volume> <fpage>W122</fpage>&#x2013;<lpage>W129</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkx382</pub-id> <pub-id pub-id-type="pmid">28472432</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y.</given-names></name> <name><surname>You</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>T.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>PCVMZM: using the probabilistic classification vector machines model combined with a zernike moments descriptor to predict protein&#x2013;protein interactions from protein sequences.</article-title> <source><italic>Int. J. Mol. Sci.</italic></source> <volume>18</volume>:<fpage>1029</fpage>. <pub-id pub-id-type="doi">10.3390/ijms18051029</pub-id> <pub-id pub-id-type="pmid">28492483</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>F.</given-names></name> <name><surname>Zhao</surname> <given-names>C.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Deng</surname> <given-names>Y.</given-names></name> <name><surname>Shi</surname> <given-names>T.</given-names></name></person-group> (<year>2011</year>). <article-title>Exploring virus relationships based on virus-host protein-protein interaction network.</article-title> <source><italic>BMC Syst. Biol.</italic></source> <volume>5</volume>:<fpage>S11</fpage>. <pub-id pub-id-type="doi">10.1186/1752-0509-5-S3-S11</pub-id> <pub-id pub-id-type="pmid">22784617</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Xia</surname> <given-names>J.-F.</given-names></name> <name><surname>Gui</surname> <given-names>J.</given-names></name></person-group> (<year>2010</year>). <article-title>Prediction of protein-protein interactions from protein sequence using local descriptors.</article-title> <source><italic>Protein Pept. Lett.</italic></source> <volume>17</volume> <fpage>1085</fpage>&#x2013;<lpage>1090</lpage>. <pub-id pub-id-type="doi">10.2174/092986610791760306</pub-id> <pub-id pub-id-type="pmid">20509850</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yi</surname> <given-names>H.-C.</given-names></name> <name><surname>You</surname> <given-names>Z.-H.</given-names></name> <name><surname>Huang</surname> <given-names>D.-S.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>T.-H.</given-names></name> <name><surname>Li</surname> <given-names>L.-P.</given-names></name></person-group> (<year>2018</year>). <article-title>A deep learning framework for robust and accurate prediction of ncRNA-protein interactions using evolutionary information.</article-title> <source><italic>Mol. Ther. Nucleic Acids</italic></source> <volume>11</volume> <fpage>337</fpage>&#x2013;<lpage>344</lpage>. <pub-id pub-id-type="doi">10.1016/j.omtn.2018.03.001</pub-id> <pub-id pub-id-type="pmid">29858068</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>X.</given-names></name> <name><surname>Zheng</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Dou</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2012</year>). <article-title>Predicting subcellular location of apoptosis proteins with pseudo amino acid composition: approach from amino acid substitution matrix and auto covariance transformation.</article-title> <source><italic>Amino Acids</italic></source> <volume>42</volume> <fpage>1619</fpage>&#x2013;<lpage>1625</lpage>. <pub-id pub-id-type="doi">10.1007/s00726-011-0848-8</pub-id> <pub-id pub-id-type="pmid">21344173</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zahiri</surname> <given-names>J.</given-names></name> <name><surname>Mohammad-Noori</surname> <given-names>M.</given-names></name> <name><surname>Ebrahimpour</surname> <given-names>R.</given-names></name> <name><surname>Saadat</surname> <given-names>S.</given-names></name> <name><surname>Bozorgmehr</surname> <given-names>J. H.</given-names></name> <name><surname>Goldberg</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>LocFuse: human protein&#x2013;protein interaction prediction via classifier fusion using protein localization information.</article-title> <source><italic>Genomics</italic></source> <volume>104</volume> <fpage>496</fpage>&#x2013;<lpage>503</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2014.10.006</pub-id> <pub-id pub-id-type="pmid">25458812</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>M.</given-names></name> <name><surname>Zhang</surname> <given-names>F.</given-names></name> <name><surname>Wu</surname> <given-names>F.-X.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>M.</given-names></name></person-group> (<year>2020</year>). <article-title>Protein&#x2013;protein interaction site prediction through combining local and global features with deep neural networks.</article-title> <source><italic>Bioinformatics</italic></source> <volume>36</volume> <fpage>1114</fpage>&#x2013;<lpage>1120</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz699</pub-id> <pub-id pub-id-type="pmid">31593229</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zeng</surname> <given-names>Y.-H.</given-names></name> <name><surname>Guo</surname> <given-names>Y.-Z.</given-names></name> <name><surname>Xiao</surname> <given-names>R.-Q.</given-names></name> <name><surname>Yang</surname> <given-names>L.</given-names></name> <name><surname>Yu</surname> <given-names>L.-Z.</given-names></name> <name><surname>Li</surname> <given-names>M.-L.</given-names></name></person-group> (<year>2009</year>). <article-title>Using the augmented Chou&#x2019;s pseudo amino acid composition for predicting protein submitochondria locations based on auto covariance approach.</article-title> <source><italic>J. Theor. Biol.</italic></source> <volume>259</volume> <fpage>366</fpage>&#x2013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1016/j.jtbi.2009.03.028</pub-id> <pub-id pub-id-type="pmid">19341746</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Gao</surname> <given-names>P.</given-names></name> <name><surname>Yuan</surname> <given-names>J. S.</given-names></name></person-group> (<year>2010</year>). <article-title>Plant protein-protein interaction network and interactome.</article-title> <source><italic>Curr. Genomics</italic></source> <volume>11</volume> <fpage>40</fpage>&#x2013;<lpage>46</lpage>. <pub-id pub-id-type="doi">10.2174/138920210790218016</pub-id> <pub-id pub-id-type="pmid">20808522</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>G.</given-names></name> <name><surname>Wu</surname> <given-names>A.</given-names></name> <name><surname>Xu</surname> <given-names>X.-J.</given-names></name> <name><surname>Xiao</surname> <given-names>P.-P.</given-names></name> <name><surname>Lu</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>PPIM: a protein-protein interaction database for maize.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>170</volume> <fpage>618</fpage>&#x2013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1104/pp.15.01821</pub-id> <pub-id pub-id-type="pmid">26620522</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>Y. M.</given-names></name> <name><surname>Peyrin</surname> <given-names>F.</given-names></name> <name><surname>Goutte</surname> <given-names>R.</given-names></name></person-group> (<year>1990</year>). <article-title>The use of a two-dimensional Hilbert transform for Wigner analysis of 2-dimensional real signals.</article-title> <source><italic>Signal Process.</italic></source> <volume>19</volume> <fpage>205</fpage>&#x2013;<lpage>220</lpage>. <pub-id pub-id-type="doi">10.1016/0165-1684(90)90113-d</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="footnote1">
<label>1</label>
<p><ext-link ext-link-type="uri" xlink:href="https://www.arabidopsis.org/">https://www.arabidopsis.org/</ext-link></p></fn>
<fn id="footnote2">
<label>2</label>
<p><ext-link ext-link-type="uri" xlink:href="https://www.ebi.ac.uk/intact/">https://www.ebi.ac.uk/intact/</ext-link></p></fn>
<fn id="footnote3">
<label>3</label>
<p><ext-link ext-link-type="uri" xlink:href="https://thebiogrid.org/">https://thebiogrid.org/</ext-link></p></fn>
<fn id="footnote4">
<label>4</label>
<p><ext-link ext-link-type="uri" xlink:href="http://comp-sysbio.org/ppim/">http://comp-sysbio.org/ppim/</ext-link></p></fn>
<fn id="footnote5">
<label>5</label>
<p><ext-link ext-link-type="uri" xlink:href="http://systemsbiology.cau.edu.cn/agriGOv2/">http://systemsbiology.cau.edu.cn/agriGOv2/</ext-link></p></fn>
<fn id="footnote6">
<label>6</label>
<p><ext-link ext-link-type="uri" xlink:href="http://bis.zju.edu.cn/prin/">http://bis.zju.edu.cn/prin/</ext-link></p></fn>
<fn id="footnote7">
<label>7</label>
<p><ext-link ext-link-type="uri" xlink:href="http://blast.ncbi.nlm.nih.gov/Blast.cgi">http://blast.ncbi.nlm.nih.gov/Blast.cgi</ext-link></p></fn>
</fn-group>
</back>
</article>