<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Microbiol.</journal-id>
<journal-title>Frontiers in Microbiology</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Microbiol.</abbrev-journal-title>
<issn pub-type="epub">1664-302X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fmicb.2022.846915</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Microbiology</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>NNAN: Nearest Neighbor Attention Network to Predict Drug&#x2013;Microbe Associations</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Zhu</surname> <given-names>Bei</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1490595/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Yi</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x2020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhao</surname> <given-names>Pengcheng</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1473580/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Yiu</surname> <given-names>Siu-Ming</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/32453/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Yu</surname> <given-names>Hui</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1618873/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Shi</surname> <given-names>Jian-Yu</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c002"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/501947/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>School of Life Sciences, Northwestern Polytechnical University</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Computer Science, The University of Hong Kong</institution>, <addr-line>Hong Kong</addr-line>, <country>China</country></aff>
<aff id="aff3"><sup>3</sup><institution>School of Computer Science, Northwestern Polytechnical University</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Qi Zhao, University of Science and Technology Liaoning, China</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Wen Zhang, Huazhong Agricultural University, China; Lihong Peng, Hunan University of Technology, China</p></fn>
<corresp id="c001">&#x002A;Correspondence: Hui Yu, <email>huiyu@nwpu.edu.cn</email></corresp>
<corresp id="c002">Jian-Yu Shi, <email>jianyushi@nwpu.edu.cn</email></corresp>
<fn fn-type="equal" id="fn002"><p><sup>&#x2020;</sup>These authors have contributed equally to this work and share first authorship</p></fn>
<fn fn-type="other" id="fn004"><p>This article was submitted to Systems Microbiology, a section of the journal Frontiers in Microbiology</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>846915</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>12</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>02</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2022 Zhu, Xu, Zhao, Yiu, Yu and Shi.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhu, Xu, Zhao, Yiu, Yu and Shi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Many drugs can be metabolized by human microbes; the drug metabolites would significantly alter pharmacological effects and result in low therapeutic efficacy for patients. Hence, it is crucial to identify potential drug&#x2013;microbe associations (DMAs) before the drug administrations. Nevertheless, traditional DMA determination cannot be applied in a wide range due to the tremendous number of microbe species, high costs, and the fact that it is time-consuming. Thus, predicting possible DMAs in computer technology is an essential topic. Inspired by other issues addressed by deep learning, we designed a deep learning-based model named Nearest Neighbor Attention Network (NNAN). The proposed model consists of four components, namely, a similarity network constructor, a nearest-neighbor aggregator, a feature attention block, and a predictor. In brief, the similarity block contains a microbe similarity network and a drug similarity network. The nearest-neighbor aggregator generates the embedding representations of drug&#x2013;microbe pairs by integrating drug neighbors and microbe neighbors of each drug&#x2013;microbe pair in the network. The feature attention block evaluates the importance of each dimension of drug&#x2013;microbe pair embedding by a set of ordinary multi-layer neural networks. The predictor is an ordinary fully-connected deep neural network that functions as a binary classifier to distinguish potential DMAs among unlabeled drug&#x2013;microbe pairs. Several experiments on two benchmark databases are performed to evaluate the performance of NNAN. First, the comparison with state-of-the-art baseline approaches demonstrates the superiority of NNAN under cross-validation in terms of predicting performance. Moreover, the interpretability inspection reveals that a drug tends to associate with a microbe if it finds its top-<italic>l</italic> most similar neighbors that associate with the microbe.</p>
</abstract>
<kwd-group>
<kwd>deep learning</kwd>
<kwd>bipartite graph network</kwd>
<kwd>link prediction</kwd>
<kwd>drug&#x2013;microbe association</kwd>
<kwd>attention matrix</kwd>
</kwd-group>
<contract-num rid="cn001">No. 61872297, PI: JYS</contract-num>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="4"/>
<table-count count="4"/>
<equation-count count="15"/>
<ref-count count="47"/>
<page-count count="10"/>
<word-count count="6967"/>
</counts>
</article-meta>
</front>
<body>
<sec id="S1" sec-type="intro">
<title>Introduction</title>
<p>The human microbiome refers to all the microbes associated with a human body, including bacteriophages, archaea, bacteria, eukaryotes, and fungi (<xref ref-type="bibr" rid="B24">Lynch and Pedersen, 2016</xref>). To assess the diversity and functions of the human microbiome, the Human Microbiome Project (HMP) was supported by the National Institutes of Health (NIH) from 2007 to 2016 (<xref ref-type="bibr" rid="B35">Turnbaugh et al., 2007</xref>). HMP provided a complete description of the microbiome in five tissues of the human body, including skin, gut, nostrils, vagina, and mouth (<xref ref-type="bibr" rid="B1">Aagaard et al., 2013</xref>). Human microbes have been verified for their close associations with human health by cell experiments, animal experiments, epidemiological studies, clinical case studies (<xref ref-type="bibr" rid="B31">Schwabe and Jobin, 2013</xref>; <xref ref-type="bibr" rid="B24">Lynch and Pedersen, 2016</xref>), etc. Previous works have revealed that abnormal microbe communities lead to metabolic disorders [e.g., non-alcoholic fatty liver disease (<xref ref-type="bibr" rid="B38">Younossi et al., 2016</xref>), obesity, and diabetes mellitus (<xref ref-type="bibr" rid="B11">Jaacks et al., 2019</xref>; <xref ref-type="bibr" rid="B43">Zheng et al., 2018</xref>)]. Oral drug administration is a typical treatment. Many drugs, however, can be metabolized by human microbes, and the drug metabolites would significantly alter pharmacological effects and result in low therapeutic efficacy for patients. For example, after being modified by gut microbes, the compounds can lead to their activation [e.g., <italic>salicylazosulfapyridine</italic> (<xref ref-type="bibr" rid="B33">Sousa et al., 2014</xref>)] or inactivation [e.g., inactivation of the cardiac drug digoxin by the intestinal actinomycete <italic>Eggerthella lenta</italic> (<xref ref-type="bibr" rid="B8">Haiser et al., 2013</xref>)], or induce toxicity [e.g., 70% toxicity of Brivudine may be attributed to intestinal microorganisms (<xref ref-type="bibr" rid="B46">Zimmermann et al., 2019b</xref>)]. The persistent findings of microbiome-induced individual pathogenesis, phenotypes, and treatment responses boost the microbiome to be an integral part of precision medicine (<xref ref-type="bibr" rid="B12">Kashyap et al., 2017</xref>). Therefore, drug&#x2013;microbe association (DMA) prediction is of great significance for therapy and medicine development. However, the acquisition of DMAs needs a large scale of assays with high costs, low efficiency, and culturing limitations, and that are time-consuming. To identify DMAs rapidly and effectively, machine learning methods, especially deep learning-based methods, have attracted many scientists due to their inspiring applications in other areas [e.g., predicting microbe&#x2013;disease associations (<xref ref-type="bibr" rid="B9">He et al., 2018</xref>; <xref ref-type="bibr" rid="B26">Peng et al., 2018</xref>), drug&#x2013;drug interactions (<xref ref-type="bibr" rid="B39">Yu et al., 2021a</xref>), lncRNA&#x2013;miRNA interactions (<xref ref-type="bibr" rid="B41">Zhang L. et al., 2021</xref>), and lncRNA&#x2013;protein interactions (<xref ref-type="bibr" rid="B20">Lihong et al., 2021</xref>; <xref ref-type="bibr" rid="B44">Zhou et al., 2021</xref>)].</p>
<p>In recent years, researchers have applied Graph Attention Network [GAT (<xref ref-type="bibr" rid="B36">Velickovic et al., 2018</xref>)] to bioinformatics with remarkable results. For instance, <xref ref-type="bibr" rid="B42">Zhang Z. et al. (2021)</xref> used fragments containing functional groups to represent molecular maps for molecular property prediction through a fragment-oriented multi-scale graph attention model. <xref ref-type="bibr" rid="B3">Bang et al. (2021)</xref> made the prediction of polypharmacy side effects with enhanced interpretability based on graph feature attention network. Constructing a bipartite network is the most popular approach to represent associations between two types of nodes. The prediction problem of DMA can then be transformed into a link prediction problem in a bipartite graph network. However, few models predict DMAs through bipartite graph networks. For example, EGATMDA (<xref ref-type="bibr" rid="B23">Long et al., 2020b</xref>) used the drug&#x2013;disease&#x2013;microbe perspective to predict the DMAs, which does not show a direct relationship between drugs and microbes and may contain noise. HMDAKATZ (<xref ref-type="bibr" rid="B45">Zhu et al., 2019</xref>) predicted the interactions between drugs and microbes based on the <xref ref-type="bibr" rid="B13">Katz (1953)</xref>; the disadvantage of this method in the node&#x2019;s information transmission (i.e., a node with a high central value transmits its high influence to all its neighbors) may not be appropriate in real life. GCNMDA (<xref ref-type="bibr" rid="B22">Long et al., 2020a</xref>) used GCN, random walk with restart, and GAT to learn node features, which relies on the parameter &#x201C;step size&#x201D; when using the restart random walk algorithm. HNERMDA (<xref ref-type="bibr" rid="B21">Long and Luo, 2020</xref>) learned the drug&#x2013;microbe heterogeneous network information by metapath2vec measure, which considered the type of nodes in the meta-path-based random walk but the skip-gram does not treat them differently during training.</p>
<p>In the field of drug&#x2013;target interaction prediction, there is a widely accepted assumption that structurally similar drugs tend to interact with the same target (<xref ref-type="bibr" rid="B14">Khalili et al., 2012</xref>). Analogously, we anticipate that if a drug (d<sub><italic>x</italic></sub>) can associate with a microbe (b<sub><italic>p</italic></sub>), the other drugs associated with the same microbe (b<sub><italic>p</italic></sub>) are usually the first <italic>l</italic> nearest neighbors of the drug (d<sub><italic>x</italic></sub>). Therefore, we propose a new model, Nearest Neighbor Attention Network (NNAN), which aggregates the information from nodes&#x2019; neighbors according to their entity types and maps them into a unified embedding space for further predicting potential DMAs. The comparison with state-of-the-art methods on two different databases demonstrates the superiority of our NNAN. Moreover, its interpretability is illustrated and validates our assumption. Finally, the case study assesses its ability to find potential associations between drugs and microbes. In general, our contribution is as follows:</p>
<list list-type="simple">
<list-item>
<label>&#x2022;</label>
<p>We make use of three networks: drug&#x2013;drug similarity network, microbe&#x2013;microbe similarity network, and a drug&#x2013;microbe bipartite graph network. Imitate the idea of KNN [K-Nearest-Neighbor (<xref ref-type="bibr" rid="B5">Cover and Hart, 1967</xref>)] to learn the substructures of the bipartite graph network, which can promote the accuracy of link prediction.</p>
</list-item>
<list-item>
<label>&#x2022;</label>
<p>We follow the idea of GAT and use multiple DNNs to learn the weights of embedding features to improve the screening efficiency of potential associations.</p>
</list-item>
<list-item>
<label>&#x2022;</label>
<p>In a quantitative way, we verify the hypothesis that &#x201C;If a drug can associate with a microbe, the other drugs that associate with the microbe are usually the first <italic>l</italic> nearest neighbors to the drug.&#x201D;</p>
</list-item>
</list>
</sec>
<sec id="S2" sec-type="materials|methods">
<title>Materials and Methods</title>
<p>In this section, we describe a model for predicting DMAs in a bipartite graph network, named NNAN as shown in <xref ref-type="fig" rid="F1">Figure 1</xref>. It consists of four components: a similarity network constructor, a nearest-neighbor aggregator, a feature attention block, and a predictor. Firstly, the similarity network constructor is mainly used to build a drug similarity network and a microbe similarity network (section &#x201C;Similarity Networks&#x201D; for details). Secondly, the nearest-neighbor aggregator generates the embedding representations of drug&#x2013;microbe pairs by integrating drug neighbors and microbe neighbors of each drug&#x2013;microbe pair in the network (section &#x201C;Nearest-Neighbor Aggregator for Drug&#x2013;Microbe Pair Embeddings&#x201D; for details). Thirdly, the feature attention block evaluates the importance of each dimension of drug&#x2013;microbe pair embedding by a set of ordinary multi-layer neural networks (section &#x201C;Feature Attention Block&#x201D; for details). Finally, we make use of a fully-connected deep neural network as a binary classifier to predict potential DMAs.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The overall framework of NNAN for drug&#x2013;microbe association prediction.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmicb-13-846915-g001.tif"/>
</fig>
<sec id="S2.SS1">
<title>Similarity Networks</title>
<sec id="S2.SS1.SSS1">
<title>Drug Similarity Network</title>
<p>We calculate drug similarities by the following steps. First, drugs are represented by Functional-Class Fingerprints [FCFPs (<xref ref-type="bibr" rid="B28">Rogers and Hahn, 2010</xref>)], which is the generalized version of Extended-Connectivity Fingerprints [ECFPs (<xref ref-type="bibr" rid="B28">Rogers and Hahn, 2010</xref>)] with more attention to atom functions. The FCFPs is implemented by RDKit (<xref ref-type="bibr" rid="B17">Landrum, 2010</xref>). Second, the similarity between drug d<sub><italic>i</italic></sub> and drug d<sub><italic>j</italic></sub> is calculated by the Tanimoto coefficient (<xref ref-type="bibr" rid="B29">Rogers and Tanimoto, 1960</xref>) as follows:</p>
<disp-formula id="S2.E1">
<label>(1)</label>
<mml:math id="M1">
<mml:mrow>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mo fence="true">||</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo fence="true">||</mml:mo>
</mml:mrow>
<mml:mo rspace="5.8pt">+</mml:mo>
<mml:mrow>
<mml:mo fence="true">||</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo fence="true">||</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>-</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo>&#x22C5;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <bold>f</bold><sub>d<sub><italic>i</italic></sub></sub> and <bold>f</bold><sub>d<sub><italic>j</italic></sub></sub> represent the FCFPs vector of drug d<sub><italic>i</italic></sub> and drug d<sub><italic>j</italic></sub>, respectively, ||&#x22C5;|| indicates the norm of the vector.</p>
<p>Fingerprint similarity provides intuitive results: why the two molecules have been determined to be similar, but this transparency tends to vanish completely when molecular fingerprints are used as input to machine learning models. Inspired by the similarity maps (<xref ref-type="bibr" rid="B27">Riniker and Landrum, 2013</xref>), we calculate the contribution of each atom to the similarity between two molecules. To make it easier to distinguish the drugs, we regard d<sub><italic>i</italic></sub> as a reference drug, d<sub><italic>j</italic></sub> as a comparison drug, and <italic>S</italic>(d<sub><italic>i</italic></sub>,d<sub><italic>j</italic></sub>) as the base similarity of this drug pair. The RDKit will automatically number each atom of the comparison drug d<sub><italic>j</italic></sub> (<italic>K</italic> = {0,1,&#x2026;,<italic>t</italic>&#x2212;1}). Then, we remove the atoms of the comparison drug one by one in the order of the atomic numbers to form multiple new comparing drugs (<inline-formula><mml:math id="INEQ33"><mml:mrow><mml:mrow><mml:mrow><mml:msubsup><mml:mtext>d</mml:mtext><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mi mathvariant="normal">K</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi mathvariant="normal">K</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mo stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="false">}</mml:mo></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula>). We calculate the new similarity between the reference drug (d<sub><italic>i</italic></sub>) and the new comparison drug (<inline-formula><mml:math id="INEQ35"><mml:msubsup><mml:mtext>d</mml:mtext><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:msubsup></mml:math></inline-formula>), and regard the difference between the new similarity and the base similarity as the weight (<inline-formula><mml:math id="INEQ36"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>w</mml:mi></mml:mpadded><mml:msubsup><mml:mi/><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula>) of each removed atom. The weight <inline-formula><mml:math id="INEQ37"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>w</mml:mi></mml:mpadded><mml:msubsup><mml:mi/><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> is formulated as:</p>
<disp-formula id="S2.E2">
<label>(2)</label>
<mml:math id="M2">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>w</mml:mi>
</mml:mpadded>
<mml:mmultiscripts>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mprescripts/>
<mml:mi>j</mml:mi>
<mml:mi>k</mml:mi>
</mml:mmultiscripts>
<mml:mo>|</mml:mo>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>-</mml:mo>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mtext>d</mml:mtext>
<mml:mi>j</mml:mi>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>We set the dimension of the FCFPs vector to 1,024 bits, of which the non-zero bits indicate the occurrences of drug feature substructures. To obtain the weight of each non-zero bit, we add up the weights of all the atoms contained in the feature substructure:</p>
<disp-formula id="S2.E3">
<label>(3)</label>
<mml:math id="M3">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi>w</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mtext>bit</mml:mtext>
</mml:mrow>
<mml:mi>q</mml:mi>
</mml:msub>
</mml:msub>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>U</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:msub>
<mml:mi>M</mml:mi>
<mml:mi>q</mml:mi>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mi>w</mml:mi>
<mml:mi>j</mml:mi>
<mml:mi>k</mml:mi>
</mml:msubsup>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <italic>w</italic><sub><italic>bit</italic><sub><italic>q</italic></sub></sub> denotes the weight of the <italic>q</italic><sub><italic>th</italic></sub> dimensional bit of the FCFPs vector, and the function <italic>SUM</italic><sub><italic>q</italic></sub>(&#x22C5;) denotes the sum of all the atomic weights contained in the feature substructure represented by the <italic>q</italic><sub><italic>th</italic></sub> dimensional bit of the FCFPs.</p>
<p>Then, the weighted Tanimoto similarity (<xref ref-type="bibr" rid="B10">Ioffe, 2010</xref>) between the reference drug and the comparison drug can be calculated as follows:</p>
<disp-formula id="S2.E4">
<label>(4)</label>
<mml:math id="M4">
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>q</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>1024</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mtext>min</mml:mtext>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>q</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mtext>bit</mml:mtext>
</mml:mrow>
<mml:mi>q</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mi>q</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mo largeop="true" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>q</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mn>1024</mml:mn>
</mml:msubsup>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mtext>max</mml:mtext>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>q</mml:mi>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mi>w</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mtext>bit</mml:mtext>
</mml:mrow>
<mml:mi>q</mml:mi>
</mml:msub>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:msubsup>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:msub>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mi>q</mml:mi>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula><mml:math id="INEQ39"><mml:msubsup><mml:mtext mathvariant="bold">f</mml:mtext><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mi>q</mml:mi></mml:msubsup></mml:math></inline-formula> and <inline-formula><mml:math id="INEQ40"><mml:msubsup><mml:mtext mathvariant="bold">f</mml:mtext><mml:msub><mml:mrow><mml:mtext>d</mml:mtext></mml:mrow><mml:mi>j</mml:mi></mml:msub><mml:mi>q</mml:mi></mml:msubsup></mml:math></inline-formula> denote the <italic>q</italic><sub><italic>th</italic></sub> dimension of the FCFPs vectors for the reference drug and the comparison drug.</p>
<p>Based on drug similarities, we can build a drug similarity network <italic>Net</italic><sub>d</sub>, where nodes are drugs. There are edges between the drugs if these drugs associate with the same microbe; the edges are weighted by drug similarities.</p>
</sec>
<sec id="S2.SS1.SSS2">
<title>Microbe Similarity Network</title>
<p>To calculate microbe similarities, we use BLAST (<xref ref-type="bibr" rid="B2">Altschul et al., 1990</xref>) to make pairwise alignments of microbial genomes. Specifically, the main function of BLAST is to discover local similarity regions between sequences and then use the local sequence alignment algorithm (<xref ref-type="bibr" rid="B32">Smith and Waterman, 1981</xref>) to calculate the similarity. For example, <inline-formula><mml:math id="INEQ42"><mml:mrow><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext>G</mml:mtext><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow></mml:msub></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mpadded width="+8.3pt"><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>A</mml:mtext></mml:mrow><mml:mi>n</mml:mi></mml:msubsup></mml:mpadded><mml:mo>&#x2062;</mml:mo><mml:mpadded width="+3.3pt"><mml:mi>and</mml:mi></mml:mpadded><mml:mo>&#x2062;</mml:mo><mml:mpadded width="+3.3pt"><mml:msub><mml:mtext>G</mml:mtext><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow></mml:msub></mml:mpadded></mml:mrow><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow><mml:mn>1</mml:mn></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>&#x2062;</mml:mo><mml:mi mathvariant="normal">&#x2026;</mml:mi><mml:mo>&#x2062;</mml:mo><mml:msubsup><mml:mtext>g</mml:mtext><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow><mml:mi>m</mml:mi></mml:msubsup></mml:mrow></mml:mrow></mml:math></inline-formula> are the genome sequences of microbe A and microbe B, where <italic>n</italic> and <italic>m</italic> are the lengths of sequences G<sub>A</sub> and G<sub>B</sub>, respectively. BLAST creates the scoring matrix <bold>H</bold><sub>(<bold>n</bold>+<bold>1</bold>)&#x00D7;(<bold>m</bold>+<bold>1</bold>)</sub> and makes the first row and column elements zero. The formula for the element <bold>H</bold><sub><italic>ij</italic></sub>(<bold>H</bold><sub><italic>ij</italic></sub> &#x2208; <bold>H</bold><sub>(<bold>n</bold>+<bold>1</bold>)&#x00D7;(<bold>m</bold>+<bold>1</bold>)</sub>,<italic>i</italic> = 1,2,&#x2026;,<italic>n</italic>;<italic>j</italic> = 1,2,&#x2026;,<italic>m</italic>) in this scoring matrix is:</p>
<disp-formula id="S2.Ex1">
<label>(5)</label>
<mml:math id="M5">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mtext>max</mml:mtext>
<mml:mrow>
<mml:mo>{</mml:mo>
<mml:mtable displaystyle="true" rowspacing="0pt">
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mtext>Score</mml:mtext>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>-</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>-</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>-</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mrow>
</mml:msub>
<mml:mo>-</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="center">
<mml:mn>0</mml:mn>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:mtext>g</mml:mtext>
</mml:mpadded>
<mml:mmultiscripts>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mprescripts/>
<mml:mrow>
<mml:mtext>A</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:mmultiscripts>
<mml:mpadded width="+3.3pt">
<mml:mtext>g</mml:mtext>
</mml:mpadded>
<mml:mmultiscripts>
<mml:mo>,</mml:mo>
<mml:mprescripts/>
<mml:mrow>
<mml:mtext>B</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:mmultiscripts>
<mml:mpadded width="+3.3pt">
<mml:mtext>Score</mml:mtext>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:mtext>g</mml:mtext>
</mml:mpadded>
<mml:mmultiscripts>
<mml:mo rspace="5.8pt">&#x2260;</mml:mo>
<mml:mprescripts/>
<mml:mrow>
<mml:mtext>A</mml:mtext>
</mml:mrow>
<mml:mi>i</mml:mi>
</mml:mmultiscripts>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mtext>g</mml:mtext>
</mml:mpadded>
<mml:mmultiscripts>
<mml:mo>,</mml:mo>
<mml:mprescripts/>
<mml:mrow>
<mml:mtext>B</mml:mtext>
</mml:mrow>
<mml:mi>j</mml:mi>
</mml:mmultiscripts>
<mml:mpadded width="+3.3pt">
<mml:mtext>Score</mml:mtext>
</mml:mpadded>
<mml:mo rspace="10.8pt">=</mml:mo>
<mml:mo>-</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>the highest value in the matrix <bold>H</bold><sub>(<italic>n</italic> + 1)&#x00D7;(<italic>m</italic> + 1)</sub> is chosen as <italic>sw</italic>(G<sub>A</sub>,G<sub>B</sub>). The similarity between microbes A and B is adopted by the same definition as <xref ref-type="bibr" rid="B37">Yamanishi et al. (2008)</xref>, as follows:</p>
<disp-formula id="S2.E6">
<label>(6)</label>
<mml:math id="M6">
<mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mtext>b</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="normal">A</mml:mi>
<mml:mo rspace="7.5pt">,</mml:mo>
<mml:mi mathvariant="normal">B</mml:mi>
<mml:mo rspace="5.8pt" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>A</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>B</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>A</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>A</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo rspace="5.8pt" stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">&#x00D7;</mml:mo>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>w</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>B</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>G</mml:mtext>
<mml:mrow>
<mml:mtext>B</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>Based on microbe similarities, we can build a microbe similarity network <italic>Net</italic><sub>b</sub>, where nodes are microbes. There are edges between the microbes if these microbes associate with the same drug; the edges are weighted by microbe similarities.</p>
</sec>
</sec>
<sec id="S2.SS2">
<title>Nearest-Neighbor Aggregator for Drug&#x2013;Microbe Pair Embeddings</title>
<p>In this section, inspired by the idea of KNN [K-Nearest-Neighbor (<xref ref-type="bibr" rid="B5">Cover and Hart, 1967</xref>)], we learn the substructures of the bipartite graph network to obtain the embedding representations of drug&#x2013;microbe pairs.</p>
<p>First, we construct the drug&#x2013;microbe bipartite graph network, G=(D,B,E), where D={d<sub>1</sub>,d<sub>2</sub>,&#x2026;,d<sub><italic>m</italic></sub>} represents <italic>m</italic> drugs, B={b<sub>1</sub>,b<sub>2</sub>,&#x2026;,b<sub><italic>n</italic></sub>} represents <italic>n</italic> microbes, and each edge (e<sub><italic>ij</italic></sub>) in edge set E connects two nodes that belong to two different sets of vertexes (i.e., <italic>i</italic> in D, <italic>j</italic> in B). We regard the DMAs as bidirectional links. That is, e<sub>d<sub><italic>x</italic></sub>&#x2192;b<sub><italic>p</italic></sub></sub> denotes the edge pointing from the drug d<sub><italic>x</italic></sub> to the microbe b<sub><italic>p</italic></sub>, and <italic>e</italic><sub><italic>b<sub>p</sub>&#x2192;d<sub>x</sub></italic></sub> denotes the edge pointing from the microbe b<sub><italic>p</italic></sub> to the drug d<sub><italic>x</italic></sub>. Correspondingly, the nearest-neighbor aggregator contains two blocks (<xref ref-type="fig" rid="F2">Figure 2</xref>), the microbe-specific drug neighbor aggregator (MsDNA), and the drug-specific microbe neighbor aggregator (DsMNA). Due to their architectures being similar, we only illustrate the MsDNA block in this section.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Nearest-Neighbor Aggregator block. <bold>(A)</bold> Microbe-specific drug neighbor aggregator (MsDNA); the embedding representation of the unidirectional edge, which is from the drug d<sub><italic>x</italic></sub> to the microbe b<sub><italic>p</italic></sub>. <bold>(B)</bold> Drug-specific microbe neighbor aggregator (DsMNA); the embedding representation of the unidirectional edge, which is from the microbe b<sub><italic>p</italic></sub> to the drug d<sub><italic>x</italic></sub>, where <inline-graphic xlink:href="fmicb-13-846915-i100.jpg"/><italic><sup>x</sup></italic>&#x2286;<inline-graphic xlink:href="fmicb-13-846915-i100.jpg"/> is a set of instantiated keywords, <inline-graphic xlink:href="fmicb-13-846915-i100.jpg"/><italic><sup>x</sup></italic> denotes the neighbors of microbe b<sub><italic>p</italic></sub> in the <italic>Net</italic><sub>b</sub>. <italic>S</italic><sub>b</sub>(b<sub><italic>p</italic></sub>,<italic>m</italic><sub><italic>j</italic></sub>) denotes the similarity of b<sub><italic>p</italic></sub> and <italic>m<sub>j</sub></italic>. <bold>h</bold><sub><italic>j</italic></sub> is the corresponding one-hot encoding vector of <italic>m<sub>j</sub></italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmicb-13-846915-g002.tif"/>
</fig>
<p>Microbe-specific drug neighbor aggregator (<xref ref-type="fig" rid="F2">Figure 2A</xref>) contains a virtual key dictionary; &#x1D4A9;={<italic>n</italic><sub>1</sub>,<italic>n</italic><sub>2</sub>,&#x2026;,<italic>n</italic><sub><italic>m</italic></sub>} indicates all the drugs. In the dictionary, we imitate the idea of KNN to learn the substructures of the bipartite graph network, where virtual keys are sorted by their semantic nearest neighbors. In simple terms, <italic>n<sub>1</sub></italic> denotes d<sub><italic>x</italic></sub> itself, its nearest neighbor is the second key, and the farthest neighbor is the last key. The embedding representation of the edge, which is from drug d<sub><italic>x</italic></sub> to microbe b<sub><italic>p</italic></sub>, is formulated as follows:</p>
<disp-formula id="S2.E7">
<label>(7)</label>
<mml:math id="M7">
<mml:mrow>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:munderover>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mrow>
<mml:mo>|</mml:mo>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D4A9;</mml:mi>
<mml:mo>|</mml:mo>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mtext mathvariant="bold">v</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mpadded>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>i</mml:mi>
<mml:mpadded width="+3.3pt">
<mml:mi>f</mml:mi>
</mml:mpadded>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>&#x2209;</mml:mo>
<mml:msup>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D4A9;</mml:mi>
<mml:mi>p</mml:mi>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mtext>d</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>n</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x1D4A9;<sup><italic>p</italic></sup>&#x2286;&#x1D4A9; is a set of instantiated keywords, and &#x1D4A9;<sup><italic>p</italic></sup> denotes the neighbors of d<sub><italic>x</italic></sub> in the <italic>Net</italic><sub>d</sub>. <italic>S</italic><sub>d</sub>(d<sub><italic>x</italic></sub>,<italic>n</italic><sub><italic>i</italic></sub>) denotes the similarity of d<sub><italic>x</italic></sub> and <italic>n<sub>i</sub></italic>, and <bold>v</bold><sub><italic>i</italic></sub> is the corresponding one-hot encoding vector of <italic>n<sub>i</sub></italic> (i.e., the one-hot encoding has a non-zero value only in the <italic>i</italic><sub><italic>th</italic></sub> element, and all other position elements are zero).</p>
<p>Similarly, DsMNA (<xref ref-type="fig" rid="F2">Figure 2B</xref>) makes the single directional embedding representation from b<sub><italic>p</italic></sub>tod<sub><italic>x</italic></sub> as <bold>a</bold>(b<sub><italic>p</italic></sub>,d<sub><italic>x</italic></sub>). Then, the representation of drug&#x2013;microbe pair could be encoded as</p>
<disp-formula id="S2.E8">
<label>(8)</label>
<mml:math id="M8">
<mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="bold">e</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo rspace="5.8pt">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mo>[</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo lspace="2.5pt" rspace="2.5pt">&#x2225;</mml:mo>
<mml:mrow>
<mml:mtext mathvariant="bold">a</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:msub>
<mml:mtext>b</mml:mtext>
<mml:mi>p</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mtext>d</mml:mtext>
<mml:mi>x</mml:mi>
</mml:msub>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <bold>e</bold>(d<sub><italic>x</italic></sub>,b<sub><italic>p</italic></sub>) is generated <italic>via</italic> the concatenation of bidirectional embedding, and &#x2225; is the concatenation operation. All the embedding representations of drug&#x2013;microbe pairs could stack as a matrix <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>, where <italic>k</italic> is the number of all the drug&#x2013;microbe pairs and <italic>g</italic> is the dimension of each embedding. The nearest-neighbor aggregator effectively learns the bipartite graph substructures, and <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> will be input into a feature attention block to select crucial features for achieving a better DMA prediction.</p>
</sec>
<sec id="S2.SS3">
<title>Feature Attention Block</title>
<p>To improve the performance of the prediction, we build the feature attention block (<xref ref-type="fig" rid="F3">Figure 3</xref>) for updating the embedding of drug&#x2013;microbe pairs.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Feature attention block. Input the representation matrix <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> into a set of DNNs, then we obtain an attention matrix <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>of drug&#x2013;microbe embedding features. After the element-wise product operation of <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> and <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>, the final feature matrix <inline-formula><mml:math id="INEQ22"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">F</mml:mtext><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the drug&#x2013;microbe pairs is obtained.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmicb-13-846915-g003.tif"/>
</fig>
<p>Recall the equation of output feature representation in GAT (<xref ref-type="bibr" rid="B36">Velickovic et al., 2018</xref>):</p>
<disp-formula id="S2.E9">
<label>(9)</label>
<mml:math id="M9">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mover accent="true">
<mml:msubsup>
<mml:mi>h</mml:mi>
<mml:mi>i</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msubsup>
<mml:mo>&#x2192;</mml:mo>
</mml:mover>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D4A6;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:msub>
<mml:mi mathvariant="normal">&#x03B1;</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2062;</mml:mo>
<mml:mtext mathvariant="bold">W</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mover accent="true">
<mml:msub>
<mml:mi>h</mml:mi>
<mml:mi>j</mml:mi>
</mml:msub>
<mml:mo>&#x2192;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03C3; is a nonlinear activation function, &#x1D4A6;<sup><italic>i</italic></sup> is the first-order neighbors of node <italic>i</italic> (including <italic>i</italic>), &#x03B1;<sub><italic>ij</italic></sub> is the coefficients computed by the attention mechanism, and <bold>W</bold> is a weight matrix. To make equation (9) easier to understand. We compute the coefficients as:</p>
<disp-formula id="S2.E10">
<label>(10)</label>
<mml:math id="M10">
<mml:mrow>
<mml:mrow>
<mml:munder>
<mml:mo largeop="true" movablelimits="false" symmetric="true">&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D4A6;</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mi mathvariant="normal">&#x03B1;</mml:mi>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">A</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mo>&#x2299;</mml:mo>
<mml:mtext mathvariant="bold">M</mml:mtext>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula><mml:math id="INEQ78"><mml:mrow><mml:mpadded width="+3.3pt"><mml:mover accent="true"><mml:mtext mathvariant="bold">A</mml:mtext><mml:mo>~</mml:mo></mml:mover></mml:mpadded><mml:mo rspace="5.8pt">=</mml:mo><mml:mrow><mml:mtext mathvariant="bold">A</mml:mtext><mml:mo>+</mml:mo><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow></mml:mrow></mml:math></inline-formula> is the adjacency matrix of the undirected graph G with added self-connections (<xref ref-type="bibr" rid="B15">Kipf and Welling, 2017</xref>), &#x2299; is the element-wise product operation, and <bold>M</bold> is the attention matrix. Then, the layer-wise propagation rules in GAT can be formulated as:</p>
<disp-formula id="S2.E11">
<label>(11)</label>
<mml:math id="M11">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msup>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>+</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03C3;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">A</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mo>&#x2299;</mml:mo>
<mml:mtext mathvariant="bold">M</mml:mtext>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="bold">H</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2062;</mml:mo>
<mml:msup>
<mml:mtext mathvariant="bold">W</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03C3; is a nonlinear activation function, and <bold>W</bold><sup>(<italic>l</italic>)</sup> is the weight matrix of the <italic>l</italic><sub><italic>th</italic></sub> neural network layer.</p>
<p>Inspired by the conception of the layer-wise propagation rules in GAT, we calculate the augmented representation matrix <inline-formula><mml:math id="INEQ81"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">F</mml:mtext><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> by</p>
<disp-formula id="S2.E12">
<label>(12)</label>
<mml:math id="M12">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">F</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>k</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">&#x00D7;</mml:mo>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mtext mathvariant="bold">E</mml:mtext>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>k</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">&#x00D7;</mml:mo>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2299;</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">M</mml:mtext>
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>k</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">&#x00D7;</mml:mo>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> is the representation matrix of the drug&#x2013;microbe pairs obtained from the nearest-neighbor aggregator, <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> is an attention matrix of <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>, and &#x2299; is the element-wise product operation. We take the representation matrix <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> as a feature matrix <bold>F</bold> (<bold>F</bold>={<bold>f</bold><sub>1</sub>,<bold>f</bold><sub>2</sub>,&#x2026;,<bold>f</bold><sub><italic>g</italic></sub>}), which is composed of <italic>g</italic> column vectors (<bold>f</bold><sub><italic>i</italic></sub>(<italic>i</italic> = 1,2,&#x2026;,<italic>g</italic>)). The feature attention block mainly uses <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> to indicate the importance of features in the <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>. Each feature dimension <bold>f</bold><sub><italic>i</italic></sub> can be labeled as &#x201C;selected&#x201D; or &#x201C;discarded&#x201D; in a hard way, or be associated with a probability to be selected in a soft way; we employ DNNs to model the mapping by</p>
<disp-formula id="S2.E13">
<label>(13)</label>
<mml:math id="M13">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:msub>
<mml:mtext mathvariant="bold">m</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mtext>DNNs</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:msub>
<mml:mtext mathvariant="bold">f</mml:mtext>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>the DNN contains an input layer for each element of the feature dimension <bold>f</bold><sub><italic>i</italic></sub> and an output layer with sigmoid as its activation function.</p>
<p>In total, we build <italic>k</italic>&#x00D7;<italic>g</italic> DNNs to obtain <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>. The final feature matrix <inline-formula><mml:math id="INEQ95"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">F</mml:mtext><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> of the drug&#x2013;microbe pairs is obtained after the element-wise product operation of <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> and <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>. <inline-formula><mml:math id="INEQ98"><mml:msub><mml:mover accent="true"><mml:mtext mathvariant="bold">F</mml:mtext><mml:mo>~</mml:mo></mml:mover><mml:mrow><mml:mpadded width="+3.3pt"><mml:mi>k</mml:mi></mml:mpadded><mml:mo rspace="5.8pt">&#x00D7;</mml:mo><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is further fed into a predictor to achieve better predictive performance.</p>
</sec>
<sec id="S2.SS4">
<title>Predictor</title>
<p>To implement the link prediction in the drug&#x2013;microbe bipartite graph network, an ordinary DNN is utilized as the binary predictor that contains an input layer for the embedding representation of drug&#x2013;microbe pairs, a hidden layer with ReLU as its activation function, and the two-neuron output layer with Sigmoid as its activation function. The output layer generates a probability that indicates the association likelihood of the drug and the microbe. The probability is formulated as:</p>
<disp-formula id="S2.E14">
<label>(14)</label>
<mml:math id="M14">
<mml:mrow>
<mml:mpadded width="+3.3pt">
<mml:mi>P</mml:mi>
</mml:mpadded>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03C6;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mrow>
<mml:mtext>ReLU</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:mi class="ltx_font_mathcaligraphic">&#x2131;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo>(</mml:mo>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">F</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where &#x03C6; is the sigmoid activation function, and &#x2131;(&#x22C5;) is the fully-connected layer.</p>
<p>The entire network of NNAN with the nearest-neighbor aggregator, feature attention weights, and DNN weights can be jointly optimized through the binary cross-entropy loss as follows:</p>
<disp-formula id="S2.E15">
<label>(15)</label>
<mml:math id="M15">
<mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>o</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi>s</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:mi>s</mml:mi>
</mml:mpadded>
</mml:mrow>
<mml:mo rspace="5.8pt">=</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="bold">Y</mml:mtext>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:mtext>log</mml:mtext>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D49F;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">F</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>-</mml:mo>
<mml:mtext mathvariant="bold">Y</mml:mtext>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2062;</mml:mo>
<mml:mpadded width="+3.3pt">
<mml:mtext>log</mml:mtext>
</mml:mpadded>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>-</mml:mo>
<mml:mrow>
<mml:mi class="ltx_font_mathcaligraphic">&#x1D49F;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mover accent="true">
<mml:mtext mathvariant="bold">F</mml:mtext>
<mml:mo>~</mml:mo>
</mml:mover>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo>+</mml:mo>
<mml:mrow>
<mml:mi mathvariant="normal">&#x03BB;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mi class="ltx_font_mathcaligraphic">&#x211B;</mml:mi>
<mml:mo>&#x2062;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi mathvariant="normal">&#x03B8;</mml:mi>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <bold>Y</bold> is the truth labels of drug&#x2013;microbe pairs, &#x1D49F;(&#x22C5;) is the DNN, &#x03B8; denotes the weight parameters in the entire network, &#x211B;(&#x22C5;) is an L<sub>2</sub>-norm, and &#x03BB; is coefficient of the regularization item.</p>
</sec>
</sec>
<sec id="S3">
<title>Experiments and Results</title>
<sec id="S3.SS1">
<title>Data</title>
<p>In our experiments, two databases are collected from MDAD (<xref ref-type="bibr" rid="B34">Sun et al., 2018</xref>) and <xref ref-type="bibr" rid="B47">Zimmermann et al. (2019a)</xref>, respectively. The former work MDAD (<xref ref-type="bibr" rid="B34">Sun et al., 2018</xref>) investigated 5,505 clinically or experimentally DMAs between 1,388 drugs and 180 microbes. After removing redundant information, these association entries are grouped into Database 1, which contains 999 drugs, 133 microbes, and 1,708 DMAs.</p>
<p>The latter work (<xref ref-type="bibr" rid="B47">Zimmermann et al., 2019a</xref>) originally studied how 76 kinds of human gut bacteria metabolize 271 oral drugs, and found that 176 out of 217 drugs are significantly consumed by at least one bacteria strain. These associations are grouped into Database 2, which includes 176 drugs, 76 bacteria, and 4,194 associations (These two databases are shown in <xref ref-type="table" rid="T1">Table 1</xref>).</p>
<table-wrap position="float" id="T1">
<label>TABLE 1</label>
<caption><p>The statistics of two databases.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="center">Drugs</td>
<td valign="top" align="center">Microbes</td>
<td valign="top" align="center">Associations</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Database 1</td>
<td valign="top" align="center">999</td>
<td valign="top" align="center">133</td>
<td valign="top" align="center">1,708</td>
</tr>
<tr>
<td valign="top" align="left">Database 2</td>
<td valign="top" align="center">176</td>
<td valign="top" align="center">76</td>
<td valign="top" align="center">4,194</td>
</tr>
</tbody>
</table></table-wrap>
</sec>
<sec id="S3.SS2">
<title>Comparison</title>
<p>Since there are few existing approaches for predicting DMAs, we compare NNAN with three state-of-the-art methods, which were raised for bipartite link prediction.</p>
<list list-type="simple">
<list-item>
<label>&#x2022;</label>
<p>LAGCN (<xref ref-type="bibr" rid="B40">Yu et al., 2021b</xref>): A layer attention graph convolutional network for the drug&#x2013;disease association prediction.</p>
</list-item>
</list>
<list list-type="simple">
<list-item>
<label>&#x2022;</label>
<p>NIMCGCN (<xref ref-type="bibr" rid="B19">Li et al., 2020</xref>): A neural inductive matrix completion with graph convolutional networks for miRNA&#x2013;disease association prediction.</p>
</list-item>
</list>
<list list-type="simple">
<list-item>
<label>&#x2022;</label>
<p>GCNMDA (<xref ref-type="bibr" rid="B22">Long et al., 2020a</xref>): Predicting human microbe&#x2013;drug associations <italic>via</italic> graph convolutional network with conditional random field.</p>
</list-item>
</list>
<p>To evaluate the performance of these methods, we regard the known DMA pairs as positive samples and unlabeled DMA pairs as negative samples (<xref ref-type="bibr" rid="B25">Peng et al., 2020</xref>; <xref ref-type="bibr" rid="B18">Li et al., 2022</xref>). We set up a 5-fold cross-validation scenario in which we randomly divide positive samples and negative samples into five groups, respectively. One group of positive samples and one group of negative samples are treated as test samples in turn for each round. The remaining groups are used for training purposes. Our model is trained by Gradient Descent Optimizer (<xref ref-type="bibr" rid="B4">Cauchy, 2009</xref>), with batch size 3,000 for 2,000 epochs, the initial learning rate is set to 0.9, and the regularization rate is set to 2e-4. We use AUROC (area under the receiver operating characteristic curve) and AUPRC (area under the precision-recall curve) as metrics to measure the DMA prediction performance. Moreover, we investigate the running time in terms of per epoch.</p>
<p>The comparison (<xref ref-type="table" rid="T2">Table 2</xref>) shows that NNAN obtains the best AUROC value (0.911) and the best AUPRC value (0.502) in Database 1. NNAN attains the next-highest AUROC value (0.902) and the best AUPRC value (0.840) in Database 2. To further present the performance of NNAN, we calculate the running time for one epoch of the baselines and NNAN, respectively. As presented, with the same computing equipment, NNAN takes the third-shortest running time in Database 1 and the shortest running time in Database 2. In general, we can see that NNAN are comparable in terms of AUROC, AUPRC, and computation time. It demonstrates that NNAN is superior to other methods on the databases we collected.</p>
<table-wrap position="float" id="T2">
<label>TABLE 2</label>
<caption><p>The performance comparison of DMA prediction.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Method</td>
<td valign="top" align="center" colspan="3">Database 1<hr/></td>
<td valign="top" align="center" colspan="3">Database 2<hr/></td>
</tr>
<tr>
<td/>
<td valign="top" align="center">AUROC</td>
<td valign="top" align="center">AUPRC</td>
<td valign="top" align="center">Time (s/epoch)</td>
<td valign="top" align="center">AUROC</td>
<td valign="top" align="center">AUPRC</td>
<td valign="top" align="center">Time (s/epoch)</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">LAGCN</td>
<td valign="top" align="center">0.861</td>
<td valign="top" align="center"><underline>0.323</underline></td>
<td valign="top" align="center"><bold>0.201</bold></td>
<td valign="top" align="center"><bold>0.944</bold></td>
<td valign="top" align="center"><underline>0.721</underline></td>
<td valign="top" align="center"><underline>0.021</underline></td>
</tr>
<tr>
<td valign="top" align="left">NIMCGCN</td>
<td valign="top" align="center">0.778</td>
<td valign="top" align="center">0.156</td>
<td valign="top" align="center">19.076</td>
<td valign="top" align="center">0.815</td>
<td valign="top" align="center">0.720</td>
<td valign="top" align="center">0.721</td>
</tr>
<tr>
<td valign="top" align="left">GCNMDA</td>
<td valign="top" align="center"><underline>0.894</underline></td>
<td valign="top" align="center">0.042</td>
<td valign="top" align="center"><underline>0.341</underline></td>
<td valign="top" align="center">0.821</td>
<td valign="top" align="center">0.177</td>
<td valign="top" align="center">0.127</td>
</tr>
<tr>
<td valign="top" align="left">NNAN</td>
<td valign="top" align="center"><bold>0.911</bold></td>
<td valign="top" align="center"><bold>0.502</bold></td>
<td valign="top" align="center">0.649</td>
<td valign="top" align="center"><underline>0.902</underline></td>
<td valign="top" align="center"><bold>0.840</bold></td>
<td valign="top" align="center"><bold>0.019</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The highest value is indicated in bold, and the next highest value is underlined.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S3.SS3">
<title>Interpretability of Nearest Neighbor Attention Network</title>
<p>How does the NNAN interpret the hypothesis that &#x201C;If a drug can associate with a microbe, the other drugs that associate with the microbe are usually the first <italic>l</italic> nearest neighbors to the drug.&#x201D;</p>
<p>The model has two significant advantages to enhance interpretability. First, each column vector <bold>m<sub>i</sub></bold> of <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> indicates the global importance of each feature dimension <bold>f</bold><sub><italic>i</italic></sub>. Moreover, the element-wise product between <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> and <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> generates the importance map of embedding features.</p>
<p>We first use the MsDNA in the nearest-neighbor aggregator block to show how the representation of drug&#x2013;microbe pairs can provide intuitive hints, on which embedding features lead to the association. For the queried drug d<sub><italic>x</italic></sub> to the microbe b<sub><italic>p</italic></sub> of associated, non-zero cells in the embedding representation of <bold>a</bold>(<italic>d</italic><sub><italic>x</italic></sub>,b<sub><italic>p</italic></sub>) stand for its attention values derived from the drugs commonly linking b<sub><italic>p</italic></sub>. Since the keys are sorted in descending order from the drug itself (<italic>n<sub>1</sub></italic>) to the farthest neighbor (<italic>n<sub>m</sub></italic>), the positions of non-zero cells are crucial to the final association.</p>
<p>Take Database 1 as an example. By calculating two average embedding vectors for approved DMAs and unlabeled drug&#x2013;microbe pairs, we obtained a distribution along with the drug key dictionary from <italic>n<sub>1</sub></italic> to <italic>n</italic><sub><italic>66</italic></sub>(<xref ref-type="fig" rid="F4">Figure 4A</xref>). As illustrated, the significantly high values of embedding features occurring among the first <italic>l</italic> nearest neighbors reveal that a drug (d<sub><italic>x</italic></sub>) associated with a specific microbe (b<sub><italic>p</italic></sub>) can always find its top-<italic>l</italic> nearest neighbors among other drugs that associate with the same microbe. This observation demonstrates that a drug is possibly associated with the microbe if it has more non-zero value cells on the positions of the first <italic>l</italic> feature dimensions. This phenomenon could be caused by the fact that over 80% of approved drugs are of &#x201C;follow-on&#x201D; or &#x201C;me-too&#x201D; drugs. Due to high cost and high risk, the design of novel drugs, except for pioneer drugs, always starts from the structures of one or several existing drugs and then slightly modify them until meeting pharmacological needs (<xref ref-type="bibr" rid="B7">DiMasi and Faden, 2011</xref>). Analogously, the results of the DsMNA block along the microbe neighbor aggregator keys reveal that a microbe associated with a specific drug usually finds its near neighbors associated with the same drug.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>Mensurable clues of embedding features to the association outcome. <bold>(A)</bold> The distribution of embedding features along with the sorted drug neighbor keys. <bold>(B)</bold> The distribution of feature importance along with sorted node neighbor keys. <bold>(C)</bold> The predictive performance with top-<italic>l</italic> features concerning <italic>l</italic> in terms of AUROC. <bold>(D)</bold> The predictive performance in terms of AUPRC.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fmicb-13-846915-g004.tif"/>
</fig>
<p>Moreover, we illustrate how the feature attention matrix <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> can provide data-driven hints on which embedding features lead to the association. Since a high-value cell in <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> stands for a crucial feature dimension contributing to determine the association between a queried drug and a microbe, the importance <bold>m</bold>(<italic>i</italic>,:) of each feature <bold>f</bold><sub><italic>i</italic></sub> can be measured by the average of value entries in the <italic>i</italic><sub><italic>th</italic></sub> column of <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> (<xref ref-type="fig" rid="F4">Figure 4B</xref>). The importance distribution along with the sorted drug neighbor keys illustrates that highly important features are usually located among the first <italic>l</italic> nearest neighbors. In addition, the predictive performance with top-<italic>l</italic> features concerning <italic>l</italic> is investigated (<xref ref-type="fig" rid="F4">Figures 4C,D</xref>). The number of top features is tuned in the list {1, 6, 11, 16,&#x2026;, 66}. As <italic>l</italic> is increasing to 16, the performance increases sharply in the top-<italic>l</italic> features. When <italic>l</italic> keeps increasing, the performance increases slowly, then even decreases at the greater value of <italic>l</italic>. Again, this illustration demonstrates that the selection of crucial features is significantly better than the set of all features.</p>
<p>In summary, both embedding feature matrix <bold>E</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub>, which is generated by the nearest-neighbor aggregator, and its feature attention matrix <bold>M</bold><sub><italic>k</italic>&#x00D7;<italic>g</italic></sub> provide mensurable clues to the association outcome.</p>
<p>To complement the verification of the interpretability of NNAN, we selected one microbe (i.e., <italic>Staphylococcus aureus</italic>, which is a common causative agent of food poisoning) and one drug (i.e., <italic>Hexyl gallate</italic>, which has strong antimalarial activity against <italic>Plasmodium falciparum</italic>) from Database 1, and there was an association between them (<xref ref-type="bibr" rid="B6">de Lima Pimenta et al., 2013</xref>). We calculated the similarities between drugs using <italic>Hexyl gallate</italic> as the reference molecule and sorted the drugs in order of their similarity to <italic>Hexyl gallate.</italic> Then, we picked the top 10 drugs and checked whether these drugs were associated with <italic>S. aureus</italic> in Database 1. Finally, we found out that 8 out of the top 10 ranked drugs for <italic>Hexyl gallate</italic> are associated with <italic>S. aureus</italic> (<xref ref-type="table" rid="T3">Table 3</xref>).</p>
<table-wrap position="float" id="T3">
<label>TABLE 3</label>
<caption><p>The associations among <italic>Staphylococcus aureus</italic> and ten drugs.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Drug name</td>
<td valign="top" align="center">Rank</td>
<td valign="top" align="center">Association</td>
<td valign="top" align="left">Drug name</td>
<td valign="top" align="center">Rank</td>
<td valign="top" align="center">Association</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>Octyl gallate</italic></td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="left"><italic>Tannic acid</italic></td>
<td valign="top" align="center">6</td>
<td valign="top" align="center">No</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Butyl gallate</italic></td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="left"><italic>Tea tree oil</italic></td>
<td valign="top" align="center">7</td>
<td valign="top" align="center">Yes</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Octadecyl gallate</italic></td>
<td valign="top" align="center">3</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="left"><italic>Pentagalloylglucose</italic></td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">Yes</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Ethyl gallate</italic></td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="left"><italic>4-Ethylcatechol</italic></td>
<td valign="top" align="center">9</td>
<td valign="top" align="center">No</td>
</tr>
<tr>
<td valign="top" align="left"><italic>Methyl gallate</italic></td>
<td valign="top" align="center">5</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="left"><italic>Hamamelitannin</italic></td>
<td valign="top" align="center">10</td>
<td valign="top" align="center">Yes</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>These ten drugs are ranked in order of their similarity to Hexyl gallate.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
<p>From <xref ref-type="table" rid="T3">Table 3</xref>, it is clear that a drug tends to associate with a microbe if it finds its top-<italic>l</italic> near neighbors associate with the same microbe. Moreover, the higher the ranks of its top-<italic>l</italic> near neighbors are, the more possible it is to associate with the microbe. This conclusion would be helpful to screen drug-like molecules.</p>
</sec>
</sec>
<sec id="S4">
<title>Case Study of Novel Prediction</title>
<p>To further confirm the effectiveness of NNAN, we apply our model on one microbe (i.e., <italic>Bacteroides fragilis</italic>) in Database 2 as a case study. Bacteroides are the major human colonic commensal microbes (<xref ref-type="bibr" rid="B16">Kuwahara et al., 2004</xref>). Although <italic>B. fragilis</italic> is rare in comparison to other Bacteroides species, it is the most prevalent clinical isolation of the genus (<xref ref-type="bibr" rid="B30">Salyers, 1984</xref>). Thus, we select <italic>B. fragilis</italic> for the case study experiment.</p>
<p>Nearest neighbor attention network predicts potential associations between drugs and <italic>B. fragilis</italic> by scoring drug&#x2013;microbe pairs (probability). The higher the score, the more likely the association between the drugs and <italic>B. fragilis</italic> exists. In the case study, we verified whether NNAN could find out potential linkages between <italic>B. fragilis</italic> and drugs. According to the ranking of potential DMAs, we validated the top 10, 20, and 50 predicted candidate drugs by a literature search. Eventually, the validation indicates that 10, 17, and 38 out of the top 10, 20, and 50 predicted drugs associated with <italic>B. fragilis</italic> were found by previously published literature. For example, 85% out of the top 20 predicted candidate drugs for <italic>B. fragilis</italic> are validated (<xref ref-type="table" rid="T4">Table 4</xref>); more details can be found in the <xref ref-type="supplementary-material" rid="TS1">Supplementary Material</xref>. These results of prediction demonstrate the ability of NNAN for predicting potential DMAs in practice.</p>
<table-wrap position="float" id="T4">
<label>TABLE 4</label>
<caption><p>Top 20 predicted drugs associated with <italic>Bacteroides fragilis.</italic></p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left">Drug name</td>
<td valign="top" align="left">Evidence</td>
<td valign="top" align="left">Drug name</td>
<td valign="top" align="left">Evidence</td>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>NATEGLINIDE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/17253883/">PMID: 17253883</ext-link></td>
<td valign="top" align="left"><italic>RAMIPRIL</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/31158845/">PMID: 31158845</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>BENAZEPRIL</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/20445573/">PMID: 20445573</ext-link></td>
<td valign="top" align="left"><italic>DILTIAZEM</italic></td>
<td valign="top" align="left">unconfirmed</td>
</tr>
<tr>
<td valign="top" align="left"><italic>VORICONAZOLE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/18034666/">PMID: 18034666</ext-link></td>
<td valign="top" align="left"><italic>CLEMASTINE FUMARATE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/31158845/">PMID: 31158845</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>FEBUXOSTAT</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/18421623/">PMID: 18421623</ext-link></td>
<td valign="top" align="left"><italic>NAPROXEN (+)</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/15058617/">PMID: 15058617</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>LOPERAMIDE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/18192961/">PMID: 18192961</ext-link></td>
<td valign="top" align="left"><italic>ERGONOVINE MALEATE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/17948937/">PMID: 17948937</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>DIGITOXIN</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/1944247/">PMID: 1944247</ext-link></td>
<td valign="top" align="left"><italic>DROSPIRENONE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/28986954/">PMID: 28986954</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>SOTALOL</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/27836712/">PMID: 27836712</ext-link></td>
<td valign="top" align="left"><italic>DICYCLOMINE</italic></td>
<td valign="top" align="left">unconfirmed</td>
</tr>
<tr>
<td valign="top" align="left"><italic>EZETIMIBE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/15871634/">PMID: 15871634</ext-link></td>
<td valign="top" align="left"><italic>PROCARBAZINE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/1316811/">PMID: 1316811</ext-link></td>
</tr>
<tr>
<td valign="top" align="left"><italic>IRBESARTAN</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/12800253/">PMID: 12800253</ext-link></td>
<td valign="top" align="left"><italic>RIZATRIPTAN BENZOATE</italic></td>
<td valign="top" align="left">unconfirmed</td>
</tr>
<tr>
<td valign="top" align="left"><italic>SUMATRIPTAN SUCCINATE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/19925626/">PMID:19925626</ext-link></td>
<td valign="top" align="left"><italic>SULPIRIDE</italic></td>
<td valign="top" align="left"><ext-link ext-link-type="uri" xlink:href="https://pubmed.ncbi.nlm.nih.gov/31158845/">PMID: 31158845</ext-link></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn><p><italic>The first column records the top 10 drugs, while the third column records the top 10&#x2013;20 drugs.</italic></p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="S5" sec-type="conclusion">
<title>Conclusion</title>
<p>This work has introduced NNAN, a deep learning-based bipartite graph network model to predict potential associations between drugs and microbes. NNAN calculates drug similarities using the weights of feature substructures. It provides an embedding representation based on the near neighbor aggregation for drug&#x2013;microbe pairs, to enhance the explanation of DMAs. In addition, the model provides a crucial feature selection attention matrix for achieving more accurate predictions. These three components of NNAN jointly reveal that a drug associated with a specific microbe can always find its top-<italic>l</italic> near neighbors among other drugs that associate with the same microbe. Moreover, they uncover that the higher the ranks of its top-<italic>l</italic> near neighbors are, the more possible it is to associate with the microbe. Under both a cross-validation setting and a realistic potential linkage discovery setting, the empirical comparison of the proposed framework with three state-of-the-art baselines demonstrates that NNAN has significant competitive performance in predicting DMA. In addition, the framework of our model can also be evaluated in more similar biological issues (e.g., miRNA&#x2013;disease, drug&#x2013;target, and compound&#x2013;protein associations prediction). Furthermore, there is still room to improve the model. We can set new experimental scenarios, which identify the DMAs for new drugs or new microbes, and can also integrate more biological databases to enrich the information of DMAs to improve the predictive ability.</p>
</sec>
<sec id="S6" sec-type="data-availability">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="supplementary-material" rid="TS1">Supplementary Material</xref>, further inquiries can be directed to the corresponding author/s.</p>
</sec>
<sec id="S7">
<title>Author Contributions</title>
<p>J-YS and HY designed and supervised the study. BZ engaged in study design, drafted the manuscript, performed experiments, and analyzed data. YX coded and implemented the model, performed experiments. PZ assisted with performing experiments. S-MY assisted with supervising the study. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="conf1" sec-type="COI-statement">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="pudiscl1" sec-type="disclaimer">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
</body>
<back>
<sec id="S8" sec-type="funding-information">
<title>Funding</title>
<p>This work was supported by Shaanxi Provincial Key R&#x0026;D Program, China (No. 2020KW-063, PI: J-YS) and National Natural Science Foundation of China (No. 61872297, PI: J-YS).</p>
</sec>
<sec id="S9" sec-type="supplementary-material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmicb.2022.846915/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmicb.2022.846915/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table_1.XLSX" id="TS1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aagaard</surname> <given-names>K.</given-names></name> <name><surname>Petrosino</surname> <given-names>J.</given-names></name> <name><surname>Keitel</surname> <given-names>W.</given-names></name> <name><surname>Watson</surname> <given-names>M.</given-names></name> <name><surname>Katancik</surname> <given-names>J.</given-names></name> <name><surname>Garcia</surname> <given-names>N.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>The human microbiome project strategy for comprehensive sampling of the human microbiome and why it matters.</article-title> <source><italic>FASEB J.</italic></source> <volume>27</volume> <fpage>1012</fpage>&#x2013;<lpage>1022</lpage>. <pub-id pub-id-type="doi">10.1096/fj.12-220806</pub-id> <pub-id pub-id-type="pmid">23165986</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>215</volume> <fpage>403</fpage>&#x2013;<lpage>410</lpage>.</citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bang</surname> <given-names>S.</given-names></name> <name><surname>Ho Jhee</surname> <given-names>J.</given-names></name> <name><surname>Shin</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Polypharmacy side effect prediction with enhanced interpretability based on graph feature attention network.</article-title> <source><italic>Bioinformatics.</italic></source> <volume>37</volume> <fpage>2955</fpage>&#x2013;<lpage>2962</lpage> <pub-id pub-id-type="doi">10.1093/bioinformatics/btab174</pub-id> <pub-id pub-id-type="pmid">33714994</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cauchy</surname> <given-names>A.-L.</given-names></name></person-group> (<year>2009</year>). <source><italic>ANALYSE MATHMATIQUE. M&#x00C8;thodc g&#x00C8;n&#x00C8;rale pour la r&#x00C8;solution des Syst&#x00CB;mes d&#x2019;&#x00C8;quations Simultan&#x00C8;es.</italic></source> <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cover</surname> <given-names>T. M.</given-names></name> <name><surname>Hart</surname> <given-names>P. E.</given-names></name></person-group> (<year>1967</year>). <article-title>Nearest neighbor pattern classification.</article-title> <source><italic>IEEE Trans. Inf. Theory</italic></source> <volume>13</volume> <fpage>21</fpage>&#x2013;<lpage>27</lpage>.</citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>de Lima Pimenta</surname> <given-names>A.</given-names></name> <name><surname>Chiaradia-Delatorre</surname> <given-names>L. D.</given-names></name> <name><surname>Mascarello</surname> <given-names>A.</given-names></name> <name><surname>de Oliveira</surname> <given-names>K. A.</given-names></name> <name><surname>Leal</surname> <given-names>P. C.</given-names></name> <name><surname>Yunes</surname> <given-names>R. A.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Synthetic organic compounds with potential for bacterial biofilm inhibition, a path for the identification of compounds interfering with quorum sensing.</article-title> <source><italic>Int. J. Antimicrob. Agents</italic></source> <volume>42</volume> <fpage>519</fpage>&#x2013;<lpage>523</lpage>. <pub-id pub-id-type="doi">10.1016/j.ijantimicag.2013.07.006</pub-id> <pub-id pub-id-type="pmid">24016798</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DiMasi</surname> <given-names>J. A.</given-names></name> <name><surname>Faden</surname> <given-names>L. B.</given-names></name></person-group> (<year>2011</year>). <article-title>Competitiveness in follow-on drug R&#x0026;D: a race or imitation?</article-title> <source><italic>Nat. Rev. Drug Discov.</italic></source> <volume>10</volume> <fpage>23</fpage>&#x2013;<lpage>27</lpage>. <pub-id pub-id-type="doi">10.1038/nrd3296</pub-id> <pub-id pub-id-type="pmid">21151030</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haiser</surname> <given-names>H. J.</given-names></name> <name><surname>Gootenberg</surname> <given-names>D. B.</given-names></name> <name><surname>Chatman</surname> <given-names>K.</given-names></name> <name><surname>Sirasani</surname> <given-names>G.</given-names></name> <name><surname>Balskus</surname> <given-names>E. P.</given-names></name> <name><surname>Turnbaugh</surname> <given-names>P. J.</given-names></name></person-group> (<year>2013</year>). <article-title>Predicting and manipulating cardiac drug inactivation by the human gut bacterium <italic>Eggerthella lenta</italic>.</article-title> <source><italic>Science</italic></source> <volume>341</volume> <fpage>295</fpage>&#x2013;<lpage>298</lpage>. <pub-id pub-id-type="doi">10.1126/science.1235872</pub-id> <pub-id pub-id-type="pmid">23869020</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>He</surname> <given-names>B. S.</given-names></name> <name><surname>Peng</surname> <given-names>L. H.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name></person-group> (<year>2018</year>). <article-title>Human microbe-disease association prediction with graph regularized non-negative matrix factorization.</article-title> <source><italic>Front. Microbiol.</italic></source> <volume>9</volume>:<issue>2560</issue>. <pub-id pub-id-type="doi">10.3389/fmicb.2018.02560</pub-id> <pub-id pub-id-type="pmid">30443240</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioffe</surname> <given-names>S.</given-names></name></person-group> (<year>2010</year>). &#x201C;<article-title>Improved consistent sampling, weighted minhash and L1 sketching</article-title>,&#x201D; in <source><italic>Proceedings of the 2010 IEEE International Conference on Data Mining</italic></source>, <publisher-loc>Sydney</publisher-loc>, <fpage>246</fpage>&#x2013;<lpage>255</lpage>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaacks</surname> <given-names>L. M.</given-names></name> <name><surname>Vandevijvere</surname> <given-names>S.</given-names></name> <name><surname>Pan</surname> <given-names>A.</given-names></name> <name><surname>McGowan</surname> <given-names>C. J.</given-names></name> <name><surname>Wallace</surname> <given-names>C.</given-names></name> <name><surname>Imamura</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2019</year>). <article-title>The obesity transition: stages of the global epidemic.</article-title> <source><italic>Lancet Diabetes Endocrinol.</italic></source> <volume>7</volume> <fpage>231</fpage>&#x2013;<lpage>240</lpage>. <pub-id pub-id-type="doi">10.1016/S2213-8587(19)30026-9</pub-id> <pub-id pub-id-type="pmid">30704950</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kashyap</surname> <given-names>P. C.</given-names></name> <name><surname>Chia</surname> <given-names>N.</given-names></name> <name><surname>Nelson</surname> <given-names>H.</given-names></name> <name><surname>Segal</surname> <given-names>E.</given-names></name> <name><surname>Elinav</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <article-title>Microbiome at the Frontier of personalized medicine.</article-title> <source><italic>Mayo Clin. Proc.</italic></source> <volume>92</volume> <fpage>1855</fpage>&#x2013;<lpage>1864</lpage>. <pub-id pub-id-type="doi">10.1016/j.mayocp.2017.10.004</pub-id> <pub-id pub-id-type="pmid">29202942</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katz</surname> <given-names>L.</given-names></name></person-group> (<year>1953</year>). <article-title>A new status index derived from sociometric analysis.</article-title> <source><italic>Psychometrika</italic></source> <volume>18</volume> <fpage>39</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1007/bf02289026</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khalili</surname> <given-names>H.</given-names></name> <name><surname>Godwin</surname> <given-names>A.</given-names></name> <name><surname>Choi</surname> <given-names>J. W.</given-names></name> <name><surname>Lever</surname> <given-names>R.</given-names></name> <name><surname>Brocchini</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Comparative binding of disulfide-bridged PEG-Fabs.</article-title> <source><italic>Bioconjug. Chem.</italic></source> <volume>23</volume> <fpage>2262</fpage>&#x2013;<lpage>2277</lpage>. <pub-id pub-id-type="doi">10.1021/bc300372r</pub-id> <pub-id pub-id-type="pmid">22994419</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kipf</surname> <given-names>T. N.</given-names></name> <name><surname>Welling</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>Semi-supervised classification with graph convolutional networks.</article-title> <source><italic>arXiv</italic></source> [<comment>Preprint</comment>] <pub-id pub-id-type="doi">10.48550/arXiv.1609.02907</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kuwahara</surname> <given-names>T.</given-names></name> <name><surname>Yamashita</surname> <given-names>A.</given-names></name> <name><surname>Hirakawa</surname> <given-names>H.</given-names></name> <name><surname>Nakayama</surname> <given-names>H.</given-names></name> <name><surname>Toh</surname> <given-names>H.</given-names></name> <name><surname>Okada</surname> <given-names>N.</given-names></name><etal/></person-group> (<year>2004</year>). <article-title>Genomic analysis of <italic>Bacteroides fragilis</italic> reveals extensive DNA inversions regulating cell surface adaptation.</article-title> <source><italic>Proc. Natl. Acad. Sci. U S A.</italic></source> <volume>101</volume> <fpage>14919</fpage>&#x2013;<lpage>14924</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0404172101</pub-id> <pub-id pub-id-type="pmid">15466707</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><collab>Landrum.</collab> (<year>2010</year>). <source><italic>RDKit: Open-Source Cheminformatics. Release 2014.03.1.</italic></source></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>F.</given-names></name> <name><surname>Dong</surname> <given-names>S.</given-names></name> <name><surname>Leier</surname> <given-names>A.</given-names></name> <name><surname>Han</surname> <given-names>M.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2022</year>). <article-title>Positive-unlabeled learning in bioinformatics and computational biology: a brief review.</article-title> <source><italic>Brief. Bioinform.</italic></source> <volume>23</volume>:<issue>bbab461</issue>. <pub-id pub-id-type="doi">10.1093/bib/bbab461</pub-id> <pub-id pub-id-type="pmid">34729589</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Ning</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>Neural inductive matrix completion with graph convolutional networks for miRNA-disease association prediction.</article-title> <source><italic>Bioinformatics</italic></source> <volume>36</volume> <fpage>2538</fpage>&#x2013;<lpage>2546</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btz965</pub-id> <pub-id pub-id-type="pmid">31904845</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lihong</surname> <given-names>P.</given-names></name> <name><surname>Wang</surname> <given-names>C.</given-names></name> <name><surname>Tian</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>K.</given-names></name></person-group> (<year>2021</year>). <article-title>Finding lncRNA-protein interactions based on deep learning with dual-net neural architecture.</article-title> <source><italic>IEEE/ACM Trans. Comput. Biol. Bioinform.</italic></source> <volume>14</volume>:<issue>1</issue>. <pub-id pub-id-type="doi">10.1109/TCBB.2021.3116232</pub-id> <pub-id pub-id-type="pmid">34587091</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Association mining to identify microbe drug interactions based on heterogeneous network embedding representation.</article-title> <source><italic>IEEE J. Biomed. Health Informatics</italic></source> <volume>25</volume> <fpage>266</fpage>&#x2013;<lpage>275</lpage>. <pub-id pub-id-type="doi">10.1109/JBHI.2020.2998906</pub-id> <pub-id pub-id-type="pmid">32750918</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Kwoh</surname> <given-names>C. K.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name></person-group> (<year>2020a</year>). <article-title>Predicting human microbe-drug associations via graph convolutional network with conditional random field.</article-title> <source><italic>Bioinformatics</italic></source> <volume>36</volume> <fpage>4918</fpage>&#x2013;<lpage>4927</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa598</pub-id> <pub-id pub-id-type="pmid">32597948</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Long</surname> <given-names>Y.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Kwoh</surname> <given-names>C. K.</given-names></name> <name><surname>Luo</surname> <given-names>J.</given-names></name> <name><surname>Li</surname> <given-names>X.</given-names></name></person-group> (<year>2020b</year>). <article-title>Ensembling graph attention networks for human microbe-drug association prediction.</article-title> <source><italic>Bioinformatics</italic></source> <volume>36(Suppl_2)</volume> <fpage>i779</fpage>&#x2013;<lpage>i786</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa891</pub-id> <pub-id pub-id-type="pmid">33381844</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lynch</surname> <given-names>S. V.</given-names></name> <name><surname>Pedersen</surname> <given-names>O.</given-names></name></person-group> (<year>2016</year>). <article-title>The human intestinal microbiome in health and disease.</article-title> <source><italic>N. Engl. J. Med.</italic></source> <volume>375</volume> <fpage>2369</fpage>&#x2013;<lpage>2379</lpage>.</citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>L.</given-names></name> <name><surname>Shen</surname> <given-names>L.</given-names></name> <name><surname>Liao</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>G.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>RNMFMDA: a microbe-disease association identification method based on reliable negative sample selection and logistic matrix factorization with neighborhood regularization.</article-title> <source><italic>Front. Microbiol.</italic></source> <volume>11</volume>:<issue>592430</issue>. <pub-id pub-id-type="doi">10.3389/fmicb.2020.592430</pub-id> <pub-id pub-id-type="pmid">33193260</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>L. H.</given-names></name> <name><surname>Yin</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>M. X.</given-names></name> <name><surname>Zhao</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Human microbe-disease association prediction based on adaptive boosting.</article-title> <source><italic>Front. Microbiol.</italic></source> <volume>9</volume>:<issue>2440</issue>. <pub-id pub-id-type="doi">10.3389/fmicb.2018.02440</pub-id> <pub-id pub-id-type="pmid">30356751</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Riniker</surname> <given-names>S.</given-names></name> <name><surname>Landrum</surname> <given-names>G. A.</given-names></name></person-group> (<year>2013</year>). <article-title>Similarity maps &#x2013; a visualization strategy for molecular fingerprints and machine-learning methods.</article-title> <source><italic>J. Cheminform.</italic></source> <volume>5</volume>:<issue>43</issue>. <pub-id pub-id-type="doi">10.1186/1758-2946-5-43</pub-id> <pub-id pub-id-type="pmid">24063533</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogers</surname> <given-names>D.</given-names></name> <name><surname>Hahn</surname> <given-names>M.</given-names></name></person-group> (<year>2010</year>). <article-title>Extended-connectivity fingerprints.</article-title> <source><italic>J. Chem. Inf. Model.</italic></source> <volume>50</volume> <fpage>742</fpage>&#x2013;<lpage>754</lpage>. <pub-id pub-id-type="doi">10.1021/ci100050t</pub-id> <pub-id pub-id-type="pmid">20426451</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogers</surname> <given-names>D. J.</given-names></name> <name><surname>Tanimoto</surname> <given-names>T. T. A.</given-names></name></person-group> (<year>1960</year>). <article-title>Computer program for classifying plants.</article-title> <source><italic>Science</italic></source> <volume>132</volume> <fpage>1115</fpage>&#x2013;<lpage>1118</lpage>. <pub-id pub-id-type="doi">10.1126/science.132.3434.1115</pub-id> <pub-id pub-id-type="pmid">17790723</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salyers</surname> <given-names>A. A.</given-names></name></person-group> (<year>1984</year>). <article-title><italic>Bacteroides</italic> of the human lower intestinal tract.</article-title> <source><italic>Annu. Rev. Microbiol.</italic></source> <volume>38</volume> <fpage>293</fpage>&#x2013;<lpage>313</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.mi.38.100184.001453</pub-id> <pub-id pub-id-type="pmid">6388494</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schwabe</surname> <given-names>R. F.</given-names></name> <name><surname>Jobin</surname> <given-names>C.</given-names></name></person-group> (<year>2013</year>). <article-title>The microbiome and cancer.</article-title> <source><italic>Nat. Rev. Cancer.</italic></source> <volume>13</volume> <fpage>800</fpage>&#x2013;<lpage>812</lpage>.</citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Smith</surname> <given-names>T. F.</given-names></name> <name><surname>Waterman</surname> <given-names>M. S.</given-names></name></person-group> (<year>1981</year>). <article-title>Identification of common molecular subsequences.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>147</volume> <fpage>195</fpage>&#x2013;<lpage>197</lpage>. <pub-id pub-id-type="doi">10.1016/0022-2836(81)90087-5</pub-id> <pub-id pub-id-type="pmid">7265238</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sousa</surname> <given-names>T.</given-names></name> <name><surname>Yadav</surname> <given-names>V.</given-names></name> <name><surname>Zann</surname> <given-names>V.</given-names></name> <name><surname>Borde</surname> <given-names>A.</given-names></name> <name><surname>Abrahamsson</surname> <given-names>B.</given-names></name> <name><surname>Basit</surname> <given-names>A. W.</given-names></name></person-group> (<year>2014</year>). <article-title>On the colonic bacterial metabolism of azo-bonded prodrugsof 5-aminosalicylic acid.</article-title> <source><italic>J. Pharm. Sci.</italic></source> <volume>103</volume> <fpage>3171</fpage>&#x2013;<lpage>3175</lpage>. <pub-id pub-id-type="doi">10.1002/jps.24103</pub-id> <pub-id pub-id-type="pmid">25091594</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sun</surname> <given-names>Y. Z.</given-names></name> <name><surname>Zhang</surname> <given-names>D. H.</given-names></name> <name><surname>Cai</surname> <given-names>S. B.</given-names></name> <name><surname>Ming</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>J. Q.</given-names></name> <name><surname>Chen</surname> <given-names>X. M. D. A. D.</given-names></name></person-group> (<year>2018</year>). <article-title>A special resource for microbe-drug associations.</article-title> <source><italic>Front. Cell. Infect. Microbiol.</italic></source> <volume>8</volume>:<issue>424</issue>. <pub-id pub-id-type="doi">10.3389/fcimb.2018.00424</pub-id> <pub-id pub-id-type="pmid">30581775</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Turnbaugh</surname> <given-names>P. J.</given-names></name> <name><surname>Ley</surname> <given-names>R. E.</given-names></name> <name><surname>Hamady</surname> <given-names>M.</given-names></name> <name><surname>Fraser-Liggett</surname> <given-names>C. M.</given-names></name> <name><surname>Knight</surname> <given-names>R.</given-names></name> <name><surname>Gordon</surname> <given-names>J. I.</given-names></name></person-group> (<year>2007</year>). <article-title>The human microbiome project.</article-title> <source><italic>Nature</italic></source> <volume>449</volume> <fpage>804</fpage>&#x2013;<lpage>810</lpage>.</citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Velickovic</surname> <given-names>P.</given-names></name> <name><surname>Cucurull</surname> <given-names>G.</given-names></name> <name><surname>Casanova</surname> <given-names>A.</given-names></name> <name><surname>Romero</surname> <given-names>A.</given-names></name> <name><surname>Li&#x00F2;</surname> <given-names>P.</given-names></name> <name><surname>Bengio</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Graph attention networks.</article-title> <source><italic>arXiv</italic></source> [<comment>Preprint</comment>]. <pub-id pub-id-type="doi">10.48550/arXiv.1710.10903</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yamanishi</surname> <given-names>Y.</given-names></name> <name><surname>Araki</surname> <given-names>M.</given-names></name> <name><surname>Gutteridge</surname> <given-names>A.</given-names></name> <name><surname>Honda</surname> <given-names>W.</given-names></name> <name><surname>Kanehisa</surname> <given-names>M.</given-names></name></person-group> (<year>2008</year>). <article-title>Prediction of drug-target interaction networks from the integration of chemical and genomic spaces.</article-title> <source><italic>Bioinformatics</italic></source> <volume>24</volume> <fpage>i232</fpage>&#x2013;<lpage>i240</lpage>.</citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Younossi</surname> <given-names>Z. M.</given-names></name> <name><surname>Koenig</surname> <given-names>A. B.</given-names></name> <name><surname>Abdelatif</surname> <given-names>D.</given-names></name> <name><surname>Fazel</surname> <given-names>Y.</given-names></name> <name><surname>Henry</surname> <given-names>L.</given-names></name> <name><surname>Wymer</surname> <given-names>M.</given-names></name></person-group> (<year>2016</year>). <article-title>Global epidemiology of nonalcoholic fatty liver disease-meta-analytic assessment of prevalence, incidence, and outcomes.</article-title> <source><italic>Hepatology</italic></source> <volume>64</volume> <fpage>73</fpage>&#x2013;<lpage>84</lpage>. <pub-id pub-id-type="doi">10.1002/hep.28431</pub-id> <pub-id pub-id-type="pmid">26707365</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>H.</given-names></name> <name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Shi</surname> <given-names>J. Y.</given-names></name></person-group> (<year>2021a</year>). <article-title>RANEDDI: Relation-aware network embedding for prediction of drug-drug interactions.</article-title> <source><italic>Inf. Sci.</italic></source> <volume>582</volume> <fpage>167</fpage>&#x2013;<lpage>180</lpage>.</citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>F.</given-names></name> <name><surname>Zhao</surname> <given-names>X.</given-names></name> <name><surname>Xiao</surname> <given-names>W.</given-names></name> <name><surname>Zhang</surname> <given-names>W.</given-names></name></person-group> (<year>2021b</year>). <article-title>Predicting drug-disease associations through layer attention graph convolutional network.</article-title> <source><italic>Brief. Bioinform.</italic></source> <volume>22</volume>:<issue>bbaa243</issue>. <pub-id pub-id-type="doi">10.1093/bib/bbaa243</pub-id> <pub-id pub-id-type="pmid">33078832</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>L.</given-names></name> <name><surname>Yang</surname> <given-names>P.</given-names></name> <name><surname>Feng</surname> <given-names>H.</given-names></name> <name><surname>Zhao</surname> <given-names>Q.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name></person-group> (<year>2021</year>). <article-title>Using network distance analysis to predict lncRNA-miRNA Interactions.</article-title> <source><italic>Interdiscip. Sci.</italic></source> <volume>13</volume> <fpage>535</fpage>&#x2013;<lpage>545</lpage>. <pub-id pub-id-type="doi">10.1007/s12539-021-00458-z</pub-id> <pub-id pub-id-type="pmid">34232474</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Guan</surname> <given-names>J.</given-names></name> <name><surname>Zhou</surname> <given-names>S.</given-names></name></person-group> (<year>2021</year>). <article-title>FraGAT: a fragment-oriented multi-scale graph attention model for molecular property prediction.</article-title> <source><italic>Bioinformatics</italic></source> <volume>37</volume> <fpage>2981</fpage>&#x2013;<lpage>2987</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btab195</pub-id> <pub-id pub-id-type="pmid">33769437</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zheng</surname> <given-names>Y.</given-names></name> <name><surname>Ley</surname> <given-names>S. H.</given-names></name> <name><surname>Hu</surname> <given-names>F. B.</given-names></name></person-group> (<year>2018</year>). <article-title>Global aetiology and epidemiology of type 2 diabetes mellitus and its complications.</article-title> <source><italic>Nat. Rev. Endocrinol.</italic></source> <volume>14</volume> <fpage>88</fpage>&#x2013;<lpage>98</lpage>. <pub-id pub-id-type="doi">10.1038/nrendo.2017.151</pub-id> <pub-id pub-id-type="pmid">29219149</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>L.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Tian</surname> <given-names>X.</given-names></name> <name><surname>Peng</surname> <given-names>L.</given-names></name></person-group> (<year>2021</year>). <article-title>LPI-deepGBDT: a multiple-layer deep framework based on gradient boosting decision trees for lncRNA-protein interaction identification.</article-title> <source><italic>BMC Bioinform.</italic></source> <volume>22</volume>:<issue>479</issue>. <pub-id pub-id-type="doi">10.1186/s12859-021-04399-8</pub-id> <pub-id pub-id-type="pmid">34607567</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhu</surname> <given-names>L.</given-names></name> <name><surname>Duan</surname> <given-names>G.</given-names></name> <name><surname>Yan</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). &#x201C;<article-title>Prediction of microbe-drug associations based on KATZ measure</article-title>,&#x201D; in <source><italic>Proceedings of the 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</italic></source>, <publisher-loc>San Diego, CA</publisher-loc>.</citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zimmermann</surname> <given-names>M.</given-names></name> <name><surname>Zimmermann-Kogadeeva</surname> <given-names>M.</given-names></name> <name><surname>Wegmann</surname> <given-names>R.</given-names></name> <name><surname>Goodman</surname> <given-names>A. L.</given-names></name></person-group> (<year>2019b</year>). <article-title>Separating host and microbiome contributions to drug pharmacokinetics and toxicity.</article-title> <source><italic>Science</italic></source> <volume>363</volume>:<issue>eaat9931</issue>. <pub-id pub-id-type="doi">10.1126/science.aat9931</pub-id> <pub-id pub-id-type="pmid">30733391</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zimmermann</surname> <given-names>M.</given-names></name> <name><surname>Zimmermann-Kogadeeva</surname> <given-names>M.</given-names></name> <name><surname>Wegmann</surname> <given-names>R.</given-names></name> <name><surname>Goodman</surname> <given-names>A. L.</given-names></name></person-group> (<year>2019a</year>). <article-title>Mapping human microbiome drug metabolism by gut bacteria and their genes.</article-title> <source><italic>Nature</italic></source> <volume>570</volume> <fpage>462</fpage>&#x2013;<lpage>467</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-019-1291-3</pub-id> <pub-id pub-id-type="pmid">31158845</pub-id></citation></ref>
</ref-list>
</back>
</article>