<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Drug. Discov.</journal-id>
<journal-title>Frontiers in Drug Discovery</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Drug. Discov.</abbrev-journal-title>
<issn pub-type="epub">2674-0338</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">768792</article-id>
<article-id pub-id-type="doi">10.3389/fddsv.2021.768792</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Drug Discovery</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Deep Learning Prediction of Adverse Drug Reactions in Drug Discovery Using Open TG&#x2013;GATEs and FAERS Databases</article-title>
<alt-title alt-title-type="left-running-head">Mohsen et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Deep Learning Prediction of ADRs</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Mohsen</surname>
<given-names>Attayeb</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1267868/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Tripathi</surname>
<given-names>Lokesh P.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/695412/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Mizuguchi</surname>
<given-names>Kenji</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/775939/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Artificial Intelligence Center for Health and Biomedical Research (ArCHER), National Institutes of Biomedical Innovation, Health and Nutrition, <addr-line>Osaka</addr-line>, <country>Japan</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>RIKEN Center for Integrative Medical Sciences, <addr-line>Yokohama</addr-line>, <country>Japan</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Institute for Protein Research, Osaka University, <addr-line>Osaka</addr-line>, <country>Japan</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1312783/overview">L. Michel Espinoza-Fonseca</ext-link>, University of Michigan, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1335229/overview">Eli Fernandez-de Gortari</ext-link>, International Iberian Nanotechnology Laboratory (INL), Portugal</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1469343/overview">Fernando Prieto-Mart&#xed;nez</ext-link>, National Autonomous University of Mexico, Mexico</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Attayeb Mohsen, <email>attayeb@nibiohn.go.jp</email>; Kenji Mizuguchi, <email>kenji@nibiohn.go.jp</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to In silico Methods and Artificial Intelligence for Drug Discovery, a section of the journal Frontiers in Drug Discovery</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>27</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>1</volume>
<elocation-id>768792</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Mohsen, Tripathi and Mizuguchi.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Mohsen, Tripathi and Mizuguchi</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Machine learning techniques are being increasingly used in the analysis of clinical and omics data. This increase is primarily due to the advancements in Artificial intelligence (AI) and the build-up of health-related big data. In this paper we have aimed at estimating the likelihood of adverse drug reactions or events (ADRs) in the course of drug discovery using various machine learning methods. We have also described a novel machine learning-based framework for predicting the likelihood of ADRs. Our framework combines two distinct datasets, drug-induced gene expression profiles from Open TG&#x2013;GATEs (Toxicogenomics Project&#x2013;Genomics Assisted Toxicity Evaluation Systems) and ADR occurrence information from FAERS (FDA [Food and Drug Administration] Adverse Events Reporting System) database, and can be applied to many different ADRs. It incorporates data filtering and cleaning as well as feature selection and hyperparameters fine tuning. Using this framework with Deep Neural Networks (DNN), we built a total of 14 predictive models with a mean validation accuracy of 89.4%, indicating that our approach successfully and consistently predicted ADRs for a wide range of drugs. As case studies, we have investigated the performances of our prediction models in the context of Duodenal ulcer and Hepatitis fulminant, highlighting mechanistic insights into those ADRs. We have generated predictive models to help to assess the likelihood of ADRs in testing novel pharmaceutical compounds. We believe that our findings offer a promising approach for ADR prediction and will be useful for researchers in drug discovery.</p>
</abstract>
<kwd-group>
<kwd>adverse drug reactions</kwd>
<kwd>gene expression profiles</kwd>
<kwd>drug discovery</kwd>
<kwd>deep learning</kwd>
<kwd>prediction</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>An adverse drug reaction (ADR) or event is defined as any unintended or undesired effect of a drug (<xref ref-type="bibr" rid="B21">Katzung et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B9">Coleman and Pontefract, 2016</xref>). ADRs are responsible for a high number of visits to emergency departments and in-hospital admissions. For instance, The Japan Adverse Drug Events (JADE) study reported around 17 adverse drug events per 1,000 patient days; 1.6% were fatal, 4.9% were life-threatening, and 33% were serious (<xref ref-type="bibr" rid="B30">Morimoto et&#x20;al., 2010</xref>). These observations underscore the importance of toxicity assessment of any medication, especially in the early stages of drug discovery.</p>
<p>Machine learning methods can play a significant role in the interpretation of various data types to predict ADRs. These methods utilize multiple kinds of input data, such as chemical structures, gene expressions as well as text mining. These data types are then processed algorithmically with random forest (machine learning) or by an artificial neural network (deep learning) to generate prediction models. (<xref ref-type="bibr" rid="B17">Ho et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B28">Mayr et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B14">Gao et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B43">Zhang et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B10">Dana et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B11">Dey et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B38">Vamathevan et&#x20;al., 2019</xref>).</p>
<p>Deep learning (<xref ref-type="bibr" rid="B39">Wang et&#x20;al., 2020</xref>), a type of machine learning in Artificial intelligence (AI), has emerged as a promising and highly effective approach that can combine and interrogate diverse biological data types to generate new hypotheses. Deep learning is used extensively in the field of drug discovery and drug repurposing; however, its application in ADR prediction using gene expression data are rather limited.</p>
<p>Open TG&#x2013;GATEs (<xref ref-type="bibr" rid="B19">Igarashi et&#x20;al., 2014</xref>) is a large&#x2013;scale toxicogenomics database that collects gene expression profiles of <italic>in vivo</italic> as well as <italic>in&#x20;vitro</italic> samples that have been treated with various drugs. These expression profiles are an outcome of the Japanese Toxicogenomics Project (<xref ref-type="bibr" rid="B37">Uehara et&#x20;al., 2009</xref>), which aimed to build an extensive database of drug toxicities for drug discovery. It also collects physiological, biochemical, and pathological measurements of the treated animals. Similar databases that aim to profile compound toxicities have also been developed (<xref ref-type="bibr" rid="B7">Chen et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B3">Alexander-Dann et&#x20;al., 2018</xref>).</p>
<p>In contrast with other databases, such as (LINCS) (<xref ref-type="bibr" rid="B34">Subramanian et&#x20;al., 2017</xref>), which have been used to predict multiple ADRs in a single study (<xref ref-type="bibr" rid="B40">Wang et&#x20;al., 2016</xref>), Open TG&#x2013;GATEs has been used to investigate individual/specific toxicities (<xref ref-type="bibr" rid="B33">Rueda-Z&#xe1;rate et&#x20;al., 2017</xref>). To the best of our knowledge, no attempts have been made to provide a general framework for predicting multiple ADRs by using Open TG&#x2013;GATEs.</p>
<p>The design of Open TG-Gates has several advantages over the LINCS database, chiefly the inclusion of <italic>in vivo</italic> samples with different doses and durations of administration. Therefore, we designed our analysis to encompass multiple samples with different dosages and duration for each compound, necessitating additional noise-removal steps in the data processing. This study describes our approach to generating deep learning-based, systematic ADR prediction models. This approach combines ADR occurrence data, including frequency details, from the FAERS (FDA Adverse Event Reporting System) database, with the gene expression profiles from Open TG-GATEs. We show how to improve the models&#x2019; performance by applying feature selection and hyperparameter optimization algorithms. The methodologies and models described in our study offer valuable tools for assessing the likelihood of ADRs in the course of drug discovery.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Overview</title>
<p>An overview of this study&#x2019;s methodology is illustrated in <xref ref-type="sec" rid="s9">Supplementary Figures S1, S2</xref>. First, we retrieved the relevant data from the above-described two databases (open TG&#x2013;GATEs and FAERS). Next, we pre-processed the gene expression data to filter out noisy profiles by using a simple classification model, and retained only the significant ADR-drug associations (<italic>p</italic>&#x20;&#x3c; 0.05; Fisher&#x2019;s exact test). Next, gene expression profile datasets were created by assigning positive and negative compounds for each ADR. We then split the datasets into training and validation sets three times. Subsequently, we used the training set data to perform feature selection and built deep neural network models with hyperparameter tuning using the Optuna package (see below). Finally, we evaluated the performances of the individual models on the validation set. We discuss these steps in detail below:</p>
</sec>
<sec id="s2-2">
<title>2.2 Data Retrieval and Processing</title>
<sec id="s2-2-1">
<title>2.2.1 Open TG&#x2013;GATEs Database</title>
<p>We extracted the <italic>in</italic>&#x2013;<italic>vivo</italic> gene expression profiles of rat liver samples from the Open TG&#x2013;GATEs database (<xref ref-type="bibr" rid="B37">Uehara et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B19">Igarashi et&#x20;al., 2014</xref>). We selected the rat <italic>in</italic>&#x2013;<italic>vivo</italic> data for our analysis chiefly because the <italic>in</italic>&#x2013;<italic>vivo</italic> dataset included more compounds and a greater number of time points as compared with the <italic>in</italic>&#x2013;<italic>vitro</italic> data (rat and human). However, our methodology can be easily extended to the other datasets.</p>
<p>This dataset was comprised of single-dose experiments and repeated-dose experiments. Single-dose experiments included administration-to-sacrifice periods of 3, 6, 9, or 24&#xa0;h, whereas, in the repeated dose experiments, drugs were administered to rats once daily for 4, 8, 15, or 29&#xa0;days. In the repeated-dose experiments, all rats were sacrificed 24&#xa0;h after the last dose (<xref ref-type="bibr" rid="B19">Igarashi et&#x20;al., 2014</xref>). In Open TG&#x2013;GATEs, gene expression profiles were measured using Microarray technology (Affymetrix GeneChip).</p>
<p>The Affymetrix CEL files were downloaded from <ext-link ext-link-type="uri" xlink:href="http://toxico.nibiohn.go.jp">http://toxico.nibiohn.go.jp</ext-link>, and were preprocessed using the affy package (<xref ref-type="bibr" rid="B15">Gautier et&#x20;al., 2004</xref>) from R Bioconductor (<ext-link ext-link-type="uri" xlink:href="https://bioconductor.org/">https://bioconductor.org/</ext-link>); Affymetrix Microarray Suite algorithm version 5 (mas5) was applied with the default parameters provided in affy, wherein normalization &#x3d; TRUE. The resulting normalized dataset&#x2013;hereafter referred to as &#x201c;the raw dataset&#x201d;&#x2013;was used for all the subsequent analyses. Next, the fold change values were calculated for each probe set by dividing the raw dataset by the mean intensities of corresponding control samples; these values were then log2 transformed, hereafter referred to as the &#x201c;log2FC dataset.&#x201d;</p>
<p>Since the experimental design included multiple dosages and durations of exposure, the drugs had varied effects on the gene expression profiles. To reduce the noise, we predicted all the samples to be either treated or control using a generalized linear model with Lasso regularization [GLMNET package from R (<xref ref-type="bibr" rid="B13">Friedman et&#x20;al., 2010</xref>)]. We used the whole raw dataset as the training set with a binary classification (treated and control). We fed through all the microarray data of the same duration to a single model, creating one model for each exposure set duration. Next, we estimated the probability of being classified as a treated sample for all the training sets. Only those samples with a probability of higher than 92% were included in our analysis. The remaining samples were considered to fall within the gray zone between treated and control, and they were discarded. We chose a cut-off of 92% because, at this threshold, no control samples were misclassified as treated.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Standardized FAERS Data</title>
<p>FAERS (FDA Adverse Event Reporting System) is &#x201c;a database that collects adverse event reports, medication error reports, and product quality complaints resulting in adverse events that were submitted to FDA&#x201d; (<ext-link ext-link-type="uri" xlink:href="https://open.fda.gov/data/faers/">https://open.fda.gov/data/faers/</ext-link>). However, since the terms used in the FAERS database are left to the reporter to decide, inaccurate descriptions may often be incorporated, such as using general, vague terms to describe adverse events or treatments (<xref ref-type="bibr" rid="B41">Wong et&#x20;al., 2015</xref>). To surmount this issue, we used the portion of the FAERS dataset standardized by <xref ref-type="bibr" rid="B4">Banda et&#x20;al. (2016)</xref>. They had curated and standardized the entries of the FAERS database for 11&#xa0;years (2004&#x2013;2015) following Medical Dictionary for Regulatory Activities (MedDRA) preferred terms (PT) (<xref ref-type="bibr" rid="B42">Wood, 1994</xref>).</p>
<p>We extracted all the compound-ADR combinations (70,553,900) from the total number of reports (4.8 million). Among the difficulties of using the FAERS database in ADRs prediction models is the presence of reports with multiple drugs used (Multipharma), which is expected in patients with chronic diseases. Such cases introduce unreliable associations added to the data noise. To solve this issue, we used only the associations in which the drug was assigned as the primary suspect (PS) (15,377,900). We calculated the number of reports for each compound-adverse drug event combination and calculated the total number of reports of the compound in question and also the total number of reports of the adverse&#x20;event.</p>
<p>We assessed the significance of the compound-ADR associations by one-sided Fisher test (<xref ref-type="bibr" rid="B16">Ghosh, 1988</xref>) using &#x201c;fisher.test&#x201d; function from R with the parameter (alternative &#x3d; &#x201c;greater&#x201d;). This option returns a significant <italic>p</italic>&#x2013;value only in the event of a positive association, in contrast to &#x201c;two.sides&#x201d; test, which assesses both positive and negative associations <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Fisher exact test: a: the number of reports of that the compound cause the ADE, b: the number of reports of the compound that does not report the cause of ADE, c: the number of all positive reports of the ADE for all compounds other than the specific compound, e: the number of all negative reports of all compound other than the specific compound.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">Positive</th>
<th align="center">Negative</th>
<th align="center">Row total</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Compound</td>
<td align="center">a</td>
<td align="center">b</td>
<td align="center">a &#x2b; b</td>
</tr>
<tr>
<td align="left">All other compounds</td>
<td align="center">c</td>
<td align="center">d</td>
<td align="center">c &#x2b; d</td>
</tr>
<tr>
<td align="left">Column total</td>
<td align="center">a &#x2b; c</td>
<td align="center">b &#x2b; d</td>
<td align="center">All reports</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s2-3">
<title>2.3 Model Building and Training</title>
<p>For a given ADR, we first designated the compounds with the most significant associations as positive compounds, (<italic>p</italic>&#x2013;value threshold &#x3c;0.05) and the least significantly associated compounds as negative compounds. In building a predictive model, we evenly balanced the number of positive and negative compounds, and retrieved the associated gene expression profiles (treated samples only as described above).</p>
<p>We then defined the training and validation sets by imposing two criteria: 1) the data-sets were balanced, i.e.,&#x20;the number of positive and negative samples were equal in both the sets, and; 2) no compounds were commonly shared between training and validation. The number of samples associated with individual compounds was highly variable, making it difficult to apply the standard cross-validation approach. To overcome this limitation, we shuffled the compounds between training and validation and sampled various training and validation set configurations. We then selected the most balanced configurations with training: validation ratios close to 80:20.</p>
<p>To prevent information leakage between the validation set and the training set, feature selection was performed using the training set only. Consequently, the validation set was used solely to identify the best performing models.</p>
<p>For feature selection, we used Boruta (<xref ref-type="bibr" rid="B24">Kursa and Rudnicki, 2010</xref>) implementation in Python <ext-link ext-link-type="uri" xlink:href="https://github.com/scikit-learn-contrib/boruta_py">https://github.com/scikit-learn-contrib/boruta_py</ext-link>, that is based on the random forest classifier from scikit-learn (<xref ref-type="bibr" rid="B31">Pedregosa et&#x20;al., 2011</xref>) python package with default parameters. Important features (<xref ref-type="bibr" rid="B18">Huynh-Thu et&#x20;al., 2012</xref>) are the variables (genes in this instance) that are essential to classify the samples as either positive or negative. Utilizing such key features for classification helps minimize data dimensionality. Moreover, these important features (genes) can offer deep insights into the biological phenomenon under study (the pathophysiology of the ADR in this study). To remove the effect of randomness and improve the accuracy of feature selection, Boruta generates additional shadow variables by shuffling the values of the original features, these additional shadow variables are added to the training set before assessing feature importance. Subsequently the importance of all the features is evaluated following the random forest algorithm, and the features with significantly higher significance than the shadow variables are considered important, while those with less importance are ignored. This procedure was repeated 100&#x20;times to detect the important features more accurately (<xref ref-type="bibr" rid="B24">Kursa and Rudnicki, 2010</xref>). Then, we used TensorFlow 2 (<xref ref-type="bibr" rid="B1">Abadi et&#x20;al., 2016</xref>) to construct deep learning models. The input of the deep learning model is constructed of two-dimensional matrix with samples in rows and genes in columns.</p>
<p>Each model consists of three groups of layers, input, output, and hidden layers (<xref ref-type="sec" rid="s9">Supplementary Figures S1, S2</xref>). We applied Optuna (<xref ref-type="bibr" rid="B2">Akiba et&#x20;al., 2019</xref>) for hyperparameters tuning. Optuna uses the trial and error method for optimization, by randomly assigning values to the model hyperparameters from a range of values or choices offered by the user for a pre-determined number of trials. Subsequently, the results of all the trials can be examined to determine the optimal parameters. The parameters that were optimized included (<xref ref-type="table" rid="T2">Table&#x20;2</xref>): &#x201c;depth&#x201d;: corresponds to the number of densely connected layers (number of hidden layers aka DNN depth); the possible values are (1, 2, 5, 10, and 30). &#x201c;width&#x201d;: corresponds to the number of nodes per layer, the possible values were (100, 250, 500, and 700). To reduce the likelihood of over fitting, we used two measures. The first measure &#x201c;drop&#x201d; is to drop some nodes before going to the next layer. It took one of these values (0.2, 0.3, 0.4, and 0.5), (0.2 means 20% of nodes are dropped). The second measure is the introduction of noise: the value of introduced Gaussian noise (0.2, 0.3, 0.4, and 0.5). Other hyperparameters included: &#x201c;activation&#x201d;: corresponds to the activation function of the final layer (output layer); possible values: (&#x201c;sigmoid,&#x201d; &#x201c;linear&#x201d;), &#x201c;learning rate&#x201d; for Adam optimizer (<xref ref-type="bibr" rid="B22">Kingma and Ba, 2014</xref>) was selected among (0.001, 0.0005, and 0.00001). The models with the highest validation set accuracies were chosen.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Optuna hyperparameter choices.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Parameter</th>
<th align="center">Choices</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Depth</td>
<td align="center">1, 2, 5, 10, 30</td>
</tr>
<tr>
<td align="left">Width</td>
<td align="center">100, 250, 500, 700</td>
</tr>
<tr>
<td align="left">Drop percentage</td>
<td align="center">0.2, 0.3, 0.4, 0.7</td>
</tr>
<tr>
<td align="left">Gaussian noise</td>
<td align="center">0.2, 0.3, 0.4, 0.5</td>
</tr>
<tr>
<td align="left">Activation</td>
<td align="center">Sigmoid, linear</td>
</tr>
<tr>
<td align="left">Learning rate</td>
<td align="center">0.001, 0.0005, 0.00001</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The maximum number of epochs was set at 800; however, we applied the &#x201c;early stopping&#x201d; strategy if the accuracy did not improve for 75 epochs. The best models were saved for each Optuna trial. The best parameters are shown in the <xref ref-type="sec" rid="s9">Supplementary Table&#x20;S1</xref>.</p>
<p>We opted not to employ the leave-p-out or k-fold cross-validation protocols to the gene expression samples because: 1) the training and validation sets should be assigned based on compound segregation, i.e.,&#x20;the same compound should not span training and validation sets, and 2) the number of samples differed from compound to compound and thus, creating balanced sets was impossible. We instead adopted the approach of creating three different training and validation sets. We also performed feature selection for each of the set combinations.</p>
<p>Overfitting is a significant issue in machine learning, especially if the amount of data is small. We attempted to minimize overfitting by combining multiple approaches; one such approach is early stopping, which stops the training when the model becomes more specific to the training set. We also used functions in TensorFlow and Keras that help generalize the models by dropping some nodes in the hidden layers by using the drop function and adding Gaussian noise in the deep layers. <xref ref-type="table" rid="T2">Table&#x20;2</xref>
</p>
</sec>
<sec id="s2-4">
<title>2.4 Evaluation and Enrichment Analysis</title>
<p>Model performances were evaluated by testing the performance of the validation set prediction. We estimated the accuracy of the validation set and the area under the ROC (Receiver operating characteristic) curve using the scikit-learn (<xref ref-type="bibr" rid="B31">Pedregosa et&#x20;al., 2011</xref>) package from Python.</p>
<p>TargetMine data analysis platform was used for enrichment analysis and gene annotation (<xref ref-type="bibr" rid="B8">Chen et&#x20;al., 2019</xref>). Databases used for enrichment analysis were KEGG, Reactome, and NCI. <italic>p</italic>&#x2013;<italic>values</italic> were calculated in TargetMine using one-tailed Fisher&#x2019;s exact test. Multiple test correction was set to Benjamini Hochberg, with <italic>p</italic>-<italic>value</italic> significance threshold of&#x20;0.05.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Data Processing</title>
<p>To reduce the data dispersion caused by multiple dose levels&#x20;and administration durations (sacrifice period) in Open TG-GATEs, we filtered out low quality/unsuitable samples. To do that, we used Lasso to classify the samples to either treated or control classes. A total of 6,619 of 10,573&#x20;treated samples, chiefly belonging to the &#x201c;Low&#x201d; dose level category, were classified as controls and eventually excluded. Samples that were correctly classified as treated (3,953 samples) remained for subsequent analysis. (<xref ref-type="table" rid="T3">Table&#x20;3</xref>).</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>The number of Open TG-GATEs samples included in the analysis after clustering using Lasso, dose level and sacrifice period details are&#x20;shown.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="left"/>
<th align="center">Original</th>
<th align="center">Included</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td rowspan="3" align="left">Dose level</td>
<td align="left">Low</td>
<td align="center">3,540</td>
<td align="center">421</td>
</tr>
<tr>
<td align="left">Middle</td>
<td align="center">3,537</td>
<td align="center">1,212</td>
</tr>
<tr>
<td align="left">High</td>
<td align="center">3,496</td>
<td align="center">2,320</td>
</tr>
<tr>
<td rowspan="8" align="left">Sacrifice period</td>
<td align="center">24&#xa0;h</td>
<td align="center">1,408</td>
<td align="center">762</td>
</tr>
<tr>
<td align="center">9&#xa0;h</td>
<td align="center">1,371</td>
<td align="center">556</td>
</tr>
<tr>
<td align="center">6&#xa0;h</td>
<td align="center">1,371</td>
<td align="center">533</td>
</tr>
<tr>
<td align="center">3&#xa0;h</td>
<td align="center">1,368</td>
<td align="center">518</td>
</tr>
<tr>
<td align="center">4&#xa0;days</td>
<td align="center">1,275</td>
<td align="center">254</td>
</tr>
<tr>
<td align="center">8&#xa0;days</td>
<td align="center">1,275</td>
<td align="center">472</td>
</tr>
<tr>
<td align="center">15&#xa0;days</td>
<td align="center">1,266</td>
<td align="center">500</td>
</tr>
<tr>
<td align="center">29&#xa0;days</td>
<td align="center">1,239</td>
<td align="center">358</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-2">
<title>3.2 Model Building and Training</title>
<p>We created a total of 14 models (<xref ref-type="table" rid="T4">Table&#x20;4</xref>). The number of compounds used to create each model ranged from 10 to 18. We equalized the number of positive compounds and negative compounds to generate balanced models.</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>The number of compounds, and samples used to create ADRs prediction models. (AGEP: Acute generalized exanthematous pustulosis, ECG: Electrocardiogram). (&#x2b;): Positive, (&#x2212;): Negative.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">ADR</th>
<th colspan="2" align="center">Drugs</th>
<th colspan="2" align="center">Samples</th>
</tr>
<tr>
<th align="center">&#x2b;</th>
<th align="center">&#x2212;</th>
<th align="center">&#x2b;</th>
<th align="center">&#x2212;</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">AGEP</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">203</td>
<td align="center">87</td>
</tr>
<tr>
<td align="left">Bone marrow failure</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">125</td>
<td align="center">178</td>
</tr>
<tr>
<td align="left">Catatonia</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">169</td>
<td align="center">130</td>
</tr>
<tr>
<td align="left">Duodenal ulcer</td>
<td align="center">6</td>
<td align="center">6</td>
<td align="center">122</td>
<td align="center">178</td>
</tr>
<tr>
<td align="left">ECG qt prolonged</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">105</td>
<td align="center">152</td>
</tr>
<tr>
<td align="left">Febrile neutropenia</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">62</td>
<td align="center">141</td>
</tr>
<tr>
<td align="left">Gastric haemorrhage</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">124</td>
<td align="center">161</td>
</tr>
<tr>
<td align="left">Hepatitis fulminant</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">193</td>
<td align="center">126</td>
</tr>
<tr>
<td align="left">Liver transplant</td>
<td align="center">9</td>
<td align="center">9</td>
<td align="center">297</td>
<td align="center">244</td>
</tr>
<tr>
<td align="left">Lymphocytosis</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">196</td>
<td align="center">128</td>
</tr>
<tr>
<td align="left">Neutropenic sepsis</td>
<td align="center">6</td>
<td align="center">6</td>
<td align="center">89</td>
<td align="center">165</td>
</tr>
<tr>
<td align="left">Optic atrophy</td>
<td align="center">6</td>
<td align="center">6</td>
<td align="center">122</td>
<td align="center">139</td>
</tr>
<tr>
<td align="left">Torsade de pointes</td>
<td align="center">7</td>
<td align="center">7</td>
<td align="center">144</td>
<td align="center">219</td>
</tr>
<tr>
<td align="left">Toxic epidermal necrolysis</td>
<td align="center">5</td>
<td align="center">5</td>
<td align="center">136</td>
<td align="center">99</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-3">
<title>3.3 Model Evaluation</title>
<p>The Average accuracy for all models was 89.94% (minimum &#x3d; 71.42%, and maximum &#x3d; 100%). The validation accuracies of the models are shown in <xref ref-type="fig" rid="F1">Figure&#x20;1</xref>. The area under the Receiver Operating Characteristic (ROC) curve is shown in <xref ref-type="fig" rid="F2">Figure&#x20;2</xref>. Matthews correlation coefficient (<xref ref-type="bibr" rid="B27">Matthews, 1975</xref>) of the predicted classes and the real classes for the modules are shown in <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Validation accuracy of the created models, red diamonds represent the mean.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Area under ROC curves for the created models, red diamonds represent the&#x20;mean.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g002.tif"/>
</fig>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Matthews correlation coefficient, red diamonds represent the&#x20;mean.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g003.tif"/>
</fig>
</sec>
<sec id="s3-4">
<title>3.4 Case Study 1: Duodenal Ulcer</title>
<p>To highlight the effectiveness of our approach, we describe below our observations on the development of the duodenal ulcer ADR prediction model. Duodenal ulcer is a type of peptic ulcer disease characterized by the emergence of open sores on the duodenum&#x2019;s inner lining of the duodenum (<xref ref-type="bibr" rid="B23">Kuna et&#x20;al., 2019</xref>). It is mainly caused by the failure of gastrointestinal system inner coating protection, and the most common causing agents are <italic>Helicobacter Pylori</italic> infection and NSAIDs (Non-steroidal anti-inflammatory drugs). It can lead to serious bleeding or perforation.</p>
<sec id="s3-4-1">
<title>3.4.1 Duodenal Ulcer Model Description</title>
<p>Using the FAERS database, the six most significantly associated drugs with duodenal ulcer were identified using Fisher test (positive drugs: Aspirin, Diclofenac, Ibuprofen, Indomethacin, Meloxicam, Naproxen) and the least associated drugs were also identified (negative drugs: Acetaminophen, Amiodarone, Bortezomib, Carbamazepine, Ciprofloxacin, Cyclophosphamide). Subsequently, the open TG&#x2013;Gates samples of these drugs were used to build the prediction models. Details of the compounds used, the number of the gene expression samples associated with these compounds, and the results of the Fisher test (see Methods) are shown in <xref ref-type="table" rid="T5">Table&#x20;5</xref>. The ROC curves and the area under them for the three models of duodenal ulcer (each trained on a different training set) are shown in <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> and the training curves are shown in <xref ref-type="fig" rid="F5">Figure&#x20;5</xref>. The performance of these models showed that the area under the curve ranges from 0.94 to 0.99. The number of the features (genes) commonly selected among the three duodenal ulcer models were&#x20;108.</p>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Details of the data set for the duodenal ulcer model (compounds, the number of the samples, the <italic>p</italic>-value of the Fisher test and the class in the training or test sets). Positive class compounds are those that can cause doudenal ulcer, while Negative class compounds are controls. Entries were ordered alphabetically.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">Fisher (<italic>p</italic>)</th>
<th align="center">No</th>
<th align="center">Class</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Acetaminophen</td>
<td align="center">4.88&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">50</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Amiodarone</td>
<td align="center">9.35&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">31</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Aspirin</td>
<td align="center">4.67&#x20;&#xd7; 10<sup>&#x2212;233</sup>&#x2009;</td>
<td align="center">45</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Bortezomib</td>
<td align="center">3.57&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">24</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Carbamazepine</td>
<td align="center">9.99&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">45</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Ciprofloxacin</td>
<td align="center">9.67&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">7</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Cyclophosphamide</td>
<td align="center">8.29&#x20;&#xd7; 10<sup>&#x2212;1</sup>&#x2009;</td>
<td align="center">21</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Diclofenac</td>
<td align="center">1.3&#x20;&#xd7; 10<sup>&#x2212;66</sup>&#x2009;</td>
<td align="center">9</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Ibuprofen</td>
<td align="center">1.22&#x20;&#xd7; 10<sup>&#x2212;112</sup>&#x2009;</td>
<td align="center">24</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Indomethacin</td>
<td align="center">5.40&#x20;&#xd7; 10<sup>&#x2212;15</sup>&#x2009;</td>
<td align="center">12</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Meloxicam</td>
<td align="center">2.15&#x20;&#xd7; 10<sup>&#x2212;35</sup>&#x2009;</td>
<td align="center">10</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Naproxen</td>
<td align="center">2.02&#x20;&#xd7; 10<sup>&#x2212;78</sup>&#x2009;</td>
<td align="center">22</td>
<td>Positive</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Area under ROC curves for duodenal ulcer models. Each color corresponds to a different&#x20;model.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g004.tif"/>
</fig>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Training and validation loss plots for Duodenal ulcer models, vertical line marks the early stopping.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g005.tif"/>
</fig>
</sec>
<sec id="s3-4-2">
<title>3.4.2 Enrichment Analysis</title>
<p>Pathway enrichment analysis (see Methods) using the duodenal ulcer-selected features (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>) clearly highlighted the involvement of bleeding cascade and complement function (<xref ref-type="table" rid="T6">Table&#x20;6</xref>).</p>
<table-wrap id="T6" position="float">
<label>TABLE 6</label>
<caption>
<p>The enrichment analysis results of duodenal ulcer model, showing the involvement of both complement and coagulation functions.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Pathway</th>
<th align="center">
<italic>P</italic> value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Complement and coagulation cascades (rno04610)</td>
<td align="center">1.86 &#xd7; 10<sup>&#x2212;11</sup>
</td>
</tr>
<tr>
<td align="left">Regulation of Complement cascade (R-RNO-977606)</td>
<td align="center">1.56 &#xd7; 10<sup>&#x2212;5</sup>
</td>
</tr>
<tr>
<td align="left">Complement cascade (R-RNO-166658)</td>
<td align="center">9.16 &#xd7; 10<sup>&#x2212;5</sup>
</td>
</tr>
<tr>
<td align="left">Pertussis (rno05133)</td>
<td align="center">3.84 &#xd7; 10<sup>&#x2212;3</sup>
</td>
</tr>
<tr>
<td align="left">Fatty acid degradation (rno00071)</td>
<td align="center">4.83 &#xd7; 10<sup>&#x2212;3</sup>
</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The manifestation of a duodenal ulcer activates the complement cascade, which is probably due to the inflammation caused by the acid effect on the intestinal mucosa. Bleeding is also linked with the duodenal ulcer disease. The enrichment of the Fatty acid degradation pathway is consistent with the fact that the majority of the compounds that cause duodenal ulcers are NSAIDs that inhibit Arachidonic acid metabolism, which is a part of the Fatty acid metabolism pathway.</p>
</sec>
</sec>
<sec id="s3-5">
<title>3.5 Case Study 2: Hepatitis Fulminant</title>
<p>From FAERS database, the five most significantly associated drugs with hepatitis fulminant (<xref ref-type="bibr" rid="B29">Morabito and Adebayo, 2014</xref>; <xref ref-type="bibr" rid="B6">Bernal and Wendon, 2013</xref>) were identified using Fisher test (positive drugs: danazol, famotidine, flutamide, mexiletine, ticlopidine) and the least associated drugs were also identified (negative drugs: carbamazepine, ciprofloxacine, ibuprofen, naproxen, simvastatin). Subsequently, the open TG&#x2013;Gates samples of these drugs were used to build the prediction models. Details of the compounds used, the number of the gene expression samples associated with these compounds, and the results of the Fisher test (see Methods) are shown in <xref ref-type="table" rid="T7">Table&#x20;7</xref>. The ROC curves and the area under them for the three models of hepatitis fulminant (each trained on a different training set) are shown in <xref ref-type="fig" rid="F6">Figure&#x20;6</xref> as well as training curves in <xref ref-type="fig" rid="F7">Figure&#x20;7</xref>. The performance of these models showed that the area under the curve ranges from 0.76 to 0.96. The number of the features (genes) commonly selected among the three hepatitis fulminant models were 108. Pathway enrichment analysis using the hepatitis fulminant-selected features (<xref ref-type="sec" rid="s9">Supplementary Table S2</xref>) clearly highlighted the involvement of Activation of NF-kappaB in B&#x20;cells and Ub-specific processing proteases (<xref ref-type="table" rid="T8">Table&#x20;8</xref>), which coincides with the previous publications that suggest NF-kappaB pathway is linked to liver pathologies, including viral hepatitis, fibrosis and liver necrosis (<xref ref-type="bibr" rid="B35">Sun and Karin, 2008</xref>).</p>
<table-wrap id="T7" position="float">
<label>TABLE 7</label>
<caption>
<p>Details of the data set for the hepatits fulminant model (compounds, the number of the samples, the <italic>p</italic>-value of the Fisher test and the class in the training or test sets). Positive class compounds are those that can cause hepatitis fulminant, while Negative class compounds are controls. Entries were ordered alphabetically.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">Fisher (<italic>p</italic>)</th>
<th align="center">No</th>
<th align="center">Class</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Carbamazepine</td>
<td align="center">7.72 &#xd7; 10<sup>&#x2212;2</sup>
</td>
<td align="center">45</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Ciprofloxacin</td>
<td align="center">2.23 &#xd7; 10<sup>&#x2212;1</sup>
</td>
<td align="center">7</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Danazol</td>
<td align="center">5.92 &#xd7; 10<sup>&#x2212;4</sup>
</td>
<td align="center">51</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Famotidine</td>
<td align="center">6.63 &#xd7; 10<sup>&#x2212;16</sup>
</td>
<td align="center">11</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Flutamide</td>
<td align="center">1.02 &#xd7; 10<sup>&#x2212;5</sup>
</td>
<td align="center">45</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Ibuprofen</td>
<td align="center">6.66 &#xd7; 10<sup>&#x2212;1</sup>
</td>
<td align="center">24</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Mexiletine</td>
<td align="center">1.41 &#xd7; 10<sup>&#x2212;5</sup>
</td>
<td align="center">29</td>
<td>Positive</td>
</tr>
<tr>
<td align="left">Naproxen</td>
<td align="center">9.79 &#xd7; 10<sup>&#x2212;1</sup>
</td>
<td align="center">22</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Simvastatin</td>
<td align="center">2.42 &#xd7; 10<sup>&#x2212;1</sup>
</td>
<td align="center">28</td>
<td>Negative</td>
</tr>
<tr>
<td align="left">Ticlopidine</td>
<td align="center">1.78 &#xd7; 10<sup>&#x2212;6</sup>
</td>
<td align="center">57</td>
<td>Positive</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Area under ROC curves for hepatitis fulminant models. Each color corresponds to a different&#x20;model.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g006.tif"/>
</fig>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Training and validation loss plots for hepatitis fulminant models, vertical line marks the early stopping.</p>
</caption>
<graphic xlink:href="fddsv-01-768792-g007.tif"/>
</fig>
<table-wrap id="T8" position="float">
<label>TABLE 8</label>
<caption>
<p>The enrichment analysis results of hepatitis fulminant model (Top pathways).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Pathway</th>
<th align="center">
<italic>P</italic> value</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Biological oxidations (computationally inferred) (R-RNO-211859)</td>
<td align="char" char=".">0.0001398</td>
</tr>
<tr>
<td align="left">Metabolism of xenobiotics by cytochrome P450 (rno00980)</td>
<td align="char" char=".">0.0004514</td>
</tr>
<tr>
<td align="left">Phase II - Conjugation of compounds (computationally inferred) (R-RNO-156580)</td>
<td align="char" char=".">0.0004669</td>
</tr>
<tr>
<td align="left">Metabolism (computationally inferred) (R-RNO-1430728)</td>
<td align="char" char=".">0.0058958</td>
</tr>
<tr>
<td align="left">Activation of NF-kappaB in B&#x20;cells (computationally inferred) (R-RNO-1169091)</td>
<td align="char" char=".">0.0078192</td>
</tr>
<tr>
<td align="left">Ub-specific processing proteases (computationally inferred) (R-RNO-5689880)</td>
<td align="char" char=".">0.0089848</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4 Discussion</title>
<p>We have described a novel approach that combined toxicogenomics gene expression profiles extracted from Open TG-GATEs and ADRs reports extracted from FAERS to predict the likelihood of ADRs. This integration of two highly distinct data types allowed us to predict ADRs successfully. Moreover, it led to creating a novel dataset that associated drug-induced gene expression profiles with&#x20;ADRs.</p>
<p>To overcome the significant challenges in combining the two datasets, we first sought to extract the individual drug-induced gene expression signature from Open TG-GATEs. Next, we extracted the ADR occurrence frequencies for these drugs and estimated their statistical significance to eventually combine the two datasets.</p>
<p>Moreover, due to multiple dose-levels and sacrifice periods and the presence of repeated and single administration events, the drug-induced gene expression profiles were fairly noisy. We generated a simple model to classify all the samples as either control or treated classes using Lasso to filter out this noise. We performed a rigorous statistical assessment to narrow down suitable samples for subsequent analyses (see Methods for details).</p>
<p>Recently, deep learning has gathered an increasing usage in the field of drug discovery (<xref ref-type="bibr" rid="B12">Eduati et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B28">Mayr et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B32">Preuer et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B43">Zhang et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B11">Dey et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B25">Lee and Chen, 2019</xref>). In this study we have used deep learning together with feature selection to reduce the data dimensionality and avoid overfitting due to limited samples.</p>
<p>Previously, <xref ref-type="bibr" rid="B40">Wang et&#x20;al. (2016)</xref> utilized multiple cell lines to develop predictive models for multiple ADRs. Another study (<xref ref-type="bibr" rid="B20">Joseph et&#x20;al., 2013</xref>) demonstrated that blood transcriptomics could be used to examine other organ toxicities. Our results have supported this notion by exhibiting robust prediction models with high accuracy using liver samples. Liver is a vital organ for drug metabolism and receives a significant amount of blood, and hence, it is widely used in drug toxicity studies. Moreover, even in the absence of pathological responses to the compound toxicity, cells still display differences in gene expression profiles.</p>
<p>This study utilized <italic>in vivo</italic> gene expression data in contrast with another study (<xref ref-type="bibr" rid="B40">Wang et&#x20;al., 2016</xref>) that utilized the data from the LINCS database (<xref ref-type="bibr" rid="B34">Subramanian et&#x20;al., 2017</xref>), a collection of <italic>in&#x20;vitro</italic> gene expression profiles from human cell lines. Our approach is easily applicable to other publicly available collections of toxicogenomics data, such as those from Drug Matrix (<xref ref-type="bibr" rid="B36">Svoboda et&#x20;al., 2019</xref>). Another difference is that <xref ref-type="bibr" rid="B40">Wang et&#x20;al. (2016)</xref> combined chemical structure and Gene Ontology (GO) term associations in their models. They also selected only a single gene expression profile to represent the compound effects; in contrast, our method analyzed multiple samples with different doses and durations of each compound&#x2019;s exposure. Hence, our method is better equipped to account for the biological variations that are inherent in drug-induced physiological and phenotypic responses. Indeed, our models performed better than Wang <italic>et&#x20;al.</italic>&#x2019;s gene expression data-only models (<xref ref-type="bibr" rid="B40">Wang et&#x20;al., 2016</xref>). This difference in performances may be attributed to our utilization of multiple samples for each compound.</p>
<p>Studies related to specific ADR or systems examined in this study have been recently published. <xref ref-type="bibr" rid="B26">Liu et&#x20;al. (2020)</xref> investigated the prediction of drug-induced liver injury utilizing data from different sources and showed comparable accuracy to ours. The highest AUC of their created models was around 0.86. However, their approach is specific to drug-induced liver injury. They also utilized chemical structures and protein-related information, among many other data types. Similarly, <xref ref-type="bibr" rid="B5">Ben Guebila and Thiele (2019)</xref> predicted Gastric ulcers using gene expression data from the LINCS database with an AUC of 0.97. They also compared using gene expression alone with adding more information to the same&#x20;model.</p>
<p>Using the Optuna optimization package (<xref ref-type="bibr" rid="B2">Akiba et&#x20;al., 2019</xref>) made our prediction models&#x2019; creation computationally expensive; hence, only a limited number of models were built. However, our approach can help generate models to serve specific applications using other data resources such as DrugMatrix, depending on the user&#x2019;s&#x20;needs.</p>
<p>One of the limitations of this study is the relatively smaller number of samples, and also the various dosages and durations of the exposure to the drugs. The effect of these limitations can be seen in the training plots (<xref ref-type="fig" rid="F5">Figures 5</xref>, <xref ref-type="fig" rid="F7">7</xref>), which show fluctuating curves. Accordingly, a few models display lower correlations with ADR prediction as evidenced from their Matthews correlation coefficient plots (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref>), however, the majority have correlation values greater than&#x20;50%.</p>
<p>In conclusion, we have developed 14 deep learning models to predict adverse drug events utilizing the publicly available Open TG&#x2013;Gates and FAERS databases. These models can be used to examine if a new drug candidate can cause these side effects. Moreover, following the same feature selection and model building and tuning steps, other models can be created for other&#x20;ADRs.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The prediction models and auxiliary information associated&#x20;with this study are available at: <ext-link ext-link-type="uri" xlink:href="https://github.com/attayeb/adr">https://github.com/attayeb/adr</ext-link>
</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>AM, and KM: concept, design and draft writing. AM: data analysis, models building, AM, LT, and KM: results analysis, and manuscript writing. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec sec-type="COI-statement" id="s7">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We would like to gratefully acknowledge Dr. Yoshinobu Igarashi (Laboratory of Toxicogenomics Informatics, NIBIOHN) for his valuable suggestions and comments. The statements made here are solely the responsibility of the authors.</p>
</ack>
<sec id="s9">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fddsv.2021.768792/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fddsv.2021.768792/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Image1.PDF" id="SM2" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Abadi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Agarwal</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Barham</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Brevdo</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Citro</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <source>Tensorflow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. <italic>arXiv preprint</italic> 1603.04467</source>. </citation>
</ref>
<ref id="B2">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Akiba</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sano</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Yanase</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ohta</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Koyama</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Optuna: A next-generation hyperparameter optimization framework. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery and data mining. 2623-2631. doi: 10.1145/3292500.3330701</source>. </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alexander-Dann</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Pruteanu</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Oerton</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Sharma</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Berindan-Neagoe</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>M&#xf3;dos</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Developments in Toxicogenomics: Understanding and Predicting Compound-Induced Toxicity from Gene Expression Data</article-title>. <source>Mol. Omics</source> <volume>14</volume>, <fpage>218</fpage>&#x2013;<lpage>236</lpage>. <pub-id pub-id-type="doi">10.1039/c8mo00042e</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Banda</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Evans</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Vanguri</surname>
<given-names>R. S.</given-names>
</name>
<name>
<surname>Tatonetti</surname>
<given-names>N. P.</given-names>
</name>
<name>
<surname>Ryan</surname>
<given-names>P. B.</given-names>
</name>
<name>
<surname>Shah</surname>
<given-names>N. H.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>A Curated and Standardized Adverse Drug Event Resource to Accelerate Drug Safety Research</article-title>. <source>Sci. Data</source> <volume>3</volume>. <pub-id pub-id-type="doi">10.1038/sdata.2016.26</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ben Guebila</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Thiele</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Predicting Gastrointestinal Drug Effects Using Contextualized Metabolic Models</article-title>. <source>Plos Comput. Biol.</source> <volume>15</volume>, <fpage>e1007100</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1007100</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernal</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wendon</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Acute Liver Failure</article-title>. <source>N. Engl. J.&#x20;Med.</source> <volume>369</volume>, <fpage>2525</fpage>&#x2013;<lpage>2534</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Borlak</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>A Decade of Toxicogenomic Research and its Contribution to Toxicological Science</article-title>. <source>Toxicol. Sci.</source> <volume>130</volume>, <fpage>217</fpage>&#x2013;<lpage>228</lpage>. <pub-id pub-id-type="doi">10.1093/toxsci/kfs223</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.-A.</given-names>
</name>
<name>
<surname>Tripathi</surname>
<given-names>L. P.</given-names>
</name>
<name>
<surname>Fujiwara</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kameyama</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Itoh</surname>
<given-names>M. N.</given-names>
</name>
<name>
<surname>Mizuguchi</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>The TargetMine Data Warehouse: Enhancement and Updates</article-title>. <source>Front. Genet.</source> <volume>10</volume>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00934</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Coleman</surname>
<given-names>J.&#x20;J.</given-names>
</name>
<name>
<surname>Pontefract</surname>
<given-names>S. K.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Adverse Drug Reactions</article-title>. <source>Clin. Med.</source> <volume>16</volume>, <fpage>481</fpage>&#x2013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.7861/clinmedicine.16-5-481</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dana</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gadhiya</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>St. Surin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Naaz</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Ali</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Deep Learning in Drug Discovery and Medicine; Scratching the Surface</article-title>. <source>Molecules</source> <volume>23</volume>, <fpage>2384</fpage>. <pub-id pub-id-type="doi">10.3390/molecules23092384</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dey</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Fokoue</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Predicting Adverse Drug Reactions through Interpretable Deep Learning Framework</article-title>. <source>BMC Bioinformatics</source> <volume>19</volume>. <pub-id pub-id-type="doi">10.1186/s12859-018-2544-0</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eduati</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mangravite</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Mangravite</surname>
<given-names>L. M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Bare</surname>
<given-names>J.&#x20;C.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Prediction of Human Population Responses to Toxic Compounds by a Collaborative Competition</article-title>. <source>Nat. Biotechnol.</source> <volume>33</volume>, <fpage>933</fpage>&#x2013;<lpage>940</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3299</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Friedman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hastie</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tibshirani</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Regularization Paths for Generalized Linear Models via Coordinate Descent</article-title>. <source>J.&#x20;Stat. Softw.</source> <volume>33</volume>, <fpage>1</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v033.i01</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gao</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Igata</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Takeuchi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Sato</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Ikegaya</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Machine Learning-Based Prediction of Adverse Drug Effects: An Example of Seizure-Inducing Compounds</article-title>. <source>J.&#x20;Pharmacol. Sci.</source> <volume>133</volume>, <fpage>70</fpage>&#x2013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1016/j.jphs.2017.01.003</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gautier</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cope</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Bolstad</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Irizarry</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>affy--analysis of Affymetrix GeneChip Data at the Probe Level</article-title>. <source>Bioinformatics</source> <volume>20</volume>, <fpage>307</fpage>&#x2013;<lpage>315</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btg405</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Ghosh</surname>
<given-names>J.&#x20;K.</given-names>
</name>
</person-group> (<year>1988</year>). <source>A Discussion on the fisher Exact Test. Statistical Information and Likelihood</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name>, <fpage>321</fpage>&#x2013;<lpage>324</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4612-3894-2_18</pub-id>
<article-title>A Discussion on the Fisher Exact Test</article-title> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ho</surname>
<given-names>T.-B.</given-names>
</name>
<name>
<surname>Le</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tran Thai</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Taewijit</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Data-driven Approach to Detect and Predict Adverse Drug Reactions</article-title>. <source>Cpd</source> <volume>22</volume>, <fpage>3498</fpage>&#x2013;<lpage>3526</lpage>. <pub-id pub-id-type="doi">10.2174/1381612822666160509125047</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huynh-Thu</surname>
<given-names>V. A.</given-names>
</name>
<name>
<surname>Saeys</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wehenkel</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Geurts</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Statistical Interpretation of Machine Learning-Based Feature Importance Scores for Biomarker Discovery</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>1766</fpage>&#x2013;<lpage>1774</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts238</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Igarashi</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Nakatsu</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Yamashita</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ono</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ohno</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Urushidani</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Open TG-GATEs: a Large-Scale Toxicogenomics Database</article-title>. <source>Nucleic Acids Res.</source> <volume>43</volume>, <fpage>D921</fpage>&#x2013;<lpage>D927</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku955</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Joseph</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Umbright</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Sellamuthu</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Blood Transcriptomics: Applications in Toxicology</article-title>. <source>J.&#x20;Appl. Toxicol.</source>, <fpage>a</fpage>&#x2013;<lpage>n</lpage>. <pub-id pub-id-type="doi">10.1002/jat.2861</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Katzung</surname>
<given-names>B. G.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Development &#x26; Regulation of Drugs</article-title>,&#x201d; in <source>Basic &#x26; Clinical Pharmacology</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Katzung</surname>
<given-names>B. G.</given-names>
</name>
<name>
<surname>Masters</surname>
<given-names>S. B.</given-names>
</name>
<name>
<surname>Trevor</surname>
<given-names>A. J.</given-names>
</name>
</person-group> (<publisher-loc>New York</publisher-loc>: <publisher-name>McGraw-Hill</publisher-name>), <fpage>69</fpage>&#x2013;<lpage>78</lpage>. <comment>chap. 5</comment>. </citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kingma</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>Ba</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2014</year>). <source>Adam: A method for stochastic optimization. <italic>arXiv preprint</italic>: 1412.6980.</source> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kuna</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jakab</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Smolic</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Raguz-Lucic</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Vcev</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Smolic</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Peptic Ulcer Disease: A Brief Review of Conventional Therapy and Herbal Treatment Options</article-title>. <source>Jcm</source> <volume>8</volume>, <fpage>179</fpage>. <pub-id pub-id-type="doi">10.3390/jcm8020179</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kursa</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Rudnicki</surname>
<given-names>W. R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Feature Selection with theBorutaPackage</article-title>. <source>J.&#x20;Stat. Soft.</source> <volume>36</volume>. <pub-id pub-id-type="doi">10.18637/jss.v036.i11</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname>
<given-names>C. Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.-P. P.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Machine Learning on Adverse Drug Reactions for Pharmacovigilance</article-title>. <source>Drug Discov. Today</source> <volume>24</volume>, <fpage>1332</fpage>&#x2013;<lpage>1343</lpage>. <pub-id pub-id-type="doi">10.1016/j.drudis.2019.03.003</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Luo</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Weng</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Machine-learning Prediction of Oral Drug-Induced Liver Injury (DILI) via Multiple Features and Endpoints</article-title>. <source>Biomed. Res. Int.</source> <volume>2020</volume>, <fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1155/2020/4795140</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Matthews</surname>
<given-names>B. W.</given-names>
</name>
</person-group> (<year>1975</year>). <article-title>Comparison of the Predicted and Observed Secondary Structure of T4 Phage Lysozyme</article-title>. <source>Biochim. Biophys. Acta (Bba) - Protein Struct.</source> <volume>405</volume>, <fpage>442</fpage>&#x2013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.1016/0005-2795(75)90109-9</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mayr</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Klambauer</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Unterthiner</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>DeepTox: Toxicity Prediction Using Deep Learning</article-title>. <source>Front. Environ. Sci.</source> <volume>3</volume>. <pub-id pub-id-type="doi">10.3389/fenvs.2015.00080</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Mohsen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tripathi</surname>
<given-names>L. P.</given-names>
</name>
<name>
<surname>Mizuguchi</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2020</year>). <source>Deep Learning Prediction of Adverse Drug Reactions Using Open TG-GATEs and FAERS Databases. <italic>arXiv preprint</italic>: 2010.05411</source>. </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morabito</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Adebayo</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Fulminant Hepatitis: Definitions, Causes and Management</article-title>. <source>Health</source>, <fpage>2014</fpage>. <pub-id pub-id-type="doi">10.4236/health.2014.610130</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Morimoto</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sakuma</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Matsui</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kuramoto</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Toshiro</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Murakami</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Incidence of Adverse Drug Events and Medication Errors in japan: the JADE Study</article-title>. <source>J.&#x20;Gen. Intern. Med.</source> <volume>26</volume>, <fpage>148</fpage>&#x2013;<lpage>153</lpage>. <pub-id pub-id-type="doi">10.1007/s11606-010-1518-3</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pedregosa</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Varoquaux</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gramfort</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Michel</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Thirion</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Grisel</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Scikit-learn: Machine Learning in python</article-title>. <source>J.&#x20;Machine Learn. Res.</source> <volume>12</volume>, <fpage>2825</fpage>&#x2013;<lpage>2830</lpage>. </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Preuer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Lewis</surname>
<given-names>R. P. I.</given-names>
</name>
<name>
<surname>Hochreiter</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bender</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bulusu</surname>
<given-names>K. C.</given-names>
</name>
<name>
<surname>Klambauer</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>DeepSynergy: Predicting Anti-cancer Drug Synergy with Deep Learning</article-title>. <source>Bioinformatics</source> <volume>34</volume>, <fpage>1538</fpage>&#x2013;<lpage>1546</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx806</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rueda-Z&#xe1;rate</surname>
<given-names>H. A.</given-names>
</name>
<name>
<surname>Imaz-Rosshandler</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>C&#xe1;rdenas-Ovando</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Castillo-Fern&#xe1;ndez</surname>
<given-names>J.&#x20;E.</given-names>
</name>
<name>
<surname>Noguez-Monroy</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rangel-Escare&#xf1;o</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Computational Toxicogenomics Approach Identifies a List of Highly Hepatotoxic Compounds from a Large Microarray Database</article-title>. <source>PLOS ONE</source> <volume>12</volume>, <fpage>e0176284</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0176284</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Subramanian</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Narayan</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Corsello</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Peck</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Natoli</surname>
<given-names>T. E.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>X.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>A Next Generation Connectivity Map: L1000 Platform and the First 1,000,000 Profiles</article-title>. <source>Cell</source> <volume>171</volume>, <fpage>1437</fpage>&#x2013;<lpage>1452</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2017.10.049</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Karin</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>NF-&#x3ba;B Signaling, Liver Disease and Hepatoprotective Agents</article-title>. <source>Oncogene</source> <volume>27</volume>, <fpage>6228</fpage>&#x2013;<lpage>6244</lpage>. <pub-id pub-id-type="doi">10.1038/onc.2008.300</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Svoboda</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Saddler</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Auerbach</surname>
<given-names>S. S.</given-names>
</name>
</person-group> (<year>2019</year>).<article-title>An Overview of National Toxicology Program&#x27;s Toxicogenomic Applications: DrugMatrix and ToxFX</article-title>. In <conf-name>Advances in Computational Toxicology</conf-name>. <publisher-name>Springer</publisher-name>, <fpage>141</fpage>&#x2013;<lpage>157</lpage>. <pub-id pub-id-type="doi">10.1007/978-3-030-16443-0_8</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Uehara</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ono</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Maruyama</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Kato</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Yamada</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Ohno</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>The Japanese Toxicogenomics Project: Application of Toxicogenomics</article-title>. <source>Mol. Nutr. Food Res.</source> <volume>54</volume>, <fpage>218</fpage>&#x2013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1002/mnfr.200900169</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vamathevan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Czodrowski</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Dunham</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Ferran</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Applications of Machine Learning in Drug Discovery and Development</article-title>. <source>Nat. Rev. Drug Discov.</source> <volume>18</volume>, <fpage>463</fpage>&#x2013;<lpage>477</lpage>. <pub-id pub-id-type="doi">10.1038/s41573-019-0024-5</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pourpanah</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Recent Advances in Deep Learning</article-title>. <source>Int. J.&#x20;Mach. Learn. Cyber.</source> <volume>11</volume>, <fpage>747</fpage>&#x2013;<lpage>750</lpage>. <pub-id pub-id-type="doi">10.1007/s13042-020-01096-5</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>N. R.</given-names>
</name>
<name>
<surname>Ma&#x2019;ayan</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Drug-induced Adverse Events Prediction with the LINCS L1000 Data</article-title>. <source>Bioinformatics</source> <volume>32</volume>, <fpage>2338</fpage>&#x2013;<lpage>2345</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btw168</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wong</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Ho</surname>
<given-names>S. S.</given-names>
</name>
<name>
<surname>Saini</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Hibbs</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Fois</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Standardisation of the FAERS Database: a Systematic Approach to Manually Recoding Drug Name Variants</article-title>. <source>Pharmacoepidemiol. Drug Saf.</source> <volume>24</volume>, <fpage>731</fpage>&#x2013;<lpage>737</lpage>. <pub-id pub-id-type="doi">10.1002/pds.3805</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wood</surname>
<given-names>K. L.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>The Medical Dictionary for Drug Regulatory Affairs (MEDDRA) Project</article-title>. <source>Pharmacoepidem. Drug Safe.</source> <volume>3</volume>, <fpage>7</fpage>&#x2013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.1002/pds.2630030105</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>From Machine Learning to Deep Learning: Progress in Machine Intelligence for Rational Drug Discovery</article-title>. <source>Drug Discov. Today</source> <volume>22</volume>, <fpage>1680</fpage>&#x2013;<lpage>1685</lpage>. <pub-id pub-id-type="doi">10.1016/j.drudis.2017.08.010</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>