<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Mol. Biosci.</journal-id>
<journal-title>Frontiers in Molecular Biosciences</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Mol. Biosci.</abbrev-journal-title>
<issn pub-type="epub">2296-889X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">758480</article-id>
<article-id pub-id-type="doi">10.3389/fmolb.2021.758480</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Molecular Biosciences</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>PreBINDS: An Interactive Web Tool to Create Appropriate Datasets for Predicting Compound&#x2013;Protein Interactions</article-title>
<alt-title alt-title-type="left-running-head">Ikeda et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Preparing Datasets for CPI Prediction</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Ikeda</surname>
<given-names>Kazuyoshi</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1462967/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Doi</surname>
<given-names>Takuo</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1503644/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Ikeda</surname>
<given-names>Masami</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1500973/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Tomii</surname>
<given-names>Kentaro</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<xref ref-type="aff" rid="aff5">
<sup>5</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/979004/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Medicinal Chemistry Applied AI Unit, HPC- and AI-driven Drug Development Platform Division, RIKEN Center for Computational Science, <addr-line>Yokohama</addr-line>, <country>Japan</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Division of Physics for Life Functions, Keio University Faculty of Pharmacy, <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Lifematics Inc., <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff4">
<label>
<sup>4</sup>
</label>Artificial Intelligence Research Center (AIRC), National Institute of Advanced Industrial Science and Technology (AIST), <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<aff id="aff5">
<label>
<sup>5</sup>
</label>AIST-Tokyo Tech Real World Big-Data Computation Open Innovation Laboratory (RWBC-OIL), National Institute of Advanced Industrial Science and Technology (AIST), <addr-line>Tokyo</addr-line>, <country>Japan</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/655018/overview">Masahito Ohue</ext-link>, Tokyo Institute of Technology, Japan</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/44919/overview">Jijun Tang</ext-link>, University of South Carolina, United&#x20;States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/962088/overview">Tunca Dogan</ext-link>, Hacettepe University, Turkey</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Kentaro Tomii, <email>k-tomii@aist.go.jp</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Biological Modeling and Simulation, a section of the journal Frontiers in Molecular Biosciences</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>8</volume>
<elocation-id>758480</elocation-id>
<history>
<date date-type="received">
<day>14</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>15</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Ikeda, Doi, Ikeda and Tomii.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Ikeda, Doi, Ikeda and Tomii</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Given the abundant computational resources and the huge amount of data of compound&#x2013;protein interactions (CPIs), constructing appropriate datasets for learning and evaluating prediction models for CPIs is not always easy. For this study, we have developed a web server to facilitate the development and evaluation of prediction models by providing an appropriate dataset according to the task. Our web server provides an environment and dataset that aid model developers and evaluators in obtaining a suitable dataset for both proteins and compounds, in addition to attributes necessary for deep learning. With the web server interface, users can customize the CPI dataset derived from ChEMBL by setting positive and negative thresholds to be adjusted according to the user&#x2019;s definitions. We have also implemented a function for graphic display of the distribution of activity values in the dataset as a histogram to set appropriate thresholds for positive and negative examples. These functions enable effective development and evaluation of models. Furthermore, users can prepare their task-specific datasets by selecting a set of target proteins based on various criteria such as Pfam families, ChEMBL&#x2019;s classification, and sequence similarities. The accuracy and efficiency of <italic>in silico</italic> screening and drug design using machine learning including deep learning can therefore be improved by facilitating access to an appropriate dataset prepared using our web server (<ext-link ext-link-type="uri" xlink:href="https://binds.lifematics.work/">https://binds.lifematics.work/</ext-link>).</p>
</abstract>
<kwd-group>
<kwd>compound-protein interaction</kwd>
<kwd>CHEMBL</kwd>
<kwd>machine learning</kwd>
<kwd>interactive web server</kwd>
<kwd>deep learning</kwd>
<kwd>datasets</kwd>
</kwd-group>
<contract-sponsor id="cn001">Japan Agency for Medical Research and Development<named-content content-type="fundref-id">10.13039/100009619</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Identification of disease-causing proteins and compounds that act on those diseases is an important starting point in the drug discovery process (<xref ref-type="bibr" rid="B12">Hughes et&#x20;al., 2011</xref>). Over the last two decades, the amounts of compound&#x2013;protein interaction (CPI) data have been increasing rapidly because of advances in experimental techniques such as high-throughput screening (HTS) (<xref ref-type="bibr" rid="B4">Bleicher et&#x20;al., 2003</xref>; <xref ref-type="bibr" rid="B16">Macarron et&#x20;al., 2011</xref>). Considering the social effects of the spread of infectious diseases, as exemplified by COVID-19, early discovery of therapeutic agents is highly anticipated (<xref ref-type="bibr" rid="B6">Cui et&#x20;al., 2019</xref>). Improving drug development efficiency using known CPI data is necessary because it can shorten times to market and reduce&#x20;costs.</p>
<p>Machine learning (ML) methods using CPI data have already been regarded as effective means for the hit-to-lead stage (<xref ref-type="bibr" rid="B10">Ghasemi et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B8">Ferreira and Andricopulo, 2019</xref>). In recent years, the development of artificial intelligence (AI) using deep learning has been remarkable. In fact, AI prediction models have already been applied to various issues; further enhanced efficiency of drug discovery is expected (<xref ref-type="bibr" rid="B27">Tsubaki et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B2">Beker et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B14">Kojima et&#x20;al., 2020</xref>). In the field of ML-based CPI prediction research, some widely used benchmark datasets and development methods have been proposed (<xref ref-type="bibr" rid="B11">He et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B30">Wu et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B23">Rifaioglu et&#x20;al., 2020</xref>, <xref ref-type="bibr" rid="B22">2021</xref>).</p>
<p>The accuracy of prediction models is generally known to depend heavily on the quality and quantity of training data. Nevertheless, it is often not easy to obtain a suitable dataset for creating a reliable prediction model. Particularly in the fields of biochemistry and medicinal chemistry, it is difficult for non-specialists to obtain high-quality CPI data and to distinguish between positive and negative&#x20;cases.</p>
<p>A widely used database of bioactive molecules is ChEMBL. Because the collected data are derived mainly from the literature, they include various CPI data points with different activity types and values (<xref ref-type="bibr" rid="B17">Mendez et&#x20;al., 2019</xref>). Although the ChEMBL interface provides a useful function for searching and extracting downloadable activity data, it is unsuitable for directly generating positive/negative datasets for ML. PubChem provides free access to obtain large amounts of CPI data from results of screening experiments (<xref ref-type="bibr" rid="B13">Kim et&#x20;al., 2021</xref>). However, these data are not provided to users as a binary format that can be used easily for ML. The chemical structures obtained from these databases are provided by the Simplified Molecular Input Line Entry System (SMILES) (<xref ref-type="bibr" rid="B28">Weininger, 1988</xref>) and MDL molfile (<xref ref-type="bibr" rid="B7">Dalby et&#x20;al., 1992</xref>). Also, LIT-PCBA (<xref ref-type="bibr" rid="B26">Tran-Nguyen et&#x20;al., 2020</xref>) was released as a structure-based virtual screening benchmark dataset. It also provided for MOL2 and SMILES formats. Therefore, to use it as input for ML, one must use chemical calculation programs such as CDK (<xref ref-type="bibr" rid="B29">Willighagen et&#x20;al., 2017</xref>)<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>. Then the data must be encoded into physicochemical properties and structural descriptors (fingerprints). As a result, users must bear heavy burdens to prepare suitable CPI datasets for their research.</p>
<p>For this study, we have developed a web server that simplifies creation of CPI datasets for the development and evaluation of prediction models. Because the compound data relies on curated ChEMBL data, it includes high-quality chemical data such as drug-like small molecule compounds. The target data are based on proteins from the UniProt database (<xref ref-type="bibr" rid="B1">Bateman et&#x20;al., 2021</xref>), which is well known as a reviewed protein sequence database. We also provide classifications of target proteins based on the Pfam clan (<xref ref-type="bibr" rid="B18">Mistry et&#x20;al., 2021</xref>), ChEMBL&#x2019;s protein target classification, and sequence similarity. This web server provides a new environment that enables developers to build and evaluate their prediction models more effectively and which might reduce costs of finding drug candidates.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Preparation of Compound Attributes</title>
<p>We used ChEMBL to obtain chemical data of compounds. From ChEMBL release 28, bioactive compounds associated with single protein targets were collected in the molfile format as compound structural data. Next, using Open Babel (<xref ref-type="bibr" rid="B19">O&#x2019;Boyle et&#x20;al., 2011</xref>) implemented in RDKit (ver. 2020.09.3), a library for chemical calculations, the compound structure data was converted to four fingerprints: Extended Connectivity Fingerprints four and 6 (ECFP-4 and ECFP-6) (<xref ref-type="bibr" rid="B24">Rogers and Hahn, 2010</xref>), MACCS keys, and&#x20;FP2.</p>
</sec>
<sec id="s2-2">
<title>2.2 Preparation of CPI Data and Protein Attributes</title>
<p>Based on the ChEMBL activity data, we extracted target proteins with protein IDs (i.e.,&#x20;UniProt accession numbers). We collected only activity data that were assigned pChEMBL values (negative logarithms of activity values in nM units). Next, sequence data were obtained from the UniProt Knowledgebase. Protein families and domains for the target proteins were annotated with the cross-reference information from the Pfam database (release 33.1). Each target sequence was labeled with its accession number. Because the attribute (feature) of a target protein, its Position-Specific Scoring Matrix (PSSM) was calculated using the Blast&#x2b; 2.11.0 program with adjustment (<xref ref-type="bibr" rid="B20">Oda et&#x20;al., 2017</xref>) search against the UniRef90 database (<xref ref-type="bibr" rid="B25">Suzek et&#x20;al., 2007</xref>), which provides hiding of redundant sets of sequences from UniProt.</p>
</sec>
<sec id="s2-3">
<title>2.3 Clustering Methods</title>
<p>To provide a diverse set of compounds, we clustered the compounds using the MiniBatch&#x2013;KMeans method of the scikit-learn program (0.22.1) based on ECFP4 fingerprint similarity. Here, we set <italic>K</italic>&#x20;&#x3d; 10,000. Target proteins were clustered in terms of their sequence similarity. We clustered the collected target proteins using different similarity thresholds (40, 50, 90%) with Cluster database at high identity with tolerance (CD-HIT, v4.8.1; <xref ref-type="bibr" rid="B15">Li and Godzik, 2006</xref>), which is a fast and efficient program for clustering large-scale protein sequences. Next, we pre-computed and registered the clustering data into the internal database. As a result, users can easily and quickly select a set of representative target proteins having sequence identity higher than a specified similarity threshold on the interface.</p>
</sec>
<sec id="s2-4">
<title>2.4 Representative Negative Sample Generation</title>
<p>When selecting negative examples, we used a method that generates a set of negatives that do not resemble the known compounds of a given target. The ChEMBL data are generally known to include more positive data than negative data, which is unbalanced as a training dataset for activity prediction. Therefore, we modified a method for generating reliable negative data (<xref ref-type="bibr" rid="B31">Liu et&#x20;al., 2015</xref>). A credible negative sample is based on the assumption that proteins differing from the known and predicted targets of a particular compound are unlikely to be the targets of that compound. A representative negative set was extracted from a cluster to which positive data do not belong. The negative prediction model was made using support vector machine (SVM). The scikit-learn package sklearn.svm.SVC was used for the SVM calculation.</p>
</sec>
<sec id="s2-5">
<title>2.5 Schema and Data Contents</title>
<p>The database schema comprises 11 tables (<xref ref-type="sec" rid="s10">Supplementary Figure S1</xref>). There are five table groups: ChEMBL derived data (compounds, activities, assays), UniProt target annotation data, Pfam family data, clustering results of proteins, and ligand mapping data. As shown in <xref ref-type="table" rid="T1">Table&#x20;1</xref>, this web server contains 1.33 million compounds (chemical molecules) and 8,341 single protein targets. These protein targets were classified into 2,733 Pfam families, and not all but many of them belong to 635 Pfam clans. UniChem (<xref ref-type="bibr" rid="B5">Chambers et&#x20;al., 2013</xref>) was used to identify the ligands found in the Protein Data Bank (<xref ref-type="bibr" rid="B3">Berman et&#x20;al., 2000</xref>). Currently, 12,980 PDB ligands have been registered in this system.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Data contents.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Type</th>
<th align="center">Count</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Compounds</td>
<td align="center">1,331,700</td>
</tr>
<tr>
<td align="left">Assays</td>
<td align="center">338,454</td>
</tr>
<tr>
<td align="left">Targets</td>
<td align="center">8,341</td>
</tr>
<tr>
<td align="left">Pfam families</td>
<td align="center">2,733</td>
</tr>
<tr>
<td align="left">Pfam clans</td>
<td align="center">635</td>
</tr>
<tr>
<td align="left">CD-HIT clusters (90%)</td>
<td align="center">6,649</td>
</tr>
<tr>
<td align="left">CD-HIT clusters (50%)</td>
<td align="center">4,504</td>
</tr>
<tr>
<td align="left">CD-HIT clusters (40%)</td>
<td align="center">3,890</td>
</tr>
<tr>
<td align="left">PDB ligands</td>
<td align="center">12,980</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and Discussion</title>
<p>The web server was built as a relational database management system using Python 3.8.10 programming language, Django 3.2.2, Vue 2.6.11, and SQLite. It provides a simple and interactive graphical user interface (GUI).</p>
<sec id="s3-1">
<title>3.1 Web Interface</title>
<p>Users can use the interface easily by displaying the top page without logging in (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref>). The interface allows users to search for protein targets and to create datasets for obtaining data. The procedure on this web tool includes the following steps: 1) selection of target proteins (&#x201c;Target Selection&#x201d;), 2) selection of activity types (&#x201c;Types&#x201d;), 3) setting of criteria for activity values (&#x201c;Criteria&#x201d;), 4) selection of attributes (&#x201c;Attributes&#x201d;), 5) generation of negative data samples, and 6) download of output data. The example data button (load example setting) has been implemented. In the load example setting function, the target is set to A<sub>2A</sub> receptor or CDK2. The appropriate target proteins, activity types, activity values and margin, attributes, and negative sample setting are given. With this feature, users can test the web server. We also implemented a download button to provide output data. With this feature, users can readily understand the results from the web server in advance.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Partial view of the web server interface, which is designed to be simple and easy to use, allowing users to select and retrieve compound&#x2013;protein interaction datasets quickly and interactively. Users configures their dataset in five steps: <bold>(A)</bold> Select protein targets (&#x201c;Target Selection&#x201d;), <bold>(B)</bold> Select activity types (&#x201c;Types&#x201d;), <bold>(C)</bold> Set activity data thresholds (&#x201c;Criteria&#x201d;), <bold>(D)</bold> Select compound and target attributes for output (&#x201c;Attributes&#x201d;), and <bold>(E)</bold> Set and generate negative data (&#x201c;Representative Negative Sample Generation&#x201d;). At the top of the interface, the &#x201c;How to use&#x201d; link leads to a brief instruction of this web server. &#x201c;About us&#x201d; and &#x201c;Cite us&#x201d; links respectively lead to information of developers and citations of this web server. &#x201c;Load Example Setting&#x201d; buttons provide sample setting parameters for generating datasets of two representative targets (A2AR and CDK2) and download links to obtain the output files. At the bottom of the interface, users can find the estimated time to complete the submitted dataset generation job in advance.</p>
</caption>
<graphic xlink:href="fmolb-08-758480-g001.tif"/>
</fig>
<sec id="s3-1-1">
<title>3.1.1 Selection of Target Proteins</title>
<p>Users can customize the list of target proteins to generate a CPI dataset (<xref ref-type="fig" rid="F2">Figures 2A&#x2013;C</xref>) based on classification of three types below. The Pfam family is classified based on the similarity of domains, which are functional regions in proteins. The Pfam clan is a group of Pfam families that have been inferred as evolutionarily related. The collected targets were annotated using domain and clan/family data extracted from the Pfam database (ver. 33.1). When creating a CPI dataset in terms of the evolutionary relevance of target proteins, users can select the target protein(s) of interest in the list of Pfam families according to Pfam&#x2019;s classification (<xref ref-type="fig" rid="F2">Figure&#x20;2A</xref>). When creating a CPI dataset in terms of biological function and pharmacology, users can select the target protein(s) of interest according to the ChEMBL target classification. The classification relies on the ChEMBL&#x2019;s protein target classification, which is a hierarchical classification and manually defined by experts in the field of drug discovery based on protein functions, enzymatic reactions, pharmacological actions, and so on (<xref ref-type="bibr" rid="B9">Gaulton et&#x20;al., 2012</xref>). In addition, users can use pre-computed CD-HIT clustering results to select target proteins by sequence similarity and generation of a list of target proteins (<xref ref-type="fig" rid="F2">Figure&#x20;2C</xref>). The resulting target proteins are listed as UniProt accessions. We have also implemented the ability to delete each UniProt accession. With this deletion capability, a user can easily customize the list of UniProt accessions for the target proteins.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Selecting target proteins. <bold>(A)</bold> Selection of protein families according to Pfam clans. A list of protein families is displayed by the selected Pfam clan. Users can select one of the Pfam IDs and press the &#x201c;ADD&#x201d; button to provide the corresponding protein accessions to the form at the bottom. <bold>(B)</bold> Selection of target proteins according to ChEMBL target classification. <bold>(C)</bold> Selection of target proteins by sequence similarity based on pre-computed CD-HIT clustering results. The user can select the sequence similarity threshold for clustering from 90, 50, and 40% to provide the corresponding protein accessions.</p>
</caption>
<graphic xlink:href="fmolb-08-758480-g002.tif"/>
</fig>
</sec>
<sec id="s3-1-2">
<title>3.1.2 Selection of Activity Types</title>
<p>Users can customize datasets by specifying activity types and the range of activity values on the interface (<xref ref-type="fig" rid="F3">Figures 3A&#x2013;C</xref>). Users select a single or multiple activity type(s) among EC<sub>50</sub>, IC<sub>50</sub>, AC<sub>50</sub>, Kd, Ki, and potency for the activity data associated with the selected targets (<xref ref-type="fig" rid="F3">Figure&#x20;3A</xref>). In addition, the "All&#x201d; button allows the user to select all options easily.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Customizing compound&#x2013;protein interaction datasets. Activity data are filtered by activity types and a range of activity values. <bold>(A)</bold> Selection of activity types from EC<sub>50</sub>, IC<sub>50</sub>, AC<sub>50</sub>, Kd, Ki, and potency. <bold>(B)</bold> Setting thresholds for negative logarithms of activity values in nM units. Users can easily set thresholds for positive and negative data using a slide bar if selecting &#x201c;&#x2212;Log&#x201d; of radio button. The graphical display of the histogram of the activity values and the counter of the number of activity values help the user to set the appropriate threshold values. <bold>(C)</bold> Users can directly input concentration thresholds for activity values (nM) of positive and negative data if checking &#x201c;Conc&#x201d; of radio button. <bold>(D)</bold> Selection of compound and protein attributes for output. Users can select compound attributes from ECFP4, ECFP6, FP2, and MACCS keys (Canonical SMILES by default), as well as protein attributes as PSSM (protein sequence by default). <bold>(E)</bold> Generation of representative negative samples. Users can customize the amount of negative data for output. This function allows users to select the amount of negative sample generated from 1, 3, and 5&#x20;times the amount of the positive sample.</p>
</caption>
<graphic xlink:href="fmolb-08-758480-g003.tif"/>
</fig>
</sec>
<sec id="s3-1-3">
<title>3.1.3 Setting of Criteria for Activity Values</title>
<p>By inputting the upper and lower thresholds of the activity value, positive and negative data are definable. On the interface, the activity value is displayed as the negative logarithm (-Log) (<xref ref-type="fig" rid="F3">Figure&#x20;3B</xref>) or the concentration of dose&#x2013;response experiment (nM) (<xref ref-type="fig" rid="F3">Figure&#x20;3C</xref>). Users can use a slider bar to set their thresholds easily for positive and negative data. In addition, the range of the margin between positive and negative datasets can be set arbitrarily by assignment of upper and lower thresholds of concentration. A counter has also been incorporated, showing the exact number of bioactive data points in the positive and negative datasets. This feature makes it easy for the user to visualize using the histogram with the number of active values for the training set. It is expected to be particularly helpful when adapting to ML. For many targets, setting a low threshold of activity value increases the number of compounds for learning, but it might include non-specific binders. It is generally known that proper thresholds and margins differ depending on targets. Therefore, this function helps moderate a dataset for effective prediction models.</p>
</sec>
<sec id="s3-1-4">
<title>3.1.4 Selection of Attributes for Output</title>
<p>Users choose attributes of proteins and compounds in the CPI data for output (<xref ref-type="fig" rid="F3">Figure&#x20;3D</xref>). As a compound attribute, users can select from a single or multiple molecular fingerprint(s) (ECFP4, ECFP6, FP2, and MACCS key). The canonical SMILES of the selected compounds were outputted without selecting any option. As a target attribute, protein sequence is output by default; PSSM can be additionally selected.</p>
</sec>
<sec id="s3-1-5">
<title>3.1.5 Representative Negative Sample Generation</title>
<p>Users can customize the amount of negative data to output. This function allows the users to choose between 1, 3, and 5&#x20;times the amount of generated negative samples compared to number of positive samples. Details of the procedure are presented in the <italic>Materials and Methods</italic> section.</p>
</sec>
<sec id="s3-1-6">
<title>3.1.6 Download of Output Data</title>
<p>As described above, a non-redundant CPI dataset can be prepared according to attributes and thresholds. Users click the &#x201c;Generate&#x201d; button at the bottom of the top page to output a customized dataset. It might take some time to request generation of the dataset, so users can confirm the selected items and the status on the result page (<xref ref-type="fig" rid="F4">Figure 4</xref>). Finally, after clicking the download button, users can obtain a compressed file formatted dataset that includes protein and compound attributes.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Display of result page. The result page shows the progress of the submitted dataset generation job and the input parameters specified by the user. A &#x201c;Download&#x201d; button appears at the bottom of this page when the job is completed, allowing the user to download the dataset.</p>
</caption>
<graphic xlink:href="fmolb-08-758480-g004.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Conclusion</title>
<p>A web server has been developed for generating datasets of compound&#x2013;protein interactions. This web server can provide a ready-to-use CPI dataset for ML, including deep learning in drug discovery and development. Obtaining the latest compound and activity data from the ChEMBL and other public databases is important for developing accurate prediction models. For this reason, our web server shall be updated regularly. Additional information will be imported promptly from external resources. The web server is expected to be useful for developing and evaluating ML models for predicting protein&#x2013;compound interactions, and for discovery of new bioactive molecules.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://binds.lifematics.work/">https://binds.lifematics.work/</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>KI and KT conceived and designed the project. KT managed and supervised the project. MI regularly planned and managed the discussions necessary for this research. TD coded and developed the web server. KI, MI, and KT drafted the manuscript. All authors have given approval to the final version of the manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This research was partially supported by Platform Project for Supporting Drug Discovery and Life Science Research (Basis for Supporting Innovative Drug Discovery and Life Science Research (BINDS)) from AMED under Grant Number JP21am0101110.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>Author TD was employed by Lifematics&#x20;Inc.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors would like to thank Toshiyuki Oda for his assistance in calculating PSSMs.</p>
</ack>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fmolb.2021.758480/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fmolb.2021.758480/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material>
<label>Supplementary Figure S1</label>
<caption>
<p>Database schema of this web server.</p>
</caption>
</supplementary-material>
<supplementary-material xlink:href="Image1.TIF" id="SM1" mimetype="application/TIF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="http://www.rdkit.org[2020.09.3]">http://www.rdkit.org[2020.09.3]</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bateman</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Orchard</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Magrane</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Agivetova</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ahmad</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>UniProt: The Universal Protein Knowledgebase in 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D480</fpage>&#x2013;<lpage>D489</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa1100</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beker</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Wo&#x142;os</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Szymku&#x107;</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Grzybowski</surname>
<given-names>B. A.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Minimal-uncertainty Prediction of General Drug-Likeness Based on Bayesian Neural Networks</article-title>. <source>Nat. Mach. Intell.</source> <volume>2</volume>, <fpage>457</fpage>&#x2013;<lpage>465</lpage>. <pub-id pub-id-type="doi">10.1038/s42256-020-0209-y</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Berman</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Westbrook</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Gilliland</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Bhat</surname>
<given-names>T. N.</given-names>
</name>
<name>
<surname>Weissig</surname>
<given-names>H.</given-names>
</name>
<etal/>
</person-group> (<year>2000</year>). <article-title>The Protein Data Bank</article-title>. <source>Nucleic Acids Res.</source> <volume>28</volume>, <fpage>235</fpage>&#x2013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.235</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bleicher</surname>
<given-names>K. H.</given-names>
</name>
<name>
<surname>B&#xf6;hm</surname>
<given-names>H.-J.</given-names>
</name>
<name>
<surname>M&#xfc;ller</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Alanine</surname>
<given-names>A. I.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Hit and lead Generation: beyond High-Throughput Screening</article-title>. <source>Nat. Rev. Drug Discov.</source> <volume>2</volume>, <fpage>369</fpage>&#x2013;<lpage>378</lpage>. <pub-id pub-id-type="doi">10.1038/nrd1086</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chambers</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Davies</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gaulton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hersey</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Velankar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Petryszak</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>UniChem: a Unified Chemical Structure Cross-Referencing and Identifier Tracking System</article-title>. <source>J.&#x20;Cheminform</source> <volume>5</volume>, <fpage>3</fpage>. <pub-id pub-id-type="doi">10.1186/1758-2946-5-3</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cui</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Shi</surname>
<given-names>Z.-L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Origin and Evolution of Pathogenic Coronaviruses</article-title>. <source>Nat. Rev. Microbiol.</source> <volume>17</volume>, <fpage>181</fpage>&#x2013;<lpage>192</lpage>. <pub-id pub-id-type="doi">10.1038/s41579-018-0118-9</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dalby</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Nourse</surname>
<given-names>J.&#x20;G.</given-names>
</name>
<name>
<surname>Hounshell</surname>
<given-names>W. D.</given-names>
</name>
<name>
<surname>Gushurst</surname>
<given-names>A. K. I.</given-names>
</name>
<name>
<surname>Grier</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Leland</surname>
<given-names>B. A.</given-names>
</name>
<etal/>
</person-group> (<year>1992</year>). <article-title>Description of Several Chemical Structure File&#x20;Formats Used by Computer Programs Developed at Molecular Design Limited</article-title>. <source>J.&#x20;Chem. Inf. Comput. Sci.</source> <volume>32</volume>, <fpage>244</fpage>&#x2013;<lpage>255</lpage>. <pub-id pub-id-type="doi">10.1021/ci00007a012</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ferreira</surname>
<given-names>L. L. G.</given-names>
</name>
<name>
<surname>Andricopulo</surname>
<given-names>A. D.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>ADMET Modeling Approaches in Drug Discovery</article-title>. <source>Drug Discov. Today</source> <volume>24</volume>, <fpage>1157</fpage>&#x2013;<lpage>1165</lpage>. <pub-id pub-id-type="doi">10.1016/j.drudis.2019.03.015</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gaulton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bellis</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Bento</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Chambers</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Davies</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Hersey</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>ChEMBL: a Large-Scale Bioactivity Database for Drug Discovery</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>, <fpage>D1100</fpage>&#x2013;<lpage>D1107</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr777</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ghasemi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mehridehnavi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>P&#xe9;rez-Garrido</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>P&#xe9;rez-S&#xe1;nchez</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Neural Network and Deep-Learning Algorithms Used in QSAR Studies: Merits and Drawbacks</article-title>. <source>Drug Discov. Today</source> <volume>23</volume>, <fpage>1784</fpage>&#x2013;<lpage>1790</lpage>. <pub-id pub-id-type="doi">10.1016/j.drudis.2018.06.016</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>He</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Heidemeyer</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Ban</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Cherkasov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Ester</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>SimBoost: a Read-Across Approach for Predicting Drug-Target Binding Affinities Using Gradient Boosting Machines</article-title>. <source>J.&#x20;Cheminform</source> <volume>9</volume>. <pub-id pub-id-type="doi">10.1186/s13321-017-0209-z</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hughes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rees</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Kalindjian</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Philpott</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Principles of Early Drug Discovery</article-title>. <source>Br. J.&#x20;Pharmacol.</source> <volume>162</volume>, <fpage>1239</fpage>&#x2013;<lpage>1249</lpage>. <pub-id pub-id-type="doi">10.1111/j.1476-5381.2010.01127.x</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gindulyte</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>He</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>PubChem in 2021: New Data Content and Improved Web Interfaces</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D1388</fpage>&#x2013;<lpage>D1395</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa971</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kojima</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ishida</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ohta</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Iwata</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Honma</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Okuno</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>KGCN: A Graph-Based Deep Learning Framework for Chemical Structures</article-title>. <source>J.&#x20;Cheminform.</source> <volume>12</volume>, <fpage>32</fpage>. <pub-id pub-id-type="doi">10.1186/s13321-020-00435-6</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Godzik</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Cd-hit: A Fast Program for Clustering and Comparing Large Sets of Protein or Nucleotide Sequences</article-title>. <source>Bioinformatics</source> <volume>22</volume>, <fpage>1658</fpage>&#x2013;<lpage>1659</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl158</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Guan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Improving Compound-Protein Interaction Prediction by Building up Highly Credible Negative Samples</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>i221</fpage>&#x2013;<lpage>i229</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btv256</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Macarron</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Banks</surname>
<given-names>M. N.</given-names>
</name>
<name>
<surname>Bojanic</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Burns</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Cirovic</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Garyantes</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Impact of High-Throughput Screening in Biomedical Research</article-title>. <source>Nat. Rev. Drug Discov.</source> <volume>10</volume>, <fpage>188</fpage>&#x2013;<lpage>195</lpage>. <pub-id pub-id-type="doi">10.1038/nrd3368</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mendez</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gaulton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bento</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Chambers</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>De&#xa0;Veij</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>F&#xe9;lix</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>ChEMBL: Towards Direct Deposition of Bioassay Data</article-title>. <source>Nucleic Acids Res.</source> <volume>47</volume>, <fpage>D930</fpage>&#x2013;<lpage>D940</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1075</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mistry</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chuguransky</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Qureshi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Salazar</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Sonnhammer</surname>
<given-names>E. L. L.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Pfam: The Protein Families Database in 2021</article-title>. <source>Nucleic Acids Res.</source> <volume>49</volume>, <fpage>D412</fpage>&#x2013;<lpage>D419</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkaa913</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>O&#x27;Boyle</surname>
<given-names>N. M.</given-names>
</name>
<name>
<surname>Banck</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>James</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Morley</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Vandermeersch</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Hutchison</surname>
<given-names>G. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Open Babel: An Open Chemical Toolbox</article-title>. <source>J.&#x20;Cheminform.</source> <volume>3</volume>, <fpage>33</fpage>. <pub-id pub-id-type="doi">10.1186/1758-2946-3-33</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Oda</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Lim</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Tomii</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Simple Adjustment of the Sequence Weight Algorithm Remarkably Enhances PSI-BLAST Performance</article-title>. <source>BMC Bioinformatics</source> <volume>18</volume>, <fpage>288</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1686-9</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rifaioglu</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Cetin Atalay</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Cansen Kahraman</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Do&#x11f;an</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Atalay</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>MDeePred: Novel Multi-Channel Protein Featurization for Deep Learning-Based Binding Affinity Prediction in Drug Discovery</article-title>. <source>Bioinformatics</source> <volume>37</volume>, <fpage>693</fpage>&#x2013;<lpage>704</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa858</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rifaioglu</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Nalbat</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Atalay</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Cetin-Atalay</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Do&#x11f;an</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Do&#x1e7;an, TDEEPScreen: High Performance Drug-Target Interaction Prediction with Convolutional Neural Networks Using 2-D Structural Compound Representations</article-title>. <source>Chem. Sci.</source> <volume>11</volume>, <fpage>2531</fpage>&#x2013;<lpage>2557</lpage>. <pub-id pub-id-type="doi">10.1039/c9sc03414e</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rogers</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Hahn</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Extended-connectivity Fingerprints</article-title>. <source>J.&#x20;Chem. Inf. Model.</source> <volume>50</volume>, <fpage>742</fpage>&#x2013;<lpage>754</lpage>. <pub-id pub-id-type="doi">10.1021/ci100050t</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Suzek</surname>
<given-names>B. E.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Mcgarvey</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Mazumder</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>C. H.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>UniRef: Comprehensive and Non-Redundant UniProt Reference Clusters</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>1282</fpage>&#x2013;<lpage>1288</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm098</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tran-Nguyen</surname>
<given-names>V.-K.</given-names>
</name>
<name>
<surname>Jacquemard</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Rognan</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>LIT-PCBA: An Unbiased Data Set for Machine Learning and Virtual Screening</article-title>. <source>J.&#x20;Chem. Inf. Model.</source> <volume>60</volume>, <fpage>4263</fpage>&#x2013;<lpage>4273</lpage>. <pub-id pub-id-type="doi">10.1021/acs.jcim.0c00155</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tsubaki</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tomii</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Sese</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Compound-protein Interaction Prediction with End-To-End Learning of Neural Networks for Graphs and Sequences</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>309</fpage>&#x2013;<lpage>318</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty535</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weininger</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>SMILES, a Chemical Language and Information System. 1. Introduction to Methodology and Encoding Rules</article-title>. <source>J.&#x20;Chem. Inf. Model.</source> <volume>28</volume>, <fpage>31</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1021/ci00057a005</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Willighagen</surname>
<given-names>E. L.</given-names>
</name>
<name>
<surname>Mayfield</surname>
<given-names>J.&#x20;W.</given-names>
</name>
<name>
<surname>Alvarsson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Berg</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Carlsson</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Jeliazkova</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>The Chemistry Development Kit (CDK) v2.0: Atom Typing, Depiction, Molecular Formulas, and Substructure Searching</article-title>. <source>J.&#x20;Cheminform</source> <volume>9</volume>, <fpage>33</fpage>. <pub-id pub-id-type="doi">10.1186/s13321-017-0220-4</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Ramsundar</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Feinberg</surname>
<given-names>E. N.</given-names>
</name>
<name>
<surname>Gomes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Geniesse</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pappu</surname>
<given-names>A. S.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>MoleculeNet: A Benchmark for Molecular Machine Learning</article-title>. <source>Chem. Sci.</source> <volume>9</volume>, <fpage>513</fpage>&#x2013;<lpage>530</lpage>. <pub-id pub-id-type="doi">10.1039/c7sc02664a</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>