<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="discussion" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1135320</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2023.1135320</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Opinion</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Guidelines on the performance evaluation of motif recognition methods in bioinformatics</article-title>
<alt-title alt-title-type="left-running-head">Deyneko</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fgene.2023.1135320">10.3389/fgene.2023.1135320</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Deyneko</surname>
<given-names>Igor V.</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2157147/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Laboratory of Functional Genomics</institution>, <institution>&#x41a;.&#x410;. Timiryazev Institute of Plant Physiology RAS</institution>, <addr-line>Moscow</addr-line>, <country>Russia</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/47587/overview">Yuriy L. Orlov</ext-link>, I.M.Sechenov First Moscow State Medical University, Russia</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/31459/overview">Mikhail P. Ponomarenko</ext-link>, Institute of Cytology and Genetics (RAS), Russia</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1154080/overview">Philip Machanick</ext-link>, Rhodes University, South Africa</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/54844/overview">Siegfried Weiss</ext-link>, Helmholtz Association of German Research Centers (HZ), Germany</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Igor V. Deyneko, <email>igor.deyneko@inbox.ru</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>07</day>
<month>02</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1135320</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>01</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Deyneko.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Deyneko</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<kwd-group>
<kwd>cis-regulatory modules</kwd>
<kwd>DNA motif detection</kwd>
<kwd>enhancers</kwd>
<kwd>promoters</kwd>
<kwd>DNA sequence analysis</kwd>
<kwd>gene regulation</kwd>
</kwd-group>
<contract-num rid="cn001">No. 122042700043-9</contract-num>
<contract-sponsor id="cn001">Ministry of Science and Higher Education of the Russian Federation<named-content content-type="fundref-id">10.13039/501100012190</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The accurate discovery of DNA and RNA regulatory motifs and their combinations is still a topic of active research, focusing to date mainly on the analysis of ChIP-Seq data (<xref ref-type="bibr" rid="B12">Kingsley et al., 2019</xref>; <xref ref-type="bibr" rid="B14">Kong et al., 2020</xref>), on gene co-expression analysis (<xref ref-type="bibr" rid="B17">Rouault et al., 2014</xref>; <xref ref-type="bibr" rid="B21">Teague et al., 2021</xref>) and on the general investigation of the properties of binding motifs (<xref ref-type="bibr" rid="B25">Zeitlinger 2020</xref>). Many bioinformatics methods have been (<xref ref-type="bibr" rid="B24">Zambelli et al., 2013</xref>) and are still being developed (<xref ref-type="bibr" rid="B3">Bentsen et al., 2022</xref>; <xref ref-type="bibr" rid="B8">Hammelman et al., 2022</xref>) to improve prediction accuracy and fully address the advantages of novel experimental and computational techniques, such as that based on deep learning (<xref ref-type="bibr" rid="B2">Auslander et al., 2021</xref>).</p>
<p>However, when it comes to the practical use of bioinformatic predictors, a researcher is often puzzled, first by selecting an appropriate bioinformatic program and then by a huge list of predictions that such programs usually produce. Once several programs are used to increase the chances of one at least finding a real functional motif, the list of predictions becomes too long for experimental verification (<xref ref-type="bibr" rid="B5">Deyneko et al., 2016</xref>), even though independently found similar motifs are more likely to be correct and can be given higher priority (<xref ref-type="bibr" rid="B16">Machanick and Kibet, 2017</xref>).</p>
<p>The main problem that complicates the choice of a favorable approach for a specific task is the insufficient number of comparative tests of the published methods, partly due to the difficulty of defining a universal motif assessment approach (<xref ref-type="bibr" rid="B11">Kibet and Machanick, 2016</xref>). The inadequate testing of many newly suggested algorithms has already been discussed (<xref ref-type="bibr" rid="B19">Smith et al., 2013</xref>) and can be summarized as 1) an insufficient and subjective selection of methods for comparison; 2) use of non-common metrics; and 3) use of non-standard datasets.</p>
<p>Nevertheless, many studies that present novel methods for motif detection repeatedly appear without adequate comparative evaluation. The main issues include comparison against no or only a single method, despite several comparable methods existing (<xref ref-type="bibr" rid="B1">Alvarez-Gonzalez and Erill, 2021</xref>; <xref ref-type="bibr" rid="B8">Hammelman et al., 2022</xref>), the use of only one dataset, usually with unknown true positives (<xref ref-type="bibr" rid="B15">Levitsky et al., 2022</xref>), and the use of uncommon statistical metrics (<xref ref-type="bibr" rid="B26">Zhang et al., 2019</xref>). The last can be exemplified with a criterion of the correct prediction&#x2014;if, within the top ten, there is a motif similar (not identical!) to the original, the motif is counted as positively recovered. In real applications, when the target motif is unknown, the reliability of such predictions is far from being experimentally testable. In contrast, there are many methods with well-performed comparisons, including novel deep learning methods (<xref ref-type="bibr" rid="B3">Bentsen et al., 2022</xref>; <xref ref-type="bibr" rid="B9">Iqbal et al., 2022</xref>).</p>
<p>This work is addressed not only to researchers, who may use the presented principles to better reveal the power of the software presented, but also to peer reviewers and journal editorial boards, who may use it as a starting point for their own requirements for software articles. Obviously, comprehensive comparative testing of new methods will not only reveal the best fields of application but, most importantly, will help wet-lab researchers to navigate through bioinformatics topics.</p>
</sec>
<sec id="s2">
<title>2 Guidelines on comparative testing</title>
<p>Overall, the situation can be improved by introducing the following three principles, schematically represented in <xref ref-type="fig" rid="F1">Figure 1</xref>.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Schematic representation of the proposed principles on the selection of methods, metrics, and datasets for comparative testing.</p>
</caption>
<graphic xlink:href="fgene-14-1135320-g001.tif"/>
</fig>
<sec id="s3">
<title>2.1 Selecting methods for comparison</title>
<p>There must be a clear logic as to why specific methods have been selected for benchmarking; any subjective choice of certain programs should be avoided. The easiest and most objective way is to use a review article. The classical examples are the works of <xref ref-type="bibr" rid="B18">Sandve et al. (2007</xref>) and <xref ref-type="bibr" rid="B13">Klepper et al. (2008</xref>), which additionally provide an online system for methods comparison. Other reviews worth noting are <xref ref-type="bibr" rid="B22">Tran and Huang (2014</xref>) and <xref ref-type="bibr" rid="B10">Jayaram et al. (2016</xref>).</p>
<p>Methods based on novel computational principles and/or experimental data are always welcome, provided that their performance is also properly evaluated against &#x201c;old methods.&#x201d; If, by some modification of an input (output), such methods can be adapted for testing, this should be carried out and the methods included in the comparative list. Notwithstanding, the gold standard for comparisons should comprise three methods and preferably five&#x2014;always preferring the most recent.</p>
</sec>
<sec id="s3-1">
<title>2.2 Selecting datasets</title>
<p>In its basic definition, DNA motif detection is a well-defined problem about a dataset of nucleotide sequences&#x2014;either long as promoters or short as ChIP-seq segments. Therefore, it should (almost) always be possible to run a new program on existing data, and there are many such examples (<xref ref-type="bibr" rid="B13">Klepper et al., 2008</xref>; <xref ref-type="bibr" rid="B6">Deyneko et al., 2013</xref>). Thus, the use of common and publicly available datasets should be obligatory. Once a new algorithm requires additional information, such as expression values, genome positioning, and conservation, standard datasets can be complemented with reasonable values required for a correct comparison. This will reveal how a new method works on &#x201c;old data,&#x201d; ensure a fair testing against other methods, and, most importantly, demonstrates the added value of this additional information. For example, if gene expression values are required, a sequence-only dataset can be modified by assigning &#x201c;1s&#x201d; to foreground and &#x201c;0s&#x201d; to background sequences. This will clearly show the performance gain with respect to the use of such additional information.</p>
<p>The use of self-made datasets can only be accepted as complimentary to standard ones. Even if a method is developed to address a particular problem and does not operate promisingly on standard datasets, the results should still be presented. This will clearly show where a method outperforms other methods and on which data it does not, so that an application niche is clearly defined. Authors should not be afraid of a possibly very narrow application field for their research. Instead, a clear definition will help practitioners find and use the appropriate program before they give up in disappointment after a series of unsatisfactory attempts.</p>
<p>In implementing new methods, researchers should also be cautious about integrating multiple steps into one executable. It is certainly very convenient to analyze the raw data in one go, but this will greatly reduce the field of application. For example, giving human gene names as input, instead of promoter sequences, makes it impossible to analyze the genes of mice, plants, or bacteria. Extracting specific genomic regions is today a trivial task, although it may be implemented as an option for convenience.</p>
</sec>
<sec id="s3-2">
<title>2.3 Selecting performance metrics</title>
<p>Methods including ROC curves, false-positives, true-negatives, selectivity and sensitivity, nucleotide correlation coefficients, and positive predictive values are to be used as metrics (<xref ref-type="bibr" rid="B23">Vihinen 2012</xref>; <xref ref-type="bibr" rid="B10">Jayaram et al., 2016</xref>). If a novel method or dataset does not allow standard metrics, others may be used, provided that it is clearly explained why standard metrics are not applicable. One should avoid giving subjective assertions of performance like &#x201c;Identified all 40 conserved modules reported previously&#x201d; without mentioning how many other modules (false positives) were also identified, or referring to the literature as the only measure of correct predictions (<xref ref-type="bibr" rid="B7">El-Kurdi et al., 2020</xref>). Reference to the literature is fully valid and useful, provided that comprehensive statistics are given. It is notable that statistical measures like <italic>p</italic>-values are often method-specific&#x2014;they depend on a method&#x2019;s internal calculations. So, the <italic>p</italic>-values of different methods should be compared with caution.</p>
</sec>
</sec>
<sec id="s4">
<title>3 Good practices in comparative testing</title>
<p>As examples of thorough comparative testing, two programs will be discussed&#x2014;MatrixCatch for recognizing cis-regulatory modules (<xref ref-type="bibr" rid="B6">Deyneko et al., 2013</xref>) and a predictor of acetylcytidine sites in mRNA based on novel deep learning methodology (<xref ref-type="bibr" rid="B9">Iqbal et al., 2022</xref>).</p>
<p>MatrixCatch uses a database of known composite modules as the basis for recognition. Three classes of comparisons were performed: with methods based on the same principle, with statistical methods, and on the recognition of cis-modules on a real dataset. Next, we briefly discuss the three classes of comparative testing and how they align with the suggested guidelines.</p>
<p>At the time of developing MatrixCatch, two other methods&#x2014;based on the same principle of using known examples of composite modules&#x2014;were available. They were compared against the same sequence dataset, with ROC curves as a performance characteristic.</p>
<p>The second type of comparison was performed against statistical methods for motif detection. The difference from the previous comparison is that the motifs and modules are found solely by nucleotide frequency statistics. For such &#x201c;<italic>de novo</italic>" modules, there is no experimental (or any other) evidence for their functionality. The advantage of such methods is their ability to find truly new motifs and modules. In contrast, MatrixCatch uses a library of experimentally verified modules and is therefore limited to its known repertoire. Although these two types of method use different principles, it is important from a practical point of view to know which method(s) provides the best chance of finding real motif(s) and explain, for example, a co-regulation of genes in an RNA-seq experiment. The tests were performed according to the benchmarking presented in a review by <xref ref-type="bibr" rid="B13">Klepper et al. (2008</xref>), which includes six datasets of DNA sequences, nine methods, and several performance characteristics common to all methods (methods based on reviews&#x2014;<xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
<p>Finally, testing was illustrated by the detection of cis-modules on 11 sets of tissue-specific promoters (authors&#x2019; custom data&#x2014;<xref ref-type="fig" rid="F1">Figure 1</xref>). Regulatory elements presumably existing in promoters are unknown, and therefore, measuring such factors as false positives, ROC, or otherwise cannot be calculated. The performance was measured as the specificity of the best module and equal to the ratio of the number of promoters with recognized cis-module in a positive set to the respective number in the negative set (authors&#x2019; custom metrics&#x2014;<xref ref-type="fig" rid="F1">Figure 1</xref>). Such a definition is the most indicative in real applications, where a researcher seeks to identify elements that occur preferentially in the dataset of interest. Moreover, this measure can be applied to all recognition methods despite their different search logics and output formats.</p>
<p>Another example is a method for recognition of N4-acetylcytidine sites in mRNA (<xref ref-type="bibr" rid="B9">Iqbal et al., 2022</xref>) based on novel deep learning methodology. The method was tested on publicly available reference data; ROC, precision&#x2013;recall curves, and accuracy, specificity, and sensitivity measures were used to evaluate the consistency of classification. An interesting point is that the method was compared to the three &#x201c;old-style&#x201d; machine learning methods, including regression and support vector machine. This not only serves as a bridge between new and conventional methods but also shows its advantages over, for example, regression analysis available in most statistical software.</p>
</sec>
<sec sec-type="conclusion" id="s5">
<title>4 Conclusion</title>
<p>The comprehensive testing of novel methods seems to be as laborious as the developing methods themselves and thus requires longer result sections in manuscripts. Publishing an &#x201c;application note&#x201d; or similar with an imposed page limit forces authors to search for a specific dataset or simulation settings for which their method works better than existing ones. This leads to a very subjective presentation and over-optimism in bioinformatics research&#x2014;and disillusion in practice (<xref ref-type="bibr" rid="B4">Boulesteix 2010</xref>). Following the aforementioned guidelines will simplify and unify methods benchmarking designs and will reveal their best application fields. Establishing similar practical recommendations in other subfields of bioinformatics will facilitate application by practitioners and true innovation by bioinformaticians.</p>
</sec>
</body>
<back>
<sec id="s6">
<title>Author contributions</title>
<p>The author confirms being the sole contributor of this work and has approved it for publication.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This work was funded by the Ministry of Science and Higher Education (Theme No. 122042700043-9).</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of interest</title>
<p>The author declares that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alvarez-Gonzalez</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Erill</surname>
<given-names>I.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Design of machine learning models for the prediction of transcription factor binding regions in bacterial DNA</article-title>. <source>Eng. Proc.</source> <volume>7</volume> (<issue>59</issue>), <fpage>7059</fpage>. <pub-id pub-id-type="doi">10.3390/engproc2021007059</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Auslander</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Gussow</surname>
<given-names>A. B.</given-names>
</name>
<name>
<surname>Koonin</surname>
<given-names>E. V.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Incorporating machine learning into established bioinformatics frameworks</article-title>. <source>Int. J. Mol. Sci.</source> <volume>22</volume>, <fpage>2903</fpage>. <pub-id pub-id-type="doi">10.3390/ijms22062903</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bentsen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Heger</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Schultheis</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kuenne</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Looso</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>TF-COMB - discovering grammar of transcription factor binding sites</article-title>. <source>Comput. Struct. Biotechnol. J.</source> <volume>20</volume>, <fpage>4040</fpage>&#x2013;<lpage>4051</lpage>. <pub-id pub-id-type="doi">10.1016/j.csbj.2022.07.025</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boulesteix</surname>
<given-names>A. L.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Over-optimism in bioinformatics research</article-title>. <source>Bioinformatics</source> <volume>26</volume> (<issue>3</issue>), <fpage>437</fpage>&#x2013;<lpage>439</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp648</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deyneko</surname>
<given-names>I. V.</given-names>
</name>
<name>
<surname>Kasnitz</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Leschner</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Weiss</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Composing a tumor specific bacterial promoter</article-title>. <source>PLoS One</source> <volume>11</volume> (<issue>5</issue>), <fpage>e0155338</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0155338</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Deyneko</surname>
<given-names>I. V.</given-names>
</name>
<name>
<surname>Kel</surname>
<given-names>A. E.</given-names>
</name>
<name>
<surname>Kel-Margoulis</surname>
<given-names>O. V.</given-names>
</name>
<name>
<surname>Deineko</surname>
<given-names>E. V.</given-names>
</name>
<name>
<surname>Wingender</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Weiss</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>MatrixCatch--a novel tool for the recognition of composite regulatory elements in promoters</article-title>. <source>BMC Bioinforma.</source> <volume>14</volume>, <fpage>241</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-14-241</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>El-Kurdi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Khalil</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Khazen</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Khoueiry</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>fcScan: a versatile tool to cluster combinations of sites using genomic coordinates</article-title>. <source>BMC Bioinforma.</source> <volume>21</volume> (<issue>1</issue>), <fpage>194</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3536-4</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hammelman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Krismer</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Gifford</surname>
<given-names>D. K.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>spatzie: an R package for identifying significant transcription factor motif co-enrichment from enhancer-promoter interactions</article-title>. <source>Nucleic Acids Res.</source> <volume>50</volume> (<issue>9</issue>), <fpage>e52</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkac036</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Iqbal</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Abbasi</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Bin Heyat</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Akhtar</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Abdelgeliel</surname>
<given-names>A. S.</given-names>
</name>
<name>
<surname>Albogami</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2022</year>). <article-title>Recognition of mRNA N4 acetylcytidine (ac4C) by using non-deep vs. Deep learning</article-title>. <source>Appl. Sci.</source> <volume>12</volume> (<issue>1344</issue>), <fpage>1344</fpage>. <pub-id pub-id-type="doi">10.3390/app12031344</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jayaram</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Usvyat</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ac</surname>
<given-names>R. M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Evaluating tools for transcription factor binding site prediction</article-title>. <source>BMC Bioinforma.</source> <volume>17</volume> (<issue>1</issue>), <fpage>547</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-016-1298-9</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kibet</surname>
<given-names>C. K.</given-names>
</name>
<name>
<surname>Machanick</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Transcription factor motif quality assessment requires systematic comparative [version 2; referees: 2 approved]</article-title>. <source>F1000Research</source> <volume>4</volume>, <fpage>1429</fpage>. <pub-id pub-id-type="doi">10.12688/f1000research.7408.2</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kingsley</surname>
<given-names>N. B.</given-names>
</name>
<name>
<surname>Kern</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Creppe</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hales</surname>
<given-names>E. N.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kalbfleisch</surname>
<given-names>T. S.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Functionally annotating regulatory elements in the equine genome using histone mark ChIP-seq</article-title>. <source>Genes (Basel)</source> <volume>11</volume> (<issue>1</issue>), <fpage>11010003</fpage>. <pub-id pub-id-type="doi">10.3390/genes11010003</pub-id>)</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Klepper</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Sandve</surname>
<given-names>G. K.</given-names>
</name>
<name>
<surname>Abul</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Johansen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Drablos</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Assessment of composite motif discovery methods</article-title>. <source>BMC Bioinforma.</source> <volume>9</volume>, <fpage>123</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-9-123</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kong</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Chang</surname>
<given-names>P. K.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Zheng</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Identification of AflR binding sites in the genome of Aspergillus flavus by ChIP-seq</article-title>. <source>J. Fungi (Basel)</source> <volume>6</volume> (<issue>2</issue>), <fpage>6020052</fpage>. <pub-id pub-id-type="doi">10.3390/jof6020052</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Levitsky</surname>
<given-names>V. G.</given-names>
</name>
<name>
<surname>Mukhin</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Oshchepkov</surname>
<given-names>D. Y.</given-names>
</name>
<name>
<surname>Zemlyanskaya</surname>
<given-names>E. V.</given-names>
</name>
<name>
<surname>Lashin</surname>
<given-names>S. A.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Web-MCOT server for motif Co-occurrence search in ChIP-seq data</article-title>. <source>Int. J. Mol. Sci.</source> <volume>23</volume> (<issue>16</issue>), <fpage>23168981</fpage>. <pub-id pub-id-type="doi">10.3390/ijms23168981</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Machanick</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Kibet</surname>
<given-names>C. K.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Challenges with modelling transcription factor binding</article-title>,&#x201d; in <conf-name>1st International Conference on Next Generation Computing Applications (NextComp)</conf-name>, <conf-loc>Mauritius</conf-loc>, <conf-date>19-21 July 2017</conf-date>, <fpage>68</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1109/NEXTCOMP.2017.8016178</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rouault</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Santolini</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Schweisguth</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Hakim</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Imogene: Identification of motifs and cis-regulatory modules underlying gene co-regulation</article-title>. <source>Nucleic Acids Res.</source> <volume>42</volume> (<issue>10</issue>), <fpage>6128</fpage>&#x2013;<lpage>6145</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gku209</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sandve</surname>
<given-names>G. K.</given-names>
</name>
<name>
<surname>Abul</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Walseng</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Drablos</surname>
<given-names>F.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Improved benchmarks for computational motif discovery</article-title>. <source>BMC Bioinforma.</source> <volume>8</volume>, <fpage>193</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-8-193</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Smith</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ventura</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Prince</surname>
<given-names>J. T.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Novel algorithms and the benefits of comparative validation</article-title>. <source>Bioinformatics</source> <volume>29</volume> (<issue>12</issue>), <fpage>1583</fpage>&#x2013;<lpage>1585</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt176</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teague</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Barrows</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Baafi</surname>
<given-names>C. A.</given-names>
</name>
<name>
<surname>Van Dyke</surname>
<given-names>M. W.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Discovering the DNA-binding consensus of the thermus thermophilus HB8 transcriptional regulator TTHA1359</article-title>. <source>Int. J. Mol. Sci.</source> <volume>22</volume> (<issue>18</issue>), <fpage>10042</fpage>. <pub-id pub-id-type="doi">10.3390/ijms221810042</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tran</surname>
<given-names>N. T.</given-names>
</name>
<name>
<surname>Huang</surname>
<given-names>C. H.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>A survey of motif finding Web tools for detecting binding site motifs in ChIP-Seq data</article-title>. <source>Biol. Direct</source> <volume>9</volume>, <fpage>4</fpage>. <pub-id pub-id-type="doi">10.1186/1745-6150-9-4</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vihinen</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>How to evaluate performance of prediction methods? Measures and their interpretation in variation effect analysis</article-title>. <source>BMC Genomics</source> <volume>13</volume>, <fpage>S2</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-13-S4-S2</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zambelli</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Pesole</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Pavesi</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Motif discovery and transcription factor binding sites before and after the next-generation sequencing era</article-title>. <source>Brief. Bioinform</source> <volume>14</volume> (<issue>2</issue>), <fpage>225</fpage>&#x2013;<lpage>237</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbs016</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zeitlinger</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Seven myths of how transcription factors read the cis-regulatory code</article-title>. <source>Curr. Opin. Syst. Biol.</source> <volume>23</volume>, <fpage>22</fpage>&#x2013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1016/j.coisb.2020.08.002</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>FisherMP: Fully parallel algorithm for detecting combinatorial motifs from large ChIP-seq datasets</article-title>. <source>DNA Res.</source> <volume>26</volume> (<issue>3</issue>), <fpage>231</fpage>&#x2013;<lpage>242</lpage>. <pub-id pub-id-type="doi">10.1093/dnares/dsz004</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>