<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1250907</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2023.1250907</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Fully automated annotation of mitochondrial genomes using a cluster-based approach with de Bruijn graphs</article-title>
<alt-title alt-title-type="left-running-head">Fiedler et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fgene.2023.1250907">10.3389/fgene.2023.1250907</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Fiedler</surname>
<given-names>Lisa</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/2363296/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Middendorf</surname>
<given-names>Martin</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/198929/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Bernt</surname>
<given-names>Matthias</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="fn" rid="fn1">
<sup>&#x2020;</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/138649/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Computer Science</institution>, <institution>Leipzig University</institution>, <addr-line>Leipzig</addr-line>, <country>Germany</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Helmholtz Centre for Environmental Research&#x2014;UFZ</institution>, <addr-line>Leipzig</addr-line>, <country>Germany</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1076532/overview">Min Zeng</ext-link>, Central South University, China</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/832518/overview">Junwei Luo</ext-link>, Henan Polytechnic University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/2372436/overview">Xingyu Liao</ext-link>, King Abdullah University of Science and Technology, Saudi Arabia</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Lisa Fiedler, <email>lfiedler@informatik.uni-leipzig.de</email>
</corresp>
<fn fn-type="equal" id="fn1">
<p>
<sup>&#x2020;</sup>These authors have contributed equally to this work</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>10</day>
<month>08</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1250907</elocation-id>
<history>
<date date-type="received">
<day>30</day>
<month>06</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>24</day>
<month>07</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Fiedler, Middendorf and Bernt.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Fiedler, Middendorf and Bernt</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>A wide range of scientific fields, such as forensics, anthropology, medicine, and molecular evolution, benefits from the analysis of mitogenomic data. With the development of new sequencing technologies, the amount of mitochondrial sequence data to be analyzed has increased exponentially over the last few years. The accurate annotation of mitochondrial DNA is a prerequisite for any mitogenomic comparative analysis. To sustain with the growth of the available mitochondrial sequence data, highly efficient automatic computational methods are, hence, needed. Automatic annotation methods are typically based on databases that contain information about already annotated (and often pre-curated) mitogenomes of different species. However, the existing approaches have several shortcomings: 1) they do not scale well with the size of the database; 2) they do not allow for a fast (and easy) update of the database; and 3) they can only be applied to a relatively small taxonomic subset of all species. Here, we present a novel approach that does not have any of these aforementioned shortcomings, (1), (2), and (3). The reference database of mitogenomes is represented as a richly annotated de Bruijn graph. To generate gene predictions for a new user-supplied mitogenome, the method utilizes a clustering routine that uses the mapping information of the provided sequence to this graph. The method is implemented in a software package called DeGeCI <bold>(De</bold> Bruijn graph <bold>Ge</bold>ne <bold>C</bold>luster <bold>I</bold>dentification). For a large set of mitogenomes, for which expert-curated annotations are available, DeGeCI generates gene predictions of high conformity. In a comparative evaluation with MITOS2, a state-of-the-art annotation tool for mitochondrial genomes, DeGeCI shows better database scalability while still matching MITOS2 in terms of result quality and providing a fully automated means to update the underlying database. Moreover, unlike MITOS2, DeGeCI can be run in parallel on several processors to make use of modern multi-processor systems.</p>
</abstract>
<kwd-group>
<kwd>annotation</kwd>
<kwd>gene prediction</kwd>
<kwd>mitochondria</kwd>
<kwd>genome</kwd>
<kwd>mitogenome</kwd>
<kwd>Metazoa</kwd>
<kwd>de Bruijn graph</kwd>
<kwd>clustering</kwd>
</kwd-group>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Computational Genomics</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Mitochondria are spherical organelles found in most eukaryotic cells. Their genome, the mitogenome, differs in various aspects from their nuclear counterpart, which include their size, structure, and composition. In Metazoa, the mitogenome is commonly organized as a double-stranded circular DNA molecule with an average length of approximately 16, 500&#xa0;nt and a small core set of 37 genes, comprised of 13 protein-coding genes, 22 tRNAs, two rRNAs, and one non-coding region, which contains most of the regulatory elements (<xref ref-type="bibr" rid="B23">Wolstenholme, 1992</xref>). Although the gene content is generally well conserved, the gene arrangement varies greatly among animal mitogenomes. This renders them an attractive target for a variety of comparative analyses, such as phylogenetic reconstruction or genome rearrangement studies. To facilitate such analyses to be performed systematically on a large scale, automated, standardized annotation of the mitogenome is an indispensable prerequisite.</p>
<p>The widest selection of publicly available mitochondrial genome data can be found in the GenBank (<xref ref-type="bibr" rid="B3">Benson et al., 2000</xref>) and RefSeq (<xref ref-type="bibr" rid="B19">Pruitt et al., 2007</xref>) databases. GenBank offers access to original sequence data, whereas RefSeq provides a non-redundant expert-curated collection of original GenBank entries. Several databases and tool sets have been built on these data repositories to generate (<italic>de novo</italic>) annotations for user-supplied sequence data, e.g., DOGMA (<xref ref-type="bibr" rid="B24">Wyman et al., 2004</xref>), MOSAS (<xref ref-type="bibr" rid="B20">Sheffield et al., 2010</xref>), MitoFish (<xref ref-type="bibr" rid="B13">Iwasaki et al., 2013</xref>), and MITOS (<xref ref-type="bibr" rid="B4">Bernt et al., 2013</xref>). These approaches identify genes using either the sequence similarity search against sequence databases, containing gene sequences of published mitogenomes, or search with curated (hidden Markov/covariance) gene models. All the aforementioned approaches identify protein-coding genes using BLASTX and/or BLASTN searches against an internal database. DOGMA, MOSAS, and MitoFish further apply this technique to rRNA gene detection, whereas MITOS uses Infernal (<xref ref-type="bibr" rid="B9">Eddy, 2002</xref>; <xref ref-type="bibr" rid="B16">Nawrocki et al., 2009</xref>) and covariance models for mitochondrial rRNAs to serve this purpose. For tRNA annotation, MITOS uses covariance models (<xref ref-type="bibr" rid="B11">Eddy and Durbin, 1994</xref>), DOGMA employs COVE, MOSAS applies ARWEN and tRNAscan-SE (<xref ref-type="bibr" rid="B15">Lowe and Eddy, 1997</xref>), and MitoFish makes use of MiTFi. Meanwhile, an updated and improved version, MITOS2, was developed, which is based on a more current RefSeq release and allows us to search for protein-coding genes with profile hidden Markov models (HMMs) (<xref ref-type="bibr" rid="B8">Donath et al., 2019</xref>). One drawback of DOGMA and MOSAS is that they require some manual improvements on the result set. MOSAS&#x2019;s restriction to insects and MitoFish&#x2019;s restriction to fishes limit their scope of application. The fast-growing amount of available mitogenomes creates two problems for all of the aforementioned approaches: 1) the runtime for the sequence similarity search increases approximately linearly with the database size and 2) the necessary curation of gene models impedes automatic updates that allow the inclusion of new sequence data that becomes available over time.</p>
<p>de Bruijn graphs (<xref ref-type="bibr" rid="B7">Bruijn, 1946</xref>; <xref ref-type="bibr" rid="B12">Good, 1946</xref>) are an important data structure for compact sequence data representation. To this end, sequences are decomposed into small segments, the so-called <italic>k</italic>-mers, which form the vertices of this graph. Two vertices are connected if the suffix of length <italic>k</italic> &#x2212; 1 of the first vertex is equal to the prefix of length <italic>k</italic> &#x2212; 1 of the second vertex. In the field of bioinformatics, de Bruijn graphs have often been used for DNA fragment assembly, such as in <xref ref-type="bibr" rid="B18">Pevzner et al. (2001</xref>), <xref ref-type="bibr" rid="B17">Pevzner et al. (2004</xref>), and <xref ref-type="bibr" rid="B25">Zerbino and Birney (2008</xref>). The latter employs a modified de Bruijn graph, the A-Bruijn graph, which can also be used for repeat classifications. Another variant is the manifold de Bruijn graph (<xref ref-type="bibr" rid="B14">Lin and Pevzner, 2014</xref>), which allows using (<italic>k</italic> &#x2b; 1)-mers of variable lengths, choosing larger values for high-coverage regions and smaller values for low-coverage regions. However, the focus of these applications has been on nuclear genomes. Their huge sequence length, as opposed to mitochondrial genomes, explains the emergence of vastly compressed storage structures proposed in the literature to keep the required amount of memory as small as possible. One such structure is introduced by <xref ref-type="bibr" rid="B6">Bowe et al. (2012)</xref>. <xref ref-type="bibr" rid="B1">Almodaresi et al. (2017)</xref> extended this approach by additionally allowing us to store a single property, the &#x201c;color,&#x201d; per edge. For many applications, such as variant detection, the approach is sufficient where keeping track of the identity of each of the contributing sequences is the only focus. However, if several properties need to be considered, this storage structure cannot be used. Another downside, which also applies to A-Bruijn and manifold de Bruijn graphs, is that they are all generated based on a fixed set of input genomes. When additional sequences need to be embedded or some contained sequences need to be removed, the entire graph must be reconstructed, which is already, for a moderate amount of genomes and/or long sequence lengths, a time-consuming task.</p>
<p>This work presents DeGeCI (<bold>De</bold> Bruijn graph <bold>Ge</bold>ne <bold>C</bold>luster <bold>I</bold>dentification), a new method for the efficient automatic gene detection of mitochondrial genomes. This method uses a collection of mitogenomes, whose sequence data are represented as a richly annotated <bold>M</bold>itochondrial <bold>D</bold>e <bold>B</bold>ruijn <bold>G</bold>raph (<italic>MDBG</italic>). To annotate an input sequence <italic>r</italic>
<sub>in</sub>, a subgraph <inline-formula id="inf1">
<mml:math id="m1">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> induced by all (<italic>k</italic> &#x2b; 1)-mers of <italic>r</italic>
<sub>in</sub> is initially constructed. Unmapped sequence portions result in disconnected components in this subgraph, which are bridged in the following step. To this end, alternative trails in the <italic>MDBG</italic>, exhibiting a high sequence similarity to the respective unmapped subsequences of <italic>r</italic>
<sub>in</sub>, are identified and added to <inline-formula id="inf2">
<mml:math id="m2">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> (<xref ref-type="sec" rid="s2-2-2">Section 2.2.2</xref>). Using a clustering approach, DeGeCI aggregates annotations of the subgraph to obtain gene predictions for the input sequence (<xref ref-type="sec" rid="s2-2-3">Section 2.2.3</xref>). In this study, we use a comprehensive set of all 8,015 mitogenomes contained in RefSeq 89, covering all major metazoan taxonomic groups, to construct the database graph. Gene predictions are computed for a large and taxonomically representative sample of mitogenomes and are compared to existing expert-curated annotations and MITOS2 (<xref ref-type="sec" rid="s4-3">Section 4.3</xref>).</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2 Methods</title>
<sec id="s2-1">
<title>2.1 Graph structure</title>
<p>Given a string (i.e., a sequence of characters), a <italic>k</italic>-mer is a substring of length <italic>k</italic>. A string can be disassembled into all of its (<italic>k</italic> &#x2b; 1)-mers by sliding a window of length (<italic>k</italic> &#x2b; 1) over the string while retaining duplicates. A genome <italic>r</italic> is a string composed of nucleotides A, C, T, and G, where circular genomes are linearized by &#x201c;cutting&#x201d; the genome at an arbitrary but fixed location.</p>
<p>In the proposed de Bruijn graph, <italic>MDBG</italic> (<italic>V</italic>, <italic>E</italic>), over a set of genomes <italic>G</italic> with the vertex set <italic>V</italic> and edge set <italic>E</italic>, each (<italic>k</italic> &#x2b; 1)-mer <italic>x</italic>
<sub>1</sub>
<italic>x</italic>
<sub>2</sub> &#x2026; <italic>x</italic>
<sub>
<italic>k</italic>&#x2b;1</sub> of every genome in <italic>G</italic> leads to two vertices <italic>v</italic> and <italic>v</italic>&#x2032; in <italic>V</italic>, representing the <italic>k</italic>-prefix <italic>x</italic>
<sub>1</sub>
<italic>x</italic>
<sub>2</sub> &#x2026; <italic>x</italic>
<sub>
<italic>k</italic>
</sub> and <italic>k</italic>-suffix <italic>x</italic>
<sub>2</sub> &#x2026; <italic>x</italic>
<sub>
<italic>k</italic>&#x2b;1</sub>, respectively. These are connected by directed edges (<italic>v</italic> and <italic>v</italic>&#x2032;), which represent the (<italic>k</italic> &#x2b; 1)-mer itself. For circular genomes, the (<italic>k</italic> &#x2b; 1)-mers that connect both sides of the linearized string representation are also included. Thus, while each linear genome of length &#x7c;<italic>r</italic>&#x7c; contributes &#x7c;<italic>r</italic>&#x7c; &#x2212; <italic>k</italic> many (<italic>k</italic> &#x2b; 1)-mers, a circular genome of this length contributes &#x7c;<italic>r</italic>&#x7c; many (<italic>k</italic> &#x2b; 1)-mers to the graph. The complementary DNA strand (negative strand), with an opposite reading direction, is taken into account by adding the reverse complement of each (<italic>k</italic> &#x2b; 1)-mer to <italic>MDBG</italic>. Each edge (<italic>v</italic> and <italic>v</italic>&#x2032;) is annotated with a label <italic>r</italic>, denoting the genome from which this edge originates, the strandedness <italic>&#x3c3;</italic> &#x2208; { &#x2b;, &#x2212; }, and the position <italic>p</italic> (with respect to the positive strand) of nucleotide <italic>x</italic>
<sub>
<italic>k</italic>&#x2b;1</sub> of (<italic>v</italic> and <italic>v</italic>&#x2032;) in genome <italic>r</italic>. This allows for an unambiguous reconstruction of each genome in <italic>G</italic> from <italic>MDBG</italic> (<italic>V</italic>, <italic>E</italic>). It should be noted that each (<italic>k</italic> &#x2b; 1)-mer that is contained in multiple different genomes or is contained multiple times in the same genome (i.e., due to repeats) results in a pair of vertices that is connected by multiple edges, the so-called parallel edges. The <italic>MDBG</italic> is thus a multigraph. <xref ref-type="fig" rid="F1">Figure 1</xref> illustrates an example of the de Bruijn graph of a circular genome <italic>r</italic> with the sequence ACTGAA for <italic>k</italic> &#x3d; 3 on the positive strand.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>de Bruijn graph of a circular genome <italic>r</italic> with the sequence ACTGAA for <italic>k</italic> &#x3d; 3 and positive strand <italic>&#x3c3;</italic> &#x3d; &#x2b;. The SGT (3, 2, <italic>r</italic>, &#x2b;) is exemplarily shown. The corresponding edges in the graph are highlighted in bold.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g001.tif"/>
</fig>
<p>A trail in a graph is a sequence of distinct edges that joins a sequence of vertices. Let (<italic>i</italic>, <italic>j</italic>, <italic>r</italic>, <italic>&#x3c3;</italic>) be the trails in <italic>MDBG</italic>, denoting a <italic>single genome trail</italic> (SGT), which is composed of edges corresponding to the subsequence from the position <italic>i</italic> to <italic>j</italic> in the genome <italic>r</italic> located on the strand <italic>&#x3c3;</italic>. For a circular genome, <italic>i</italic> &#x3e;<italic>j</italic>, if the associated subsequence extends over the string boundary of the linearized genome representation. In the de Bruijn graph depicted in <xref ref-type="fig" rid="F1">Figure 1</xref>, one such SGT is exemplarily highlighted.</p>
</sec>
<sec id="s2-2">
<title>2.2 Workflow</title>
<p>For the <italic>de novo</italic> annotation of an input genome <italic>r</italic>
<sub>in</sub>, DeGeCI requires only its nucleotide sequence. The DeGeCI pipeline consists of six major stages, which are summarized in <xref ref-type="fig" rid="F2">Figure 2</xref>. The following sections present the individual steps involved.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>DeGeCI workflow for the <italic>de novo</italic> annotation of an input genome <italic>r</italic>
<sub>in</sub>.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g002.tif"/>
</fig>
<sec id="s2-2-1">
<title>2.2.1 Subgraph construction</title>
<p>Initially, DeGeCI generates the set <inline-formula id="inf3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of all (<italic>k</italic> &#x2b; 1)-mers of the input genome <italic>r</italic>
<sub>in</sub>. Next, the subgraph <inline-formula id="inf4">
<mml:math id="m4">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, which is induced by all (<italic>k</italic> &#x2b; 1)-mers in the database graph <italic>MDBG</italic> that are also contained in <inline-formula id="inf5">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, is constructed. For each such <italic>matching</italic> (<italic>k</italic> &#x2b; 1)-mer in <inline-formula id="inf6">
<mml:math id="m6">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, an edge is added to <inline-formula id="inf7">
<mml:math id="m7">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and labeled with <italic>r</italic>
<sub>in</sub>, the related sequence position <italic>p</italic>, and the strand <inline-formula id="inf8">
<mml:math id="m8">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of <italic>r</italic>
<sub>in</sub>. Thus, there are at least two edges between each pair of vertices in the subgraph: one of <italic>r</italic>
<sub>in</sub> and one of a database genome <italic>r</italic>.</p>
</sec>
<sec id="s2-2-2">
<title>2.2.2 Connected component bridging</title>
<p>Even if dense taxon sampling is provided in the database graph, input species with a poorly conserved gene content can lead to (<italic>k</italic> &#x2b; 1)-mers in <inline-formula id="inf9">
<mml:math id="m9">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> that do not map to any edge in the <italic>MDBG</italic> for a reasonable value of <italic>k</italic> (see <xref ref-type="sec" rid="s4-2-1">Section 4.2.1</xref>). These non-matching (<italic>k</italic> &#x2b; 1)-mers decompose the genome&#x2019;s continuous sequence of (<italic>k</italic> &#x2b; 1)-mers in the consecutive blocks of matching (<italic>k</italic> &#x2b; 1)-mers. Consequently, the subgraph <inline-formula id="inf10">
<mml:math id="m10">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is composed of smaller subgraphs, each of which is induced by one such subsequence block of (<italic>k</italic> &#x2b; 1)-mers. Going forward, these subgraphs will be called connected components (CCs). Thus, two vertices are part of the same CC if there is an SGT of the input genome <italic>r</italic>
<sub>in</sub> that connects them.</p>
<p>While two CCs of <inline-formula id="inf11">
<mml:math id="m11">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> are not connected by SGTs of <italic>r</italic>
<sub>in</sub>, there may be SGTs of other genomes in the database graph <italic>MDBG</italic>, which connect them. More precisely, there might be an SGT <italic>t</italic>
<sub>1</sub> of some genome <italic>r</italic> &#x2208; <italic>G</italic> in one CC <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and another SGT <italic>t</italic>
<sub>2</sub> of this genome in another CC <italic>C</italic>
<sub>
<italic>T</italic>
</sub> in <inline-formula id="inf12">
<mml:math id="m12">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, which might be connected by a third SGT <italic>t</italic>
<sub>3</sub> of <italic>r</italic> in <italic>MDBG</italic>. Such <italic>connecting</italic> trails <italic>t</italic>
<sub>3</sub> between two <italic>seeding</italic> SGTs, <italic>t</italic>
<sub>1</sub> and <italic>t</italic>
<sub>2</sub>, could constitute suitable alternative trails for the unmapped sequence segments in <inline-formula id="inf13">
<mml:math id="m13">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, thereby bridging pairs of CCs <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and <italic>C</italic>
<sub>
<italic>T</italic>
</sub>.</p>
<p>To identify such connecting trails, an algorithm called <sc>cc-bridging</sc> was developed. The pseudocode is given in <xref ref-type="statement" rid="algorithm_1">Algorithm 1</xref>. For each CC <italic>C</italic>
<sub>
<italic>I</italic>
</sub>, induced by subsequence <italic>s</italic>
<sub>
<italic>I</italic>
</sub>, bridging is initially attempted with CC <italic>C</italic>
<sub>
<italic>T</italic>
</sub>, induced by subsequence <italic>s</italic>
<sub>
<italic>T</italic>
</sub>, which among all inducing subsequences of CCs in <inline-formula id="inf14">
<mml:math id="m14">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is (in the reading direction) located closest to <italic>s</italic>
<sub>
<italic>I</italic>
</sub> in the input genome (line 6). This serves to retain sequence locality. For this pair of CCs, the algorithm identifies bridging trails between suitable pairs of seeding trails (line 16). The individual steps of this bridging routine are outlined in <xref ref-type="sec" rid="s2-2-2-1">Section 2.2.2.1</xref>. To prevent a large number of mostly unsuitable seeding trails being validated, only a small portion of both CCs is considered at first. The portions are only extended if no appropriate bridging trails can be identified within them. This is controlled by a parameter <inline-formula id="inf15">
<mml:math id="m15">
<mml:mi mathvariant="double-struck">N</mml:mi>
<mml:mo>&#x220b;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
</mml:math>
</inline-formula>, which might get adapted during program execution (line 20). If at least one of the two CCs was already searched completely, <italic>C</italic>
<sub>
<italic>T</italic>
</sub> is updated to the next closest CC (line 23) and the aforementioned routine starts over again. This is repeated until valid bridging trails are found or if the distance <inline-formula id="inf16">
<mml:math id="m16">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> between the inducing subsequences is more than twice as large as either of their lengths (line 14). The latter is a rough filter for speed-up purposes, with the reasoning that the larger the distance between the inducing subsequences and the smaller their lengths, the less likely it is to find suitable bridging trails between their CCs. <xref ref-type="fig" rid="F3">Figure 3</xref> visualizes this setup.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Identification of bridging trails for components <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and <italic>C</italic>
<sub>
<italic>T</italic>
</sub>. In this example, there are three pairs of seeding SGTs <inline-formula id="inf17">
<mml:math id="m17">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf18">
<mml:math id="m18">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> of genomes <italic>r</italic>&#x2032;&#x2032;&#x2032;, <italic>r</italic>&#x2033;, and <italic>r</italic>&#x2032;, respectively. The corresponding bridging trails are <inline-formula id="inf19">
<mml:math id="m19">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
<mml:mo>&#x2032;</mml:mo>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf20">
<mml:math id="m20">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g003.tif"/>
</fig>
<p>
<statement content-type="algorithm" id="algorithm_1">
<label>Algorithm 1. CC-BRIDGING.</label>
<p>
<inline-graphic xlink:href="fgene-14-1250907-fx1.tif"/>
</p>
</statement>
</p>
<sec id="s2-2-2-1">
<title>2.2.2.1 Validation of seeding trails</title>
<p>The validation of the seeding trails for two CCs, <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and <italic>C</italic>
<sub>
<italic>T</italic>
</sub> (line 16 in <xref ref-type="statement" rid="algorithm_1">Algorithm 1</xref>), consists of the following two steps:</p>
<p>
<statement content-type="step" id="step_1">
<label>Step 1:</label>
<p>Identification of bridging candidates. The method searches for seeding SGTs (<italic>m</italic>&#x2032;, <italic>n</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>) &#x2208; <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and (<italic>o</italic>&#x2032;, <italic>l</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>) &#x2208; <italic>C</italic>
<sub>
<italic>T</italic>
</sub> so that there are <italic>mapping</italic> input genome SGTs <inline-formula id="inf21">
<mml:math id="m21">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf22">
<mml:math id="m22">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>l</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>C</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, and it is not possible to further extend these mappings. Here, mapping means that all (<italic>k</italic> &#x2b; 1)-mers of the respective database SGT coincide with those of the input SGT. For each such pair of seeding SGTs, a possible bridging trail is given by (<italic>n</italic>&#x2032; &#x2b; 1, <italic>o</italic>&#x2032; &#x2212; 1, <italic>r</italic>, <italic>&#x3c3;</italic>).</p>
<p>The two input genome SGTs should be as close as possible to each other to ensure sequence locality. The smallest possible distance between them is <inline-formula id="inf23">
<mml:math id="m23">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> since the inducing subsequences <italic>s</italic>
<sub>
<italic>I</italic>
</sub> and <italic>s</italic>
<sub>
<italic>T</italic>
</sub> are separated by this value. Here, we accept SGTs with a distance of up to <inline-formula id="inf24">
<mml:math id="m24">
<mml:mi>g</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, where <italic>g</italic> is the integer parameter introduced in the previous section. This allows us to adapt the cutoff distance in consecutive method iterations.</p>
<p>Another important aspect to consider is the sequence similarity of the bridging trail with the unmapped subsequence of <italic>r</italic>
<sub>in</sub>. By requesting a small relative distance &#x7c;(<italic>o</italic> &#x2212; <italic>n</italic>) &#x2212; (<italic>o</italic>&#x2032; &#x2212; <italic>n</italic>&#x2032;)&#x7c;/(<italic>o</italic> &#x2212; <italic>n</italic>) of both SGT pairs, the likelihood of a high sequence similarity can be increased without actually evaluating it at this point of the algorithm. In this contribution, we impose an upper bound of 0.2. Examples illustrating the application of the aforementioned two criteria are depicted in <xref ref-type="fig" rid="F4">Figure 4</xref>.</p>
<p>If there is more than one pair of seeding SGTs of the same database genome <italic>r</italic>, the validation routine only retains the pair that is the closest, together with respect to the mapping input genome SGTs, and, in case of a tie, has a lower relative distance.</p>
<p>The algorithm operates on the positions of SGT mappings. Two SGTs of the same genome that correspond to the same sequence, i.e., are a repeat in that genome, have different positions because they are located in different regions of the genome. The algorithm thus treats them exactly the same as any other SGT. Therefore, repeats are not a special case for the algorithm.</p>
</statement>
</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Two scenarios for a pair of seeding SGTs of database genome <italic>r</italic> in CCs <italic>C</italic>
<sub>
<italic>I</italic>
</sub> and <italic>C</italic>
<sub>
<italic>T</italic>
</sub>, for which a distance of <inline-formula id="inf25">
<mml:math id="m25">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>12</mml:mn>
</mml:math>
</inline-formula> shall be assumed. The distance between the input genome SGTs is 20 in both cases. This means that even with the initial setting of <italic>g</italic> &#x3d; <italic>g</italic>
<sub>0</sub> &#x3d; 2, these trails are sufficiently close to one another, i.e., closer than <inline-formula id="inf26">
<mml:math id="m26">
<mml:msub>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>24</mml:mn>
</mml:math>
</inline-formula>. Thus, only the relative distance determines whether to retain SGTs. In the top figure, database and input genome SGTs have a relative distance of zero. The seeding SGTs will, hence, be considered further. In the bottom figure, the relative distance of both SGT pairs is &#x7c;20 &#x2212; 30&#x7c;/20 &#x3d; 0.5, which exceeds the upper bound of 0.2. These trails are, hence, rejected.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g004.tif"/>
</fig>
<p>
<statement content-type="step" id="step_2">
<label>Step 2:</label>
<p>Pairwise sequence alignments. For each of the remaining SGT pairs, the corresponding bridging trails are examined for sequence similarity to the input genome. To this end, the algorithm conducts local pairwise sequence alignments with affine gap costs (cf. <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s5">Section 5</xref> for details) between the unmapped input sequence segment and sequences that correspond to the bridging trails. Alignments are accepted if they have an E-value of at most 10<sup>&#x2013;3</sup>.</p>
</statement>
</p>
</sec>
</sec>
<sec id="s2-2-3">
<title>2.2.3 Gene annotation</title>
<p>The basis for the annotation of the input genome <italic>r</italic>
<sub>in</sub> is a collection of gene annotations <inline-formula id="inf27">
<mml:math id="m27">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> of the database genomes. An element (<italic>n</italic>, <italic>m</italic>, <italic>r</italic>, <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub>, <italic>g</italic>) of <inline-formula id="inf28">
<mml:math id="m28">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> denotes that gene <italic>g</italic> is annotated on the strand <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub> from position <italic>n</italic> to <italic>m</italic> in the genome <italic>r</italic> &#x2208; <italic>G</italic>. A special gene <italic>g</italic>
<sub>0</sub> is used to label regions where no gene is annotated. This serves to avoid a bias toward a small number of random predictions in a later stage of the method. Moreover, each copy or fragment of the same gene gets a different label <italic>g</italic> and, hence, results in a new annotation in <inline-formula id="inf29">
<mml:math id="m29">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula>.</p>
<p>For each pair of mapping SGTs <inline-formula id="inf30">
<mml:math id="m30">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and (<italic>i</italic>&#x2032;, <italic>j</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>) that have been obtained from the subgraph <inline-formula id="inf31">
<mml:math id="m31">
<mml:mi>M</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>B</mml:mi>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">[</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">K</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">]</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> or the bridging routine in the previous steps, the position overlap of [<italic>i</italic>&#x2032;, <italic>j</italic>&#x2032;] with all annotations <inline-formula id="inf32">
<mml:math id="m32">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>r</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> is evaluated. This results in sub-annotations (<italic>n</italic>&#x2032;, <italic>m</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub>, <italic>g</italic>), where [<italic>n</italic>&#x2032;, <italic>m</italic>&#x2032;] &#x3d; [<italic>i</italic>&#x2032;, <italic>j</italic>&#x2032;] &#x2229; [<italic>n</italic>, <italic>m</italic>], yielding an annotation for <italic>r</italic>
<sub>in</sub> from <italic>i</italic> &#x2b; (<italic>n</italic>&#x2032; &#x2212; <italic>i</italic>&#x2032;) to <italic>j</italic> &#x2212; (<italic>j</italic>&#x2032; &#x2212; <italic>m</italic>&#x2032;), which is located on the strand <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub> if <inline-formula id="inf33">
<mml:math id="m33">
<mml:mi>&#x3c3;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and otherwise on the opposing strand. This is because <inline-formula id="inf34">
<mml:math id="m34">
<mml:mi>&#x3c3;</mml:mi>
<mml:mo>&#x2260;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> means that the mapping database SGT and input SGT reside on different strands, indicating that the encoding sequences of both genes are located on opposite strands. <xref ref-type="fig" rid="F5">Figure 5</xref> illustrates an example of the aforementioned annotation process for one database and input genome SGT mapping.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>Input genome SGT <inline-formula id="inf35">
<mml:math id="m35">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>5,11</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> mapping to a database SGT (<italic>i</italic>&#x2032;, <italic>j</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>) &#x3d;(20,26, <italic>r</italic>, &#x2212;) (vertical black lines). The example assumes an annotation (<italic>n</italic>, <italic>m</italic>, <italic>r</italic>, <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub>, <italic>g</italic>) &#x3d; (1, 25, <italic>r</italic>, &#x2b; ,<italic>g</italic>) in <inline-formula id="inf36">
<mml:math id="m36">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula>. In other words, the encoding sequence for <italic>g</italic> is located on the opposing strand of the mapping database SGT. This leads to a sub-annotation (<italic>n</italic>&#x2032;, <italic>m</italic>&#x2032;, <italic>r</italic>, <italic>&#x3c3;</italic>
<sub>
<italic>g</italic>
</sub>, <italic>g</italic>) &#x3d; (20, 25, <italic>r</italic>, &#x2b; ,<italic>g</italic>) (dark blue box on the top), resulting in an annotation of gene <italic>g</italic> from position 5 to 10 in <italic>r</italic>
<sub>in</sub> on the negative strand (light blue box at the bottom), since <inline-formula id="inf37">
<mml:math id="m37">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2b;</mml:mo>
<mml:mo>&#x2260;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3c3;</mml:mi>
</mml:math>
</inline-formula>.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g005.tif"/>
</fig>
<p>An annotation in <inline-formula id="inf38">
<mml:math id="m38">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> generally results in multiple smaller sub-annotations for <italic>r</italic>
<sub>in</sub>, which get separated by small sequence dissimilarities between the input genome and the database genome. If two such fragments originate from the same annotation in <inline-formula id="inf39">
<mml:math id="m39">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> and are separated by a distance smaller than the shortest gene that is typically present in the class of species under consideration (e.g., approximately 70&#xa0;nt for metazoan mitogenomes), it can be assumed that the missing region in between also encodes the same gene. Thus, DeGeCI iteratively replaces such fragments with a single longer annotation of the respective gene.</p>
<p>It is to be expected that flawed annotations in <inline-formula id="inf40">
<mml:math id="m40">
<mml:mi mathvariant="script">A</mml:mi>
</mml:math>
</inline-formula> and/or random (<italic>k</italic> &#x2b; 1)-mer matches cause some incorrect gene predictions for <italic>r</italic>
<sub>in</sub>. DeGeCI thus utilizes aggregations of annotations, so that the predominant and, therefore, more likely predictions can be identified. To this end, the following clustering routine <sc>clusterG</sc> was developed. The goal of the routine is to identify as large a region in the input genome as possible where the same few genes are most likely to occur, thereby filtering out noise contained in database gene annotations. Starting from the database gene annotations for each position of <italic>r</italic>
<sub>in</sub>, the routine builds a hierarchy of gene prediction clusters by greedily merging the annotations of neighboring positions in a bottom-up fashion. The frequency with which annotations for specific genes occur at a given position determines the likelihood of these genes being correct predictions for that position. Moreover, positions with a similar gene distribution are likely to belong together. This distribution is thus used as a quality measure for identifying the merge quality of two neighboring clusters. For merged clusters, the gene annotations that occur in only one of the clusters are removed to derive an increasingly consistent prediction for combined position ranges of merged clusters. Lastly, gene annotations with a high probability are retrieved from different levels of the generated hierarchy of clusters. The detailed steps are described in the following sections.</p>
<sec id="s2-2-3-1">
<title>2.2.3.1 Aggregating gene predictions</title>
<p>The pseudocode of <sc>clusterG</sc> is given in <xref ref-type="statement" rid="algorithm_2">Algorithm 2</xref>. The algorithm is performed individually with gene predictions <inline-formula id="inf41">
<mml:math id="m41">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> for the positive strand <italic>&#x3c3;</italic> &#x3d; &#x2b; and gene predictions <inline-formula id="inf42">
<mml:math id="m42">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> for the negative strand <italic>&#x3c3;</italic> &#x3d; &#x2212; of the input genome <italic>r</italic>
<sub>in</sub>.</p>
<p>
<statement content-type="algorithm" id="algorithm_2">
<label>Algorithm 2. CLUSTERG.</label>
<p>
<inline-graphic xlink:href="fgene-14-1250907-fx2.tif"/>
</p>
<p>Initially, every position <italic>i</italic> of <italic>r</italic>
<sub>in</sub> constitutes a cluster <inline-formula id="inf43">
<mml:math id="m43">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. Each such cluster is initialized with a relative frequency distribution (RFD) using gene predictions <inline-formula id="inf44">
<mml:math id="m44">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">A</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">i</mml:mi>
<mml:mi mathvariant="normal">n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> (line 7). This RFD is defined as<disp-formula id="e1">
<mml:math id="m45">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c9;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x2255;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf45">
<mml:math id="m46">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is the sum of the lengths of gene <italic>g</italic> predictions for the strand <italic>&#x3c3;</italic> that include position <italic>i</italic> and <italic>&#x3c9;</italic>
<sub>
<italic>g</italic>
</sub> is the reciprocal of the length of the gene on average. This assigns higher probabilities to longer and, hence, more trustworthy predictions while preventing bias toward long genes. To warrant a certain degree of reliability, genes are only included in distribution 1 if there are at least two predictions of different database genomes. Moreover, clusters without predictions are removed (line 5).</p>
<p>Once the clusters are initialized, the iterative merging begins (line 9). At each iteration point, the pair of clusters with the highest quality (defined in the following steps) among all mergeable clusters is identified (lines 15&#x2013;20). Two clusters <inline-formula id="inf46">
<mml:math id="m47">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf47">
<mml:math id="m48">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> are considered mergeable if they are neighboring clusters, i.e., <italic>m</italic> &#x2212; <italic>j</italic> &#x3d; 1, and their distributions share at least one gene (line 14).</p>
<p>To evaluate the quality of a merge of two clusters <inline-formula id="inf48">
<mml:math id="m49">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf49">
<mml:math id="m50">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, the joined RFD,<disp-formula id="e2">
<mml:math id="m51">
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>&#x2227;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>&#x2227;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2260;</mml:mo>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:math>
<label>(2)</label>
</disp-formula>is considered. Here, <italic>&#x3c7;</italic>
<sub>
<italic>P</italic>
</sub> equals one if predicate <italic>P</italic> is true and zero otherwise, setting the probability of a gene to zero if it does not appear in both contributing distributions. This is to prevent the annotation of a gene at positions that are not predicted by any of the genomes in the database.</p>
<p>Given such a joined RFD of <italic>n</italic> genes, the smallest number <italic>t</italic> &#x2264;<italic>n</italic> of genes necessary to obtain a cumulated relative frequency of at least 0.95 is identified. Merged clusters with a low value of <italic>t</italic> feature a joined RFD where much weight falls on few genes. This indicates that the two entering clusters can be combined to a consistent prediction, assigning a higher quality to the merge.</p>
<p>At some point, there will most likely be more than one pair of clusters with the smallest value of <italic>t</italic> among all mergeable clusters. In such a case, clusters with a higher maximum score <inline-formula id="inf50">
<mml:math id="m52">
<mml:mi>max</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> are preferred. If there is a tie, the pair with the largest number of predictions is selected. Finally, in the rare case of remaining ties, the routine selects a pair at random (line 23).</p>
<p>Merging is repeated until no clusters are left that can be joined (line 21). For every merged cluster <inline-formula id="inf51">
<mml:math id="m53">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, references to the two contributing clusters <inline-formula id="inf52">
<mml:math id="m54">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (line 26) and <inline-formula id="inf53">
<mml:math id="m55">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (line 27) are maintained. It should be noted that merging generally stops before all positions are merged into a single cluster. This happens, for example, because there might be neighboring clusters that do not share a single common gene and therefore will never be merged. The outcome is, hence, a family of dendrograms <inline-formula id="inf54">
<mml:math id="m56">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula>.</p>
<p>The last step is to retrieve annotations from clusters in <inline-formula id="inf55">
<mml:math id="m57">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula> (line 29), described as follows: initially, clusters <inline-formula id="inf56">
<mml:math id="m58">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula> are selected where the highest relative frequency max{<italic>p</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub>(<italic>g</italic>)}&#x2254;<italic>p</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> (<italic>g</italic>
<sup>&#x22c6;</sup>) of their RFD is at least 0.7. A high value of <italic>p</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> (<italic>g</italic>
<sup>&#x22c6;</sup>) means that the vast majority of predictions in <inline-formula id="inf57">
<mml:math id="m59">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> are for <italic>g</italic>
<sup>&#x22c6;</sup>, increasing the likelihood that <italic>g</italic>
<sup>&#x22c6;</sup> is the correct annotation for the sequence segment associated with the cluster. Hence, gene <italic>g</italic>
<sup>&#x22c6;</sup> is considered as a putative prediction for this segment.</p>
<p>Next, clusters <inline-formula id="inf58">
<mml:math id="m60">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are discarded if there are other clusters <inline-formula id="inf59">
<mml:math id="m61">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> with the same gene <italic>g</italic>
<sup>&#x22c6;</sup> and [<italic>i</italic>, <italic>j</italic>] &#x2282; [<italic>m</italic>, <italic>n</italic>]. This is because <inline-formula id="inf60">
<mml:math id="m62">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> does not lead to any information gain regarding the region where <italic>g</italic>
<sup>&#x22c6;</sup> is presumably encoded. However, if <italic>g</italic>
<sup>&#x22c6;</sup> is different in both clusters, <inline-formula id="inf61">
<mml:math id="m63">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>j</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> is retained to allow for overlapping predictions. <xref ref-type="fig" rid="F6">Figure 6</xref> shows an example of this retrieval routine.</p>
<p>For each of the remaining clusters <inline-formula id="inf62">
<mml:math id="m64">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, gene <italic>g</italic>
<sup>&#x22c6;</sup> is accepted as the prediction from <italic>m</italic> to <italic>n</italic> for the strand <italic>&#x3c3;</italic>.</p>
</statement>
</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>Gene prediction retrieval from the clusters in <inline-formula id="inf63">
<mml:math id="m65">
<mml:mi mathvariant="script">D</mml:mi>
</mml:math>
</inline-formula>. All clusters with maximum relative frequencies of at least 0.7 are marked with an asterisk. Of these, clusters <inline-formula id="inf64">
<mml:math id="m66">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1,2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf65">
<mml:math id="m67">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1,3</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> are discarded, since both have the same gene <italic>g</italic>
<sup>&#x22c6;</sup> &#x3d; <italic>g</italic> as <inline-formula id="inf66">
<mml:math id="m68">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1,5</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x22c6;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and [1, 2] &#x2282; [1, 3] &#x2282; [1, 5]. Thus, only clusters <inline-formula id="inf67">
<mml:math id="m69">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1,5</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf68">
<mml:math id="m70">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">P</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>4,5</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (framed by rectangles) are selected.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g006.tif"/>
</fig>
</sec>
<sec id="s2-2-3-2">
<title>2.2.3.2 Handling unannotated regions</title>
<p>We recall that in the initial RFD Eq. <xref ref-type="disp-formula" rid="e1">1</xref>, each prediction of gene <italic>g</italic> is scaled with weight <italic>&#x3c9;</italic>
<sub>
<italic>g</italic>
</sub>, which is the reciprocal of the length of such a gene on average. However, there is no reasonable definition of such a length for unannotated regions, i.e., where <italic>g</italic> &#x3d; <italic>g</italic>
<sub>0</sub>. Thus, prior to clustering, positions <italic>p</italic> of <italic>r</italic>
<sub>in</sub>, presumably not encoding any actual gene, are identified as follows: there are more predictions with <italic>g</italic> &#x3d; <italic>g</italic>
<sub>0</sub> than with the total number of predictions of the two (or one if there is only one) best scoring genes <italic>g</italic> &#x2260; <italic>g</italic>
<sub>0</sub> at <italic>p</italic> or the cumulative relative frequency of these genes falls below 0.8. The rationale behind taking the two best scoring genes into account is that two genes may well overlap, but the precise gene boundaries may vary slightly among different genomes.</p>
</sec>
</sec>
</sec>
<sec id="s2-3">
<title>2.3 Implementation</title>
<p>DeGeCI is available as a free open-source software package<xref ref-type="fn" rid="fn2">
<sup>1</sup>
</xref>. It is implemented in Apache Spark (<xref ref-type="bibr" rid="B22">Veith and de Assuncao, 2019</xref>) using its Java API, facilitating parallel program execution. The database graph <italic>MDBG</italic> is stored in an indexed PostgreSQL database (<xref ref-type="bibr" rid="B21">Stonebraker, 1987</xref>). PostgreSQL dump files for the database population are available for RefSeq 89 and RefSeq 204<xref ref-type="fn" rid="fn3">
<sup>2</sup>
</xref>.</p>
<p>To initiate the annotation of an input genome, its nucleotide sequence can be provided either as a FASTA file or as a sequence string. The final gene annotation is provided as a bed file.</p>
</sec>
</sec>
<sec sec-type="materials" id="s3">
<title>3 Materials</title>
<p>To create the database graph <italic>MDBG</italic>, a comprehensive set of all 8,015 metazoan mitogenomes contained in RefSeq 89 was used. The <italic>k</italic>-mer size was set to <italic>k</italic> &#x3d; 16, following a careful empirical analysis presented in <xref ref-type="sec" rid="s4-2-1">Section 4.2.1</xref>. The curated GenBank files of the RefSeq dataset served as a basis for the gene annotation of SGT mappings. To achieve a consistent nomenclature among all entries, an important prerequisite to derive joint annotation, each GenBank entry was parsed following the guidelines suggested by <xref ref-type="bibr" rid="B5">Boore (2006</xref>). The details are compiled in <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s2">Section 2</xref>.</p>
<p>To assess the quality of the proposed method, DeGeCI was applied to a sample of 100 genomes of RefSeq 89 (for a complete list, see <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s6">Section 6</xref>), for which expert-curated annotations exist. These annotations served as ground truth data, allowing us to assess the accuracy of the produced gene predictions. This sample was drawn by randomly selecting different numbers of genomes from each of the major metazoan groups contained in this RefSeq release. The number of species per group was chosen with respect to the frequency in which they occurred in RefSeq 89 (and, therefore, <italic>MDBG</italic>). This comprises seven Spiralia, 27 Arthropoda, 32 Actinopterygii, four Amphibia, 15 Mammalia, 13 Sauropsida, and two non-bilaterian species. In order to not to bias the gene predictions of an input genome of the sample with its annotation in the database graph, all edges of this genome in the database graph were excluded during the subgraph construction step of DeGeCI.</p>
</sec>
<sec sec-type="results" id="s4">
<title>4 Results</title>
<p>This contribution presents a new de Bruijn graph-based method, DeGeCI, for <italic>de novo</italic> gene annotations of mitochondrial genomes. In the following, we compare DeGeCI with MITOS2 (<xref ref-type="bibr" rid="B4">Bernt et al., 2013</xref>), a widely used state-of-the-art annotation tool for mitochondrial genomes. Both MITOS2 and DeGeCI are based on an internal database of mitogenomic sequences of the RefSeq database and require only nucleotide sequences as input, allowing for a fair comparison.</p>
<p>While both tools are based on an internal database of mitogenomic sequences, the underlying approaches are essentially different. MITOS2 uses profile hidden Markov models in combination with methods from <xref ref-type="bibr" rid="B8">Donath et al. (2019</xref>) and the HMMER software suite (<xref ref-type="bibr" rid="B10">Eddy, 2011</xref>) or BLASTX searches (<xref ref-type="bibr" rid="B2">Altschul et al., 1990</xref>) for the annotation of protein-coding genes. To detect non-coding RNAs, i.e., tRNAs and rRNAs, MITOS2 uses Infernal (<xref ref-type="bibr" rid="B9">Eddy, 2002</xref>; <xref ref-type="bibr" rid="B16">Nawrocki et al., 2009</xref>) in the glocal search mode and curated covariance models. Contrarily, DeGeCI is a stand-alone application that relies on no third-party bioinformatics software. Both proteins and RNAs are annotated using the mapping information of the input genome to a database graph, in combination with a subsequent clustering approach.</p>
<sec id="s4-1">
<title>4.1 Benchmarking procedure</title>
<p>To allow for a comparative evaluation with MITOS2, the following definitions are adopted from <xref ref-type="bibr" rid="B4">Bernt et al. (2013</xref>). For each DeGeCI/MITOS2 prediction, the corresponding RefSeq prediction is the annotation that has the largest position overlap with the DeGeCI/MITOS2 prediction, given that at least 75% of the DeGeCI/MITOS2 positions are shared with the RefSeq predictions. Each such allocation of a DeGeCI/MITOS2 annotation to a RefSeq annotation is classified as equal if both predict the same gene on the same strand, classified as different if both predict different genes, and classified as having a strand difference if both predict the same gene but on opposing strands. DeGeCI/MITOS2 predictions, where no corresponding RefSeq prediction is found, are classified as false positives (FPs). Analogously, RefSeq predictions without corresponding DeGeCI/MITOS2 predictions are classified as false negatives (FNs).</p>
</sec>
<sec id="s4-2">
<title>4.2 Parameter settings</title>
<p>For MITOS2, the default parameter setting and appropriate genetic codes were used for each of the 100 considered species. The parameters for DeGeCI were set as follows.</p>
<sec id="s4-2-1">
<title>4.2.1 (<italic>k</italic> &#x2b; 1)-mer size</title>
<p>To determine a suitable value for the (<italic>k</italic> &#x2b; 1)-mer size of the database graph, the right balance has to be found between a too-small value, leading to a great number of random (<italic>k</italic> &#x2b; 1)-mer matches, and a too-large value, concealing many sequence similarities among the genomes.</p>
<p>To this end, the following experiment was conducted 100 times for every <inline-formula id="inf69">
<mml:math id="m71">
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">K</mml:mi>
<mml:mo>&#x2254;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mn>6,8,10,12,14,16,18</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> in turn. First, the multiset (i.e., a set allowing for duplicate entries) <inline-formula id="inf70">
<mml:math id="m72">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of all (<italic>k</italic> &#x2b; 1)-mers of the 8,015 mitochondrial sequences in RefSeq 89 was generated. Next, all sequences were concatenated into a single long sequence, which was subsequently randomly shuffled. From this sequence, a multiset <inline-formula id="inf71">
<mml:math id="m73">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of random (<italic>k</italic> &#x2b; 1)-mers was then constructed. By creating the random set in this way, it features the same nucleotide composition and an almost identical size <inline-formula id="inf72">
<mml:math id="m74">
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x3d;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> as <inline-formula id="inf73">
<mml:math id="m75">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, where <italic>N</italic> is the number of cyclic genomes in RefSeq 89. The average (taken over all 100 experiments) fraction <inline-formula id="inf74">
<mml:math id="m76">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>hit</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of (<italic>k</italic> &#x2b; 1)-mers in <inline-formula id="inf75">
<mml:math id="m77">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> that are also contained in <inline-formula id="inf76">
<mml:math id="m78">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> at least <italic>w</italic>
<sub>min</sub> &#x2265;1 times can then be used to estimate the likelihood of randomly finding (<italic>k</italic> &#x2b; 1)-mers of the true genome sequences in unrelated sequences of same composition.</p>
<p>
<xref ref-type="fig" rid="F7">Figure 7</xref> depicts these ratios for all <inline-formula id="inf77">
<mml:math id="m79">
<mml:mi>k</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi mathvariant="script">K</mml:mi>
</mml:math>
</inline-formula> and <italic>w</italic>
<sub>min</sub> &#x2208; {1, 2, 3, 4}. For <italic>k</italic> &#x2264;8, each (<italic>k</italic> &#x2b; 1)-mer in <inline-formula id="inf78">
<mml:math id="m80">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is also contained in <inline-formula id="inf79">
<mml:math id="m81">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> at least four times (i.e., <inline-formula id="inf80">
<mml:math id="m82">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>hit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>100</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula>). Even all 4<sup>
<italic>k</italic>&#x2b;1</sup> distinct (<italic>k</italic> &#x2b; 1)-mers that can be generated from an alphabet of nucleotides A, C, T, and G are contained within <inline-formula id="inf81">
<mml:math id="m83">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> in this case, as shown in <xref ref-type="fig" rid="F8">Figure 8</xref>. This demonstrates that larger <italic>k</italic> values need to be used to achieve that the (<italic>k</italic> &#x2b; 1)-mers carry at least some meaningful information. However, it is not until <italic>k</italic> &#x3d; 16 that a significance level of 1% (dashed line) is achieved, at least for <italic>w</italic>
<sub>min</sub> &#x3e;1 (<xref ref-type="fig" rid="F7">Figure 7</xref>), and hardly any of all 4<sup>
<italic>k</italic>&#x2b;1</sup> distinct (<italic>k</italic> &#x2b; 1)-mers occurs in <inline-formula id="inf82">
<mml:math id="m84">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (<xref ref-type="fig" rid="F8">Figure 8</xref>). For <italic>k</italic> &#x2265;18, <inline-formula id="inf83">
<mml:math id="m85">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>hit</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3c;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> even for <italic>w</italic>
<sub>min</sub> &#x3d; 1. However, now, more than three-quarters of the (<italic>k</italic> &#x2b; 1)-mers in <inline-formula id="inf84">
<mml:math id="m86">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> are unique (see <xref ref-type="fig" rid="F9">Figure 9</xref>). This conceals many sequence similarities between the genomes in the database graph, rendering this choice unsuitable for the presented approach. Thus, <italic>k</italic> &#x3d; 16 is used instead.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption>
<p>Average fraction <inline-formula id="inf85">
<mml:math id="m87">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext>hit</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of random (<italic>k</italic> &#x2b; 1)-mers that are also contained in the true (<italic>k</italic> &#x2b; 1)-mer multiset <inline-formula id="inf86">
<mml:math id="m88">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> at least <italic>w</italic>
<sub>min</sub> times. Error bars are hardly noticeable, indicating small statistical fluctuations.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g007.tif"/>
</fig>
<fig id="F8" position="float">
<label>FIGURE 8</label>
<caption>
<p>Percentage of unique true (<italic>k</italic> &#x2b; 1)-mers on all possible 4<sup>
<italic>k</italic>&#x2b;1</sup> unique (<italic>k</italic> &#x2b; 1)-mers.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g008.tif"/>
</fig>
<fig id="F9" position="float">
<label>FIGURE 9</label>
<caption>
<p>Number of unique true (<italic>k</italic> &#x2b; 1)-mers on the log scale.</p>
</caption>
<graphic xlink:href="fgene-14-1250907-g009.tif"/>
</fig>
</sec>
<sec id="s4-2-2">
<title>4.2.2 Sequence alignments</title>
<p>For sequence alignments, we applied match costs of 1, mismatch costs of &#x2212;2, and gap penalties of &#x2212;2 for opening and extending a gap. The <italic>E</italic>-value threshold was set at 10<sup>&#x2013;3</sup>. These are common settings, which assume 95% of sequence conservation.</p>
</sec>
</sec>
<sec id="s4-3">
<title>4.3 Comparison with MITOS2</title>
<p>
<xref ref-type="table" rid="T1">Table 1</xref> summarizes the annotation quality of DeGeCI and MITOS2 for all predictions together and also for the different groups of genes: proteins, tRNAs, and rRNAs. RefSeq 89 was used as a reference database for both tools. DeGeCI identified all of the rRNA and protein and 99% of tRNA RefSeq predictions with an equal gene and strand annotation. Thus, DeGeCI obtains an even larger number of correct predictions than MITOS2 (only 99.1% of the protein and 98.7% of the tRNA RefSeq predictions). A genewise and taxonomic breakdown of the results of both tools is compiled in <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Tables S1&#x2013;S11</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Comparison of DeGeCI and MITOS2 predictions with RefSeq 89 annotations. Here, the number of RefSeq predictions <inline-formula id="inf87">
<mml:math id="m89">
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="monospace">R</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">f</mml:mi>
<mml:mi mathvariant="monospace">S</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">q</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, equal predictions (equal), predictions with different strand annotations (&#x394;&#xb1;), predictions where gene annotations are different (different), false negatives (FNs), and false positives (FPs) of both tools for each type of gene (protein, tRNA, rRNA, and all) are shown. The percentage of equal DeGeCI/MITOS2 predictions with respect to RefSeq predictions is given in parentheses. Results that have better agreement with RefSeq predictions are highlighted in bold.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="right">
<inline-formula id="inf88">
<mml:math id="m90">
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="monospace">R</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">f</mml:mi>
<mml:mi mathvariant="monospace">S</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">q</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</th>
<th align="left"/>
<th align="right">Equal</th>
<th align="right">&#x394;&#xb1;</th>
<th align="right">Different</th>
<th align="right">FN</th>
<th align="right">FP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Protein</td>
<td align="right">1,302</td>
<td align="left">DeGeCI</td>
<td align="right">
<bold>1,302</bold> (100%)</td>
<td align="right">0</td>
<td align="right">
<bold>0</bold>
</td>
<td align="right">
<bold>0</bold>
</td>
<td align="right">
<bold>0</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left"/>
<td align="left">MITOS2</td>
<td align="right">1,290 (99.1%)</td>
<td align="right">0</td>
<td align="right">17</td>
<td align="right">12</td>
<td align="right">13</td>
</tr>
<tr>
<td align="left">tRNA</td>
<td align="right">2,162</td>
<td align="left">DeGeCI</td>
<td align="right">
<bold>2,141</bold> (99.0%)</td>
<td align="right">11</td>
<td align="right">
<bold>1</bold>
</td>
<td align="right">
<bold>9</bold>
</td>
<td align="right">3</td>
</tr>
<tr>
<td align="left"/>
<td align="left"/>
<td align="left">MITOS2</td>
<td align="right">2,134 (98.7%)</td>
<td align="right">11</td>
<td align="right">4</td>
<td align="right">13</td>
<td align="right">
<bold>0</bold>
</td>
</tr>
<tr>
<td align="left">rRNA</td>
<td align="right">200</td>
<td align="left">DeGeCI</td>
<td align="right">200 (100%)</td>
<td align="right">0</td>
<td align="right">0</td>
<td align="right">0</td>
<td align="right">0</td>
</tr>
<tr>
<td align="left"/>
<td align="left"/>
<td align="left">MITOS2</td>
<td align="right">200 (100%)</td>
<td align="right">0</td>
<td align="right">0</td>
<td align="right">0</td>
<td align="right">0</td>
</tr>
<tr>
<td align="left">All</td>
<td align="right">3,664</td>
<td align="left">DeGeCI</td>
<td align="right">
<bold>3,643</bold> (99.4%)</td>
<td align="right">11</td>
<td align="right">
<bold>1</bold>
</td>
<td align="right">
<bold>9</bold>
</td>
<td align="right">
<bold>3</bold>
</td>
</tr>
<tr>
<td align="left"/>
<td align="left"/>
<td align="left">MITOS2</td>
<td align="right">3,624 (98.9%)</td>
<td align="right">11</td>
<td align="right">21</td>
<td align="right">25</td>
<td align="right">13</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The main cause for the few RefSeq tRNA predictions without equal DeGeCI predictions is the annotation of opposing strands, i.e., corresponding RefSeq and DeGeCI predictions, which are classified as having a strand difference. This affects 11 predictions (see <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Table S12</xref> for details). In each such case, more than 95% of DeGeCI positions are shared with RefSeq positions, suggesting a high agreement of both predictions. DeGeCI also always annotates the negative strand, whereas RefSeq annotates the positive strand. Since a special &#x201c;complement&#x201d; tag needs to be set for a RefSeq annotation to indicate that the gene is located on the negative strand, there is a reason to presume that this tag was simply forgotten in the corresponding RefSeq entries. The fact that all 11 predictions were also annotated with the opposite strand by MITOS2 supports this hypothesis. Furthermore, nine of the RefSeq predictions stem from the same genome, and not a single gene of this genome is marked with the complement tag, indicating a systematic error in these RefSeq annotations.</p>
<p>There is one DeGeCI prediction that annotates a different gene rather than the corresponding RefSeq prediction, i.e., which has the classification as &#x201c;different.&#x201d; It involves the annotation of a different anticodon type of a leucine tRNA. As discussed in <xref ref-type="bibr" rid="B4">Bernt et al. (2013</xref>), there are inconsistencies in the naming scheme of RefSeq annotations, resulting in misannotations of the anticodon types of serine and leucine tRNAs (also cf. <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s2">Section 2</xref>). Again, MITOS2 identifies the same discrepancy, leaving little doubt that the anticodon type of the RefSeq prediction should be different (cf. <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Table S13</xref>).</p>
<p>The remaining nine RefSeq predictions without the corresponding equal DeGeCI prediction are FNs (25 in MITOS2), accounting for only 0.2% of the RefSeq entries (<inline-formula id="inf89">
<mml:math id="m91">
<mml:mo>&#x2248;</mml:mo>
<mml:mn>0.7</mml:mn>
<mml:mi>%</mml:mi>
</mml:math>
</inline-formula> for MITOS2). Three of them are also not predicted by MITOS2 (i.e., three FNs are shared by both tools), indicating an increased likelihood that they are misannotations in RefSeq.</p>
<p>There are very few (three) DeGeCI annotations that are classified as FPs, all of which are tRNAs. MITOS2, on the other hand, predicts 13 FPs, which are all proteins. For a complete list of FNs of both tools, see <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Tables S14, S15</xref> and for their FPs, see <xref ref-type="sec" rid="s11">Supplementary Table S16</xref>.</p>
</sec>
<sec id="s4-4">
<title>4.4 Accuracy of gene predictions</title>
<p>The majority of the DeGeCI predictions are in much better agreement with the RefSeq predictions than the threshold of 75% overlap requires (cf. <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Table S17</xref>). More precisely, the average fractions of the DeGeCI positions that are shared with those predicted by RefSeq exceed 98.7%, and the average fractions of RefSeq positions that are shared with those predicted by DeGeCI exceed 97.3% for all types of genes (similar for MITOS2).</p>
<p>DeGeCI also predicts start and end positions with fairly high precision. <xref ref-type="table" rid="T2">Table 2</xref> shows that in 96% of the tRNA, 85% of the protein, and 78% of rRNA predictions, the start and stop positions vary by less than 5&#xa0;nt from the RefSeq annotations. As can be seen from the table, 8% more of DeGeCI&#x2032;s rRNA predictions were generated with this precision. For proteins and tRNAs, slightly more of the MITOS2 position predictions have a maximum deviation of 5&#xa0;nt (5% and 3%). One has to keep in mind, however, that MITOS2 also produced more FP and FN predictions for both of these categories (see <xref ref-type="table" rid="T1">Table 1</xref>). Moreover, for proteins, a possible explanation for this discrepancy could be as follows: after producing protein position predictions, MITOS2 searches the proximity of every start and end position for start and stop codons, respectively. If valid start or stop codons are found, the gene boundary predictions are adapted accordingly. While this approach presumably improves the accuracy of the protein boundary predictions, there are several lineage-specific variations in the standard genetic code for mitogenomes. Thus, to detect appropriate start and end codons for the input genome, the adequate genetic code table must be specified by the user. Since DeGeCI tries to minimize the required amount of knowledge about the input genome and/or user expertise, this is currently not implemented in the DeGeCI pipeline. However, we would like to point out that such an extension would be possible. As soon as slightly larger deviations are considered (i.e., &#x394; nt &#x2265;10), DeGeCI again produces more protein predictions than MITOS2 (see bold entries in <xref ref-type="table" rid="T2">Table 2</xref>).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Percentages of corresponding DeGeCI/MITOS2 and RefSeq predictions with start and stop positions of less than &#x394; nt. The higher rates are highlighted in bold.</p>
</caption>
<table>
<thead valign="top">
<tr>
<td rowspan="2" align="right">&#x394; nt</td>
<th colspan="2" align="center">Protein [%]</th>
<th colspan="2" align="center">tRNA [%]</th>
<th colspan="2" align="center">rRNA [%]</th>
</tr>
<tr>
<td align="right">DeGeCI</td>
<td align="right">MITOS2</td>
<td align="right">DeGeCI</td>
<td align="right">MITOS2</td>
<td align="right">DeGeCI</td>
<td align="right">MITOS2</td>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="right">5</td>
<td align="right">85</td>
<td align="right">
<bold>90</bold>
</td>
<td align="right">96</td>
<td align="right">
<bold>99</bold>
</td>
<td align="right">
<bold>78</bold>
</td>
<td align="right">70</td>
</tr>
<tr>
<td align="right">10</td>
<td align="right">
<bold>94</bold>
</td>
<td align="right">92</td>
<td align="right">98</td>
<td align="right">
<bold>100</bold>
</td>
<td align="right">
<bold>80</bold>
</td>
<td align="right">79</td>
</tr>
<tr>
<td align="right">25</td>
<td align="right">
<bold>97</bold>
</td>
<td align="right">94</td>
<td align="right">100</td>
<td align="right">100</td>
<td align="right">88</td>
<td align="right">
<bold>90</bold>
</td>
</tr>
<tr>
<td align="right">50</td>
<td align="right">
<bold>99</bold>
</td>
<td align="right">97</td>
<td align="right">100</td>
<td align="right">100</td>
<td align="right">94</td>
<td align="right">
<bold>96</bold>
</td>
</tr>
<tr>
<td align="right">70</td>
<td align="right">
<bold>99</bold>
</td>
<td align="right">98</td>
<td align="right">100</td>
<td align="right">100</td>
<td align="right">
<bold>100</bold>
</td>
<td align="right">98</td>
</tr>
<tr>
<td align="right">100</td>
<td align="right">
<bold>100</bold>
</td>
<td align="right">99</td>
<td align="right">100</td>
<td align="right">100</td>
<td align="right">100</td>
<td align="right">100</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4-5">
<title>4.5 RefSeq 204 as a reference database</title>
<p>We also validated DeGeCI&#x2032;s performance using the more recent RefSeq release RefSeq 204, which contains 9,877 species, for the database graph. Using this larger database, two previous FNs and one FP could be eliminated in the sample set. None of the remaining predictions was impaired (for details, see <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s11">Supplementary Table S18</xref>). Since MITOS2 currently only offers prepared databases for RefSeq 39, RefSeq 63, and RefSeq 89, a comparative analysis with MITOS2 could not be carried out for RefSeq 204.</p>
</sec>
<sec id="s4-6">
<title>4.6 Runtime and scalability</title>
<p>To generate its reference database, MITOS2 needs to retrieve amino acid sequences from the RefSeq release to be used. From these sequences, a new BLAST database needs to be built. Moreover, HMM models need to be generated that use these sequences together with their phylogenetic classification. Lastly, new covariance models need to be built for RNAs, which require manual user interaction. All these steps taken together render database updates a rather tedious task. Contrarily, DeGeCI allows for a fully automated effortless inclusion of additional species to the existing database or the creation of a new database (for details, see <xref ref-type="sec" rid="s11">Supplementary Material</xref>: <xref ref-type="sec" rid="s4">Section 4</xref>), facilitating keeping pace with the increasing amount of available mitochondrial sequence data.</p>
<p>Another important aspect to consider is the scalability of the time requirements for the annotation process with respect to the database size. To compare the impact of the database size on the runtime of both tools, DeGeCI and MITOS2, the sample set was annotated using different RefSeq releases with a different number of species as a reference database. <xref ref-type="table" rid="T3">Table 3</xref> shows the average runtimes for the annotation of one input genome on a computer with an AMD Ryzen<sup>TM</sup> 7 1700 processor with 3&#xa0;GHz. Since MITOS2 cannot be run in parallel, DeGeCI was run in the single-thread mode for a fair comparison. MITOS2 shows a clear increase in runtime with larger RefSeq releases, whereas DeGeCI is hardly affected by the database growth. DeGeCI also almost always runs considerably faster, e.g., even more than six times as fast as RefSeq 89. The only exception is for RefSeq 39. This is because the small number of species in this RefSeq release results in large unmapped regions in the subgraph, causing comparably long bridging times. DeGeCI runtimes were also measured for RefSeq 204, which includes even more species, showing that the trend of hardly impacted runtimes continues.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Comparison of runtimes. Here, the number <inline-formula id="inf90">
<mml:math id="m92">
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="monospace">R</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">f</mml:mi>
<mml:mi mathvariant="monospace">S</mml:mi>
<mml:mi mathvariant="monospace">e</mml:mi>
<mml:mi mathvariant="monospace">q</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of entries in the respective RefSeq release and the average runtime <inline-formula id="inf91">
<mml:math id="m93">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">DeGeCI</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of DeGeCI with one thread and <inline-formula id="inf92">
<mml:math id="m94">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">MITOS</mml:mtext>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> of MITOS2 (cannot be run in parallel) are shown. Except for RefSeq 39, DeGeCI runs noticeably faster than MITOS2, and increases in the database size hardly impact the runtime.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="right">Number of species</th>
<th align="right">
<inline-formula id="inf93">
<mml:math id="m95">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">DeGeCI</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (one thread) [min]</th>
<th align="right">
<inline-formula id="inf94">
<mml:math id="m96">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">MITOS</mml:mtext>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> [min]</th>
<th align="right">
<inline-formula id="inf95">
<mml:math id="m97">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">MITOS</mml:mtext>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>t</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mtext mathvariant="monospace">DeGeCI</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">RefSeq 39</td>
<td align="right">1,878</td>
<td align="right">4.22</td>
<td align="right">3.77</td>
<td align="right">0.89</td>
</tr>
<tr>
<td align="left">RefSeq 63</td>
<td align="right">3,842</td>
<td align="right">2.74</td>
<td align="right">9.11</td>
<td align="right">3.32</td>
</tr>
<tr>
<td align="left">RefSeq 89</td>
<td align="right">8,015</td>
<td align="right">2.32</td>
<td align="right">14.40</td>
<td align="right">6.21</td>
</tr>
<tr>
<td align="left">RefSeq 204</td>
<td align="right">9,877</td>
<td align="right">2.39</td>
<td align="right">&#x2013;</td>
<td align="right">&#x2013;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Nowadays, every modern CPU offers multiple hardware threads. DeGeCI has the advantage that it can (unlike MITOS2) be run in parallel. Preliminary tests on RefSeq 89 with two threads resulted in an average runtime of 1.4&#xa0;min compared to 2.32&#xa0;min for one thread. This corresponds to a speed-up value of approximately 1.66 and an efficiency value of approximately 0.83.</p>
</sec>
</sec>
<sec sec-type="discussion" id="s5">
<title>5 Discussion</title>
<p>This contribution describes a new method, DeGeCI, for an efficient, automatic <italic>de novo</italic> annotation of mitochondrial genomes. The underlying reference database, which comprises a comprehensive collection of mitogenomes, is represented as a richly annotated de Bruijn graph. To generate gene predictions for an input genome sequence, DeGeCI utilizes a clustering technique that is based on the mapping information of this sequence to the graph. DeGeCI produces gene predictions of high conformity with expert-curated annotations of RefSeq, particularly showing the high precision of gene boundaries. Compared to the standard annotation tool MITOS2, DeGeCI generates predictions of at least equal quality, requires less time for annotation when using larger databases, and features better database scalability. Different from MITOS2, the new method, DeGeCI, offers a fully automated update of the reference database and can be run in parallel. Altogether, we could demonstrate that DeGeCI is well suited for large-scale annotations of mitochondrial sequence data.</p>
<sec id="s5-1">
<title>5.1 Limitations</title>
<p>Similar to all database-based annotation approaches (e.g., MITOS, DOGMA, MOSAS, and MitoFish), DeGeCI requires a database containing annotated mitochondrial genome information that includes a certain minimum diversity of species to enable high-quality annotation of mitogenomes across a broad taxonomic spectrum. Otherwise, the annotation quality might be lower, and there might be many and/or relatively large unmapped sequence segments of the input genome. The latter leads to larger annotation times, since a comparably large amount of runtime is required to bridge these unmapped segments in the corresponding subgraph (cf. <xref ref-type="sec" rid="s2-2-2">Section 2.2.2</xref>).</p>
<p>DeGeCI has been developed with the prior aim in mind to embed complete genomic sequences of mitochondria. These are generally generated by applying a mixture of long-read and short-read sequencing techniques. Due to the higher error rates associated with long-read sequencing data, using only long-read sequencing data for graph construction might affect the quality of gene predictions obtained using DeGeCI.</p>
</sec>
<sec id="s5-2">
<title>5.2 Future work</title>
<p>The implementation of a taxonomic filter that would allow the use of only database species of a user-supplied taxonomic classification is our agenda for future work. The use of such a filter could be advantageous with respect to annotation time or quality when a specific taxonomic group of the input sequence is known.</p>
<p>As discussed in <xref ref-type="sec" rid="s4-4">Section 4.4</xref>, the result accuracy of the DeGeCI protein boundary predictions could likely be improved by scanning the proximity for start and stop codons. Since this requires the user to specify adequate genetic code tables and thus a certain degree of knowledge about the input sequence and/or user expertise, this is currently not part of DeGeCI. This could be implemented in a future version of DeGeCI, allowing the optional refinement of protein predictions.</p>
<p>The focus of this study has been on mitochondrial genomes. An application of the presented methods to nuclear genomes could, hence, be an interesting aspect to be explored in future studies. Suitable parameter settings for <italic>k</italic> could be determined similarly as suggested in this work.</p>
<p>Moreover, the implementation of a public web server version of DeGeCI is planned.</p>
</sec>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s6">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. These data can be found at: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.8101631">https://doi.org/10.5281/zenodo.8101631</ext-link> (RefSeq mitochondrial nucleotide sequences for releases 89 and 204).</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>LF, MM, and MB conceived the idea. MM and MB supervised the study and edited the manuscript. LF implemented the software, performed the computational analysis, and drafted the manuscript. All authors collaborated on the design of the algorithms and the overall workflow, the interpretation of the results, and the writing of the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>The research reported in this manuscript was funded by the Open Access Publishing Fund of Leipzig University supported by the German Research Foundation within the program Open Access Publication Funding. The authors further wish to thank the German Research Foundation for funding project 21210538.</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2023.1250907/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2023.1250907/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<sec id="s12">
<title>Abbreviations</title>
<p>DeGeCI, De Bruijn graph Gene Cluster Identification; CC, connected component; FN, false negative; FP, false positive; RFD, relative frequency distribution; SGT, single genome trail.</p>
</sec>
<fn-group>
<fn id="fn2">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://git.informatik.uni-leipzig.de/lfiedler/degeci">https://git.informatik.uni-leipzig.de/lfiedler/degeci</ext-link>
</p>
</fn>
<fn id="fn3">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.7010767">https://doi.org/10.5281/zenodo.7010767</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Almodaresi</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Pandey</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Patro</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Rainbowfish: a succinct colored de Bruijn graph representation</article-title>,&#x201d; in <conf-name>17th International Workshop on Algorithms in Bioinformatics (WABI 2017)</conf-name>, <conf-loc>Boston, MA, USA</conf-loc>, <conf-date>August 21&#x2013;23, 2017</conf-date>, <fpage>15</fpage>. <comment>1&#x2013;18</comment>. <pub-id pub-id-type="doi">10.4230/LIPIcs.WABI.2017.18</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Altschul</surname>
<given-names>S. F.</given-names>
</name>
<name>
<surname>Gish</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Myers</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Basic local alignment search tool</article-title>. <source>J. Mol. Biol.</source> <volume>215</volume>, <fpage>403</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Benson</surname>
<given-names>D. A.</given-names>
</name>
<name>
<surname>Karsch-Mizrachi</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Lipman</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Ostell</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rapp</surname>
<given-names>B. A.</given-names>
</name>
<name>
<surname>Wheeler</surname>
<given-names>D. L.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>GenBank</article-title>. <source>Nucleic Acids Res.</source> <volume>28</volume>, <fpage>15</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1093/nar/28.1.15</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernt</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Donath</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>J&#xfc;hling</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Externbrink</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Florentz</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fritzsch</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Mitos: improved de novo metazoan mitochondrial genome annotation</article-title>. <source>Mol. Phylogenetics Evol.</source> <volume>69</volume>, <fpage>313</fpage>&#x2013;<lpage>319</lpage>. <pub-id pub-id-type="doi">10.1016/j.ympev.2012.08.023</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boore</surname>
<given-names>J. L.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Requirements and standards for organelle genome databases</article-title>. <source>OMICS A J. Integr. Biol.</source> <volume>10</volume>, <fpage>119</fpage>&#x2013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1089/omi.2006.10.119</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Bowe</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Onodera</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Sadakane</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Shibuya</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2012</year>). &#x201c;<article-title>Succinct de Bruijn graphs</article-title>,&#x201d; in <source>Algorithms in bioinformatics</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Raphael</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
</person-group> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>). <comment>225&#x2013;235</comment>. <pub-id pub-id-type="doi">10.1007/978-3-642-33122-0_18</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bruijn</surname>
<given-names>de, N. G.</given-names>
</name>
</person-group> (<year>1946</year>). <article-title>A combinatorial problem</article-title>. <source>Proc. Sect. Sci. Koninklijke Nederl. Akademie van Wetenschappen te Amsterdam</source> <volume>49</volume>, <fpage>758</fpage>&#x2013;<lpage>764</lpage>.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Donath</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>J&#xfc;hling</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Al-Arab</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bernhart</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Reinhardt</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Stadler</surname>
<given-names>P. F.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Improved annotation of protein-coding genes boundaries in metazoan mitochondrial genomes</article-title>. <source>Nucleic acids Res.</source> <volume>47</volume>, <fpage>10543</fpage>&#x2013;<lpage>10552</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkz833</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>A memory-efficient dynamic programming algorithm for optimal alignment of a sequence to an rna secondary structure</article-title>. <source>BMC Bioinforma.</source> <volume>3</volume>, <fpage>1</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-3-18</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Accelerated profile hmm searches</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>, <fpage>e1002195</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1002195</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
<name>
<surname>Durbin</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>RNA sequence analysis using covariance models</article-title>. <source>Nucleic Acids Res.</source> <volume>22</volume>, <fpage>2079</fpage>&#x2013;<lpage>2088</lpage>. <pub-id pub-id-type="doi">10.1093/nar/22.11.2079</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Good</surname>
<given-names>I. J.</given-names>
</name>
</person-group> (<year>1946</year>). <article-title>Normal recurring decimals</article-title>. <source>J. Lond. Math. Soc.</source> <volume>s1-21</volume>, <fpage>167</fpage>&#x2013;<lpage>169</lpage>. <pub-id pub-id-type="doi">10.1112/jlms/s1-21.3.167</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Iwasaki</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Fukunaga</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Isagozawa</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Yamada</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Maeda</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Satoh</surname>
<given-names>T. P.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>MitoFish and MitoAnnotator: a mitochondrial genome database of fish with an accurate and automatic annotation pipeline</article-title>. <source>Mol. Biol. Evol.</source> <volume>30</volume>, <fpage>2531</fpage>&#x2013;<lpage>2540</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst141</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pevzner</surname>
<given-names>P. A.</given-names>
</name>
</person-group> (<year>2014</year>). &#x201c;<article-title>Manifold de Bruijn graphs</article-title>,&#x201d; in <source>s in bioinformatics</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Brown</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Morgenstern</surname>
<given-names>B.</given-names>
</name>
</person-group> (<publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>296</fpage>&#x2013;<lpage>310</lpage>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lowe</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence</article-title>. <source>Nucleic Acids Res.</source> <volume>25</volume>, <fpage>955</fpage>&#x2013;<lpage>964</lpage>. <pub-id pub-id-type="doi">10.1093/nar/25.5.955</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nawrocki</surname>
<given-names>E. P.</given-names>
</name>
<name>
<surname>Kolbe</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Eddy</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Infernal 1.0: inference of RNA alignments</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>1335</fpage>&#x2013;<lpage>1337</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp157</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pevzner</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tesler</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>De novo repeat classification and fragment assembly</article-title>. <source>Genome Res.</source> <volume>14</volume>, <fpage>1786</fpage>&#x2013;<lpage>1796</lpage>. <pub-id pub-id-type="doi">10.1101/gr.2395204</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pevzner</surname>
<given-names>P. A.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Waterman</surname>
<given-names>M. S.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>An Eulerian path approach to DNA fragment assembly</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>98</volume>, <fpage>9748</fpage>&#x2013;<lpage>9753</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.171285098</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pruitt</surname>
<given-names>K. D.</given-names>
</name>
<name>
<surname>Tatusova</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Maglott</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>NCBI reference sequences (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins</article-title>. <source>Nucleic Acids Res.</source> <volume>2007</volume>, <fpage>D61</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkl842</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sheffield</surname>
<given-names>N. C.</given-names>
</name>
<name>
<surname>Hiatt</surname>
<given-names>K. D.</given-names>
</name>
<name>
<surname>Valentine</surname>
<given-names>M. C.</given-names>
</name>
<name>
<surname>Song</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Whiting</surname>
<given-names>M. F.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Mitochondrial genomics in orthoptera using mosas</article-title>. <source>Mitochondrial DNA</source> <volume>21</volume>, <fpage>87</fpage>&#x2013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.3109/19401736.2010.500812</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Stonebraker</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>1987</year>). <source>The design of the Postgres storage system</source>. <comment>Tech. rep.</comment> <publisher-loc>California</publisher-loc>: <publisher-name>Univ Berkeley Electronics Research Lab California</publisher-name>.</citation>
</ref>
<ref id="B22">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Veith</surname>
<given-names>A. d. S.</given-names>
</name>
<name>
<surname>de Assuncao</surname>
<given-names>M. D.</given-names>
</name>
</person-group> (<year>2019</year>). <source>Apache Spark</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer International Publishing</publisher-name>. <pub-id pub-id-type="doi">10.1007/978-3-319-77525-8_37</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Wolstenholme</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>1992</year>). <source>Animal mitochondrial DNA: Structure and evolution</source>. <publisher-loc>United States</publisher-loc>: <publisher-name>Academic Press</publisher-name>, <fpage>173</fpage>&#x2013;<lpage>216</lpage>. <pub-id pub-id-type="doi">10.1016/S0074-7696(08)62066-5</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wyman</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Jansen</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>Boore</surname>
<given-names>J. L.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Automatic annotation of organellar genomes with DOGMA</article-title>. <source>Bioinformatics</source> <volume>20</volume>, <fpage>3252</fpage>&#x2013;<lpage>3255</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bth352</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zerbino</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Birney</surname>
<given-names>E.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Velvet: algorithms for de novo short read assembly using de Bruijn graphs</article-title>. <source>Genome Res.</source> <volume>18</volume>, <fpage>821</fpage>&#x2013;<lpage>829</lpage>. <pub-id pub-id-type="doi">10.1101/gr.074492.107</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>