<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2022.779830</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Software Choice and Sequencing Coverage Can Impact Plastid Genome Assembly&#x02013;A Case Study in the Narrow Endemic <italic>Calligonum bakuense</italic></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Giorgashvili</surname> <given-names>Eka</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Reichel</surname> <given-names>Katja</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn002"><sup>&#x02020;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1493158/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Caswara</surname> <given-names>Calvinna</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Kerimov</surname> <given-names>Vuqar</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1692594/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Borsch</surname> <given-names>Thomas</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1698351/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Gruenstaeudl</surname> <given-names>Michael</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1304594/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Systematische Botanik und Pflanzengeographie, Institut f&#x000FC;r Biologie, Freie Universit&#x000E4;t Berlin</institution>, <addr-line>Berlin</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Institute of Botany, Azerbaijan National Academy of Sciences (ANAS)</institution>, <addr-line>Baku</addr-line>, <country>Azerbaijan</country></aff>
<aff id="aff3"><sup>3</sup><institution>Botanischer Garten und Botanisches Museum Berlin, Freie Universit&#x000E4;t Berlin</institution>, <addr-line>Berlin</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Susann Wicke, Humboldt University of Berlin, Germany</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Peter Poczai, University of Helsinki, Finland; Jacob B. Landis, Cornell University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Michael Gruenstaeudl <email>m.gruenstaeudl&#x00040;fu-berlin.de</email></corresp>
<fn fn-type="other" id="fn001"><p>This article was submitted to Plant Systematics and Evolution, a section of the journal Frontiers in Plant Science</p></fn>
<fn fn-type="other" id="fn002"><p>&#x02020;These authors have contributed equally to this work and share first authorship</p></fn></author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>07</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>779830</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>09</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>06</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2022 Giorgashvili, Reichel, Caswara, Kerimov, Borsch and Gruenstaeudl.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Giorgashvili, Reichel, Caswara, Kerimov, Borsch and Gruenstaeudl</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<p>Most plastid genome sequences are assembled from short-read whole-genome sequencing data, yet the impact that sequencing coverage and the choice of assembly software can have on the accuracy of the resulting assemblies is poorly understood. In this study, we test the impact of both factors on plastid genome assembly in the threatened and rare endemic shrub <italic>Calligonum bakuense</italic>. We aim to characterize the differences across plastid genome assemblies generated by different assembly software tools and levels of sequencing coverage and to determine if these differences are large enough to affect the phylogenetic position inferred for <italic>C. bakuense</italic> compared to congeners. Four assembly software tools (FastPlast, GetOrganelle, IOGA, and NOVOPlasty) and seven levels of sequencing coverage across the plastid genome (original sequencing depth, 2,000x, 1,000x, 500x, 250x, 100x, and 50x) are compared in our analyses. The resulting assemblies are evaluated with regard to reproducibility, contig number, gene complement, inverted repeat length, and computation time; the impact of sequence differences on phylogenetic reconstruction is assessed. Our results show that software choice can have a considerable impact on the accuracy and reproducibility of plastid genome assembly and that GetOrganelle produces the most consistent assemblies for <italic>C. bakuense</italic>. Moreover, we demonstrate that a sequencing coverage between 500x and 100x can reduce both the sequence variability across assembly contigs and computation time. When comparing the most reliable plastid genome assemblies of <italic>C. bakuense</italic>, a sequence difference in only three nucleotide positions is detected, which is less than the difference potentially introduced through software choice.</p></abstract>
<kwd-group>
<kwd>assembly software</kwd>
<kwd><italic>Calligonum</italic></kwd>
<kwd>genome assembly</kwd>
<kwd>plastid genome</kwd>
<kwd>phylogenetic position</kwd>
<kwd>nucleotide differences</kwd>
<kwd>reproducibility</kwd>
<kwd>sequencing coverage</kwd>
</kwd-group>
<contract-num rid="cn001">AZ 89 950</contract-num>
<contract-sponsor id="cn001">Volkswagen Foundation<named-content content-type="fundref-id">10.13039/501100001663</named-content></contract-sponsor>
<counts>
<fig-count count="6"/>
<table-count count="3"/>
<equation-count count="0"/>
<ref-count count="84"/>
<page-count count="22"/>
<word-count count="15310"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1. Introduction</title>
<p>The comparative analysis of complete plastid genomes is performed in numerous investigations every year, even though the computational assembly of these genomes has not yet been perfected. Complete plastid genomes constitute a popular information source in various areas of plant evolutionary research, including phylogenetics (e.g., Xu et al., <xref ref-type="bibr" rid="B81">2019</xref>; Koehler et al., <xref ref-type="bibr" rid="B39">2020</xref>), phylogeography (e.g., Moner et al., <xref ref-type="bibr" rid="B50">2018</xref>; del Valle et al., <xref ref-type="bibr" rid="B17">2019</xref>), and population genetics (e.g., Yang et al., <xref ref-type="bibr" rid="B83">2013</xref>; Rogalski et al., <xref ref-type="bibr" rid="B57">2015</xref>). In recent years, the sequencing and comparison of dozens, if not hundreds, of complete plastid genomes per investigation have become commonplace (e.g., Saarela et al., <xref ref-type="bibr" rid="B59">2018</xref>; Huang et al., <xref ref-type="bibr" rid="B29">2019</xref>). Most of these studies generate complete plastid genomes from short-read whole-genome sequencing data (i.e., &#x0201C;genome skimming&#x0201D; data; Bakker, <xref ref-type="bibr" rid="B5">2017</xref>; Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>). Several specialized software tools for the <italic>de novo</italic> assembly of plastid genomes from genome skimming data exist (e.g., Coissac, <xref ref-type="bibr" rid="B16">2017</xref>; Izan et al., <xref ref-type="bibr" rid="B31">2017</xref>; McKain and Wilson, <xref ref-type="bibr" rid="B48">2017</xref>), but the process of generating complete and accurate assemblies from such data remains challenging (Wu et al., <xref ref-type="bibr" rid="B80">2015</xref>; Freudenthal et al., <xref ref-type="bibr" rid="B22">2020</xref>). For example, the use of genome skimming data for plastid genome assembly requires the separation of reads from different genomic compartments of the cell (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>). If done bioinformatically, this separation is only as accurate as the employed reference genome and its similarity to the target genome (Izan et al., <xref ref-type="bibr" rid="B31">2017</xref>; Jin et al., <xref ref-type="bibr" rid="B33">2020</xref>). Similarly, employing genome skimming data for plastid genome assembly often necessitates the use of read sets that cover the plastid genome with unequal sequencing coverage (Doorduin et al., <xref ref-type="bibr" rid="B19">2011</xref>; Izan et al., <xref ref-type="bibr" rid="B31">2017</xref>). Unequal sequencing coverage runs contrary to the implicit assumption of many genome assembly algorithms that the input reads should cover the target genome homogeneously (Peng et al., <xref ref-type="bibr" rid="B55">2012</xref>; McCorrison et al., <xref ref-type="bibr" rid="B46">2014</xref>; Olson et al., <xref ref-type="bibr" rid="B53">2019</xref>); while primarily observed for the assembly of nuclear genomes, this assumption also seems to be correct for the assembly of plastid genomes (e.g., Stadermann et al., <xref ref-type="bibr" rid="B70">2015</xref>; Soorni et al., <xref ref-type="bibr" rid="B66">2017</xref>). Moreover, the quadripartite structure of most plastid genomes, comprising a long (LSC) and a short (SSC) single-copy region separated by two inverted repeats (IR) (Ruhlman and Jansen, <xref ref-type="bibr" rid="B58">2014</xref>), often requires the manual circularization of linear assembly contigs (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>) because genome skimming data comprise an amalgamation of different reads, some of which support alternative junction sites (Jin et al., <xref ref-type="bibr" rid="B33">2020</xref>). Furthermore, the direction of the SSC often needs to be homogenized across plastid genomes before their comparison due to the structural heteroplasmy of these genomes (Walker et al., <xref ref-type="bibr" rid="B76">2015</xref>), and genome skimming data typically contain reads representing both configurations (Wang and Lanfear, <xref ref-type="bibr" rid="B77">2019</xref>). Several software tools have been developed to accommodate some of these challenges (e.g., Ankenbrand et al., <xref ref-type="bibr" rid="B2">2018</xref>; Carrion et al., <xref ref-type="bibr" rid="B13">2020</xref>; Wu et al., <xref ref-type="bibr" rid="B79">2021</xref>), but the process of plastid genome assembly from genome skimming data remains imperfect.</p>
<p>The choice of assembly software and the depth of sequencing coverage have been highlighted as potential sources for low assembly quality among plastid genomes, but a characterization of their impact has yet to be conducted. Several recent investigations have reported factors that may influence the accuracy of plastid genome assembly, including software choice (Freudenthal et al., <xref ref-type="bibr" rid="B22">2020</xref>) and sequencing coverage (reviewed in Gruenstaeudl and Jenke, <xref ref-type="bibr" rid="B26">2020</xref>). The choice of assembly software has been reported as a source of inconsistency in genome assembly by several previous studies (e.g., Magoc et al., <xref ref-type="bibr" rid="B44">2013</xref>; Morrison et al., <xref ref-type="bibr" rid="B51">2014</xref>). In the <italic>de novo</italic> assembly of plastid genomes from genome skimming data, such inconsistency may be associated with differences between assembly algorithms: while some software tools have implemented algorithms that conduct a cyclical sequence extension from a single &#x0201C;seed&#x0201D; sequence (e.g., Dierckxsens et al., <xref ref-type="bibr" rid="B18">2017</xref>), others employ a kmer-based construction of contigs, followed by the concatenation of multiple contigs based on sequence overlap and similarity to a reference genome (e.g., Bakker et al., <xref ref-type="bibr" rid="B6">2016</xref>; McKain and Wilson, <xref ref-type="bibr" rid="B48">2017</xref>). Accordingly, Freudenthal et al. (<xref ref-type="bibr" rid="B22">2020</xref>) found considerable differences among the results of different assembly software despite employing the same input sequence data. Interestingly, many of the assembly differences identified by Freudenthal et al. (<xref ref-type="bibr" rid="B22">2020</xref>) corresponded to competing locations or orientations of the four plastid genome regions rather than nucleotide polymorphisms. The question if alternative plastid genome assemblies generated for the same taxon could impact downstream analyses such as species identification or phylogenetic inference has so far not been addressed.</p>
<p>Differences in sequencing coverage have also been reported as a source for distinct plastid genome assemblies. Doorduin et al. (<xref ref-type="bibr" rid="B19">2011</xref>), for example, found that the number of SNPs across the plastid genomes of multiple individuals of <italic>Jacobaea vulgaris</italic> varied between different regions of the genome depending on the depth of sequencing coverage. Similarly, Kim et al. (<xref ref-type="bibr" rid="B38">2015</xref>) reported a correlation between cases of local misassembly and regions with exceptionally high sequencing coverage in plastid genomes of rice; regions of exceptionally high coverage are often associated with genome skimming data (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>). Moreover, Izan et al. (<xref ref-type="bibr" rid="B31">2017</xref>) found that regions with low sequencing coverage were not correctly assembled under default software settings in several angiosperm plastid genomes. Indeed, genome assemblies with unequal sequencing coverage are often characterized by high rates of sequencing error (Hubisz et al., <xref ref-type="bibr" rid="B30">2011</xref>). Based on these observations, some assembly pipelines pre-select sequence reads that represent a low but comparatively even sequencing coverage for genome assembly (e.g., 20x; Soorni et al., <xref ref-type="bibr" rid="B66">2017</xref>), and different studies indicated that a sequencing coverage of 30&#x02013;50x is needed at a minimum for reliable plastid genome assembly (e.g., Soorni et al., <xref ref-type="bibr" rid="B66">2017</xref>; Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>; Sharpe et al., <xref ref-type="bibr" rid="B62">2020</xref>). Sequencing coverage has, thus, been identified as an important indicator of assembly quality, especially in plastid genomes (Gruenstaeudl and Jenke, <xref ref-type="bibr" rid="B26">2020</xref>). Gu et al. (<xref ref-type="bibr" rid="B27">2016</xref>), for example, employed sequencing coverage as an indicator to refine the assembly of the plastid genome of <italic>Lagerstroemia fauriei</italic>. Despite the importance of sequencing coverage for the successful assembly of plastid genomes, few, if any, studies have aimed to characterize the resulting assembly differences or evaluated if those differences are large enough to impact downstream analyses.</p>
<p>In this study, we test the impact of software choice and sequencing coverage on the process of plastid genome assembly in a species for which a correct assembly is vital for conservation efforts. Specifically, we use the threatened and narrow endemic shrub <italic>Calligonum bakuense</italic> (Polygonaceae) as a test case for evaluating the variability in plastid genome assembly caused by software choice and levels of sequencing coverage. The entire species comprises only 170&#x02013;200 individuals which are currently inhabiting approximately seven localities around the Absheron Peninsula near Baku, the capital city of the Republic of Azerbaijan. <italic>Calligonum bakuense</italic> represents an exemplary case where a precise assembly of the plastid genome is of great importance to delineate the species, determine its correct phylogenetic placement relative to other members of the genus, and assess its genetic diversity at the population level. Genomic information on <italic>C. bakuense</italic> is currently absent, and documenting its complete plastid genome would be an important asset for future investigations on this rare and declining species. In this study, we use genome skimming data from two individuals of <italic>C. bakuense</italic> to characterize differences across genome assemblies in response to the choice of assembly software and levels of sequencing coverage. Specifically, we test whether the plastid genome assembly of <italic>C. bakuense</italic> is consistent across four commonly employed assembly software tools and seven different levels of sequencing coverage, and if any differences among the resulting assemblies can potentially affect the outcome of phylogenetic tree reconstruction. Based on our findings, we discuss the consequences that differences in plastid genome assembly of the magnitude detected here could have on biological conclusions and we make recommendations to optimize the assembly of complete plastid genomes.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>2. Materials and Methods</title>
<sec>
<title>2.1. Biology and Distribution of <italic>Calligonum bakuense</italic></title>
<p><italic>Calligonum bakuense</italic> LITV. is a psammophytic shrub endemic to coastal sand dune areas along the western Caspian shoreline near the city of Baku (Karjagin, <xref ref-type="bibr" rid="B35">1952</xref>; Soskov and Akhmed-Zade, <xref ref-type="bibr" rid="B67">1974</xref>). The species is a unique and declining element of the flora of Azerbaijan and of high conservation interest (Atamov, <xref ref-type="bibr" rid="B3">2008</xref>). It currently comprises a total of seven wild populations that are distributed across a distance of approximately 120 km and collectively contain roughly 170&#x02013;200 individuals (<xref ref-type="fig" rid="F1">Figure 1</xref>). Here we assemble and report the plastid genomes of two individuals that represent the northern- and the southernmost localities of the current distribution area. The evolutionary relationships of <italic>C. bakuense</italic> to other members of <italic>Calligonum</italic> are currently unknown, as is the population structure within the species. <italic>Calligonum</italic> L. is a lineage of xerophytic shrubs with an estimated 30&#x02013;40 species; it is distributed from northern Africa, the Arab Peninsula, South West Asia, the Caucasus, the Irano-Turanian region, and Central Asia to China (Brandbyge, <xref ref-type="bibr" rid="B10">1993</xref>; Abdellaoui et al., <xref ref-type="bibr" rid="B1">2011</xref>). Several species of the genus are globally red-listed and exhibit declining population sizes (Baillie et al., <xref ref-type="bibr" rid="B4">2004</xref>). Currently, there is no comprehensive molecular phylogeny of <italic>Calligonum</italic>, but complete plastid genome sequences have been shown as a promising basis for inferring phylogenetic relationships among Chinese members of the genus (Song et al., <xref ref-type="bibr" rid="B65">2020</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Habit and natural environment <bold>(A)</bold> and the current distribution area <bold>(B)</bold> of <italic>C. bakuense</italic>. The map indicates the localities of all sampled natural populations of <italic>C. bakuense</italic>, including those that individuals Cb01A and Cb04B were sampled from.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0001.tif"/>
</fig>
</sec>
<sec>
<title>2.2. DNA Extraction and Genome Skimming</title>
<p>For a conservation genetic study on <italic>C. bakuense</italic>, silica-dried tissue samples from all known individuals of the species were collected between 2013 and 2015. One individual (Cb01A) that represents the northernmost and one (Cb04B) that represents the southernmost locality of the current distribution area were selected for low-coverage whole-genome sequencing (<xref ref-type="fig" rid="F1">Figure 1</xref>). Whole-genomic DNA of each individual was extracted using a modified CTAB protocol (Borsch et al., <xref ref-type="bibr" rid="B9">2003</xref>), sheared via ultrasonication to an average fragment size of &#x0007E;300 bp, and converted to a barcoded genomic library using the Illumina TruSeq DNA sample preparation kit (Illumina, San Diego, CA, USA) under the high sample protocol of the manufacturer. The DNA of both individuals was pooled equimolarly and sequenced on a full Illumina HiSeq 4000 plate by Macrogen Inc. (Seoul, South Korea). If evenly distributed, this amount of sequence data would cover the nuclear genome (1C) of each individual with an average sequencing coverage of 18&#x02013;20x. After sequencing, low-quality bases (phred-score &#x0003C;20) and remnants of Illumina adapter sequences were trimmed from the raw reads with Cutadapt v. 1.14 (Martin, <xref ref-type="bibr" rid="B45">2011</xref>). While primarily intended for the development of genetic markers in the nuclear genome, this sequence data also comprises a high number of reads representing the plastid genome, rendering the data ideal to evaluate the impact of software choice and the depth of sequencing coverage on plastid genome assembly.</p>
</sec>
<sec>
<title>2.3. Computational Extraction of Plastid Genome Reads</title>
<p>Paired sequence reads of the plastid genome were bioinformatically extracted from the raw sequence data as input for plastid genome assembly. This extraction was primarily conducted due to the large number of raw sequence reads generated, which exceeded the maximum capacity of some of the assembly software tools employed (e.g., IOGA terminates with a memory error when operating on the raw sequence data). Hence, we mapped the raw sequence reads to a set of related, previously published plastid genomes and then extracted and retained only the successfully mapped, paired reads using script 5 of the pipeline described in Gruenstaeudl et al. (<xref ref-type="bibr" rid="B25">2018</xref>). Since the phylogenetic position of <italic>C. bakuense</italic> has not yet been evaluated on a molecular basis, we selected a taxonomically broad set of twelve plastid genomes of the Caryophyllales as reference genomes: <italic>Fagopyrum esculentum</italic> subsp. <italic>ancestrale</italic> (GenBank accession number <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_010776">NC_010776</ext-link>), <italic>Fallopia multiflora</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_041239">NC_041239</ext-link>), <italic>Rumex acetosa</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_042390">NC_042390</ext-link>), <italic>Muehlenbeckia australis</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="MG604297">MG604297</ext-link>), <italic>Oxyria sinensis</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_032031">NC_032031</ext-link>), and <italic>Rheum palmatum</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_027728">NC_027728</ext-link>, all Polygonaceae); <italic>Amaranthus hypochondriacus</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_030770">NC_030770</ext-link>) and <italic>Chenopodium quinoa</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_034949">NC_034949</ext-link>, both Amaranthaceae); <italic>Mesembryanthemum crystallinum</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_029049">NC_029049</ext-link>, Aizoaceae); <italic>Carnegiea gigantea</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_027618">NC_027618</ext-link>, Cactaceae); <italic>Dianthus caryophyllus</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_039650">NC_039650</ext-link>, Caryophyllaceae); and <italic>Nyctaginia capitata</italic> (<ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="NC_041415">NC_041415</ext-link>, Nyctaginaceae).</p>
</sec>
<sec>
<title>2.4. Capping of Sequencing Coverage</title>
<p>To determine if the process of plastid genome assembly is consistent across different levels of sequencing coverage, we created subsets of the plastid genome read set with a lower average sequencing coverage (hereafter &#x0201C;capped read sets&#x0201D;). Following Sims et al. (<xref ref-type="bibr" rid="B64">2014</xref>), we use the term "sequencing depth" to specifically denote the average sequencing coverage of a genome or genome region hereafter. We capped the sequencing coverage of the plastid genome at six different levels, which collectively represent the range of sequencing depths typically encountered in genome skimming: 2,000x, 1,000x, 500x, 250x, 100x, and 50x. The largest evaluated cap level of sequencing coverage (i.e., 2,000x) represents approximately 25% of the uncapped sequencing depth of the plastid genome and indicates roughly the maximum input capacity of the assembly software tool (i.e., IOGA) that was found to require the largest amount of primary memory for the plastid genome assembly of <italic>C. bakuense</italic>. The smallest evaluated cap level of sequencing coverage (i.e., 50x) represents less than 1% of the uncapped sequencing depth of the plastid genome and is located between the minimum and the desirable sequencing coverage for plastid genome sequencing according to Twyford and Ness (<xref ref-type="bibr" rid="B74">2017</xref>). Bioinformatically, the cap in sequencing coverage was not a hard threshold above which all additional reads were removed, but a soft threshold above which additional reads were progressively curtailed. In contrast to a hard cap, such a soft cap of sequencing coverage generates a coverage distribution similar to that of empirical sequence data. For example, under a soft cap of 1,000x, some nucleotides of the target genome are supported by more than 1,000 mapped reads and some by less, whereas under a hard cap of 1,000x, none of the nucleotides of the target genome are supported by more than 1,000 mapped reads but some still by less. Technically, this soft cap was implemented through a one-tailed normalization of the sequencing coverage using the script &#x00027;bbnorm.sh&#x00027; of the software BBtools v.33.89 (Bushnell, <xref ref-type="bibr" rid="B11">2015</xref>) under default settings and using the plastid genome of <italic>Calligonum caput-medusae</italic> (MN202600; Song et al., <xref ref-type="bibr" rid="B65">2020</xref>) as a structural reference. All capped read sets were treated identically to the uncapped read set during the genome assembly and all subsequent analyses. For greater efficiency, the read sets capped at 2,000x and 500x were employed as representatives for all six cap levels in the evaluation of the suboptimal assembly software tools FastPlast and IOGA as well as the impact of seed sequence selection on plastid genome assembly.</p>
</sec>
<sec>
<title>2.5. Genome Assembly</title>
<p>To determine if the process of plastid genome assembly for <italic>C. bakuense</italic> is consistent across different assembly software tools, we compared the assembly results of four commonly-used tools: NOVOPlasty v.3.8.3 (Dierckxsens et al., <xref ref-type="bibr" rid="B18">2017</xref>), GetOrganelle v.1.6.4 (Jin et al., <xref ref-type="bibr" rid="B33">2020</xref>), FastPlast v.1.2.8 (McKain and Wilson, <xref ref-type="bibr" rid="B48">2017</xref>), and IOGA v.38.26 (Bakker et al., <xref ref-type="bibr" rid="B6">2016</xref>). Each of these software tools had been designed for the <italic>de novo</italic> assembly of plastid genomes from short sequence reads and had demonstrated its utility in previous plastid genomic studies (reviewed in Freudenthal et al., <xref ref-type="bibr" rid="B22">2020</xref>). To improve the comparability of the assembly process across these tools, we employed each software under its default settings. To ensure a uniform software execution and to compare computation times across the tools, all assemblies of <italic>C. bakuense</italic> were conducted on the high-performance computer cluster &#x00027;Curta&#x00027; of the Freie Universit&#x000E4;t Berlin under the following settings: a single 64-bit processor, an allotment of 2 GB of RAM, and a disk I/O speed of 129 MB/s. The raw output, as well as the log file of each assembly run, are available on Zenodo under <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/record/6577786">https://zenodo.org/record/6577786</ext-link>.</p>
<p>In practice, plastid genome assembly software often generates multiple incomplete, linear contigs instead of one complete, circular genome sequence (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>). Incomplete contigs typically require manual intervention to be combined into a complete genome sequence (Gruenstaeudl et al., <xref ref-type="bibr" rid="B25">2018</xref>). In this study, two of the software tools produced incomplete contigs for <italic>C. bakuense</italic>. In such cases, we concatenated the incomplete contigs upon removing any end overhangs, followed by circularization of the resulting super-contig. The concatenation of contigs was conducted by hand in Geneious v.11.1.4 (Kearse et al., <xref ref-type="bibr" rid="B37">2012</xref>) through aligning each contig to the structural reference genome (<italic>C. caput-medusae</italic>) and then sorting the contigs according to their relative position. If adjacent contigs overlapped for at least 15 bp without differences in their nucleotide sequence, they were merged into a larger contig until all such contigs were combined into a single super-contig.</p>
<p>The identification of the endpoint of a circular genome sequence is challenging for most genome assembly algorithms (but see Wu et al., <xref ref-type="bibr" rid="B79">2021</xref>) and often results in the detection of different endpoints across tools. To avoid inflating the number of differences between assemblies due to unequal endpoints, we manually corrected super-contigs if the inferred endpoints were within 100 bp across assemblies. Specifically, we searched for the first and the last 25 bp of the super-contig of each assembly via separate motif searches, with the maximum number of mismatches set to 3 bp. Any matches within 100 bp of the super-contig ends were considered to be instances where the assembly process extended the sequence beyond its actual endpoint. Such sequence motifs were removed from one of the two ends, followed by circularization of the super-contig. Similarly, poly-N motifs in contigs are often generated by plastid genome assembly software to indicate areas of sequence uncertainty. To avoid inflating the number of differences between assemblies, we automatically corrected poly-N-motifs using the software Pilon v.1.23 (Walker et al., <xref ref-type="bibr" rid="B75">2014</xref>). Moreover, most software tools for plastid genome assembly do not automatically standardize the orientation of the SSC across assemblies, even though plastid genome isomers with alternative SSC orientations naturally exist in most land plants (Walker et al., <xref ref-type="bibr" rid="B76">2015</xref>). To avoid inflating the number of differences between assemblies, we manually homogenized the orientation of the SSC across assemblies using Geneious.</p>
</sec>
<sec>
<title>2.6. Replication of Assembly Runs</title>
<p>Several software tools for plastid genome assembly constitute multi-step pipelines rather than single applications (Gruenstaeudl et al., <xref ref-type="bibr" rid="B25">2018</xref>). These pipelines typically employ a third-party assembly tool as their core assembly engine to conduct a k-mer-based alignment of reads for the inference of de Bruijn graphs (Izan et al., <xref ref-type="bibr" rid="B31">2017</xref>). FastPlast, IOGA, and GetOrganelle, for example, utilize the assembly software SPAdes (Bankevich et al., <xref ref-type="bibr" rid="B7">2012</xref>) as core assembler, even though the full reproducibility of bacterial genome assemblies with SPAdes has been called into question (e.g., Liao et al., <xref ref-type="bibr" rid="B41">2015</xref>; Souvorov et al., <xref ref-type="bibr" rid="B69">2018</xref>). To characterize potential occurrences of spurious, non-reproducible inferences of de Bruijn graphs, we conducted every plastid genome assembly of the uncapped read set twice under the same input data and software settings (i.e., replicate run &#x00023;1 and &#x00023;2). The comparison of these replicate runs allowed us to ascertain the baseline replicability of plastid genome assemblies under these assembly tools.</p>
<p>The cyclical sequence extension that starts from a single &#x0201C;seed&#x0201D; sequence and is implemented in several assembly algorithms (e.g., Dierckxsens et al., <xref ref-type="bibr" rid="B18">2017</xref>) may represent a source of contig variability not present in other assembly algorithms. Several studies have reported minor differences in the number and sequence of assembly contigs depending on the precise seed sequence employed and have, therefore, attempted to identify universally applicable seed sequences (e.g., Lim et al., <xref ref-type="bibr" rid="B42">2018</xref>; Wu et al., <xref ref-type="bibr" rid="B79">2021</xref>). To ensure that seed selection did not inflate the number of differences between assemblies, we employed the same seed sequence for each plastid genome assembly with NOVOPlasty. Moreover, we evaluated if seed selection represented a relevant source of contig variability in our dataset by repeating plastid genome assembly with NOVOPlasty and the 2,000x and 500x capped read sets under a second seed sequence. Both seeds (i.e., seed &#x00023;1 and seed &#x00023;2) were arbitrarily selected from the read set.</p>
</sec>
<sec>
<title>2.7. Sequence Annotation</title>
<p>To enable consistent sequence annotations across all plastid genome assemblies of <italic>C. bakuense</italic>, the sequence annotations from two existing plastid genomes of <italic>Calligonum</italic> were transferred to the new assemblies using Geneious. Specifically, we automatically transferred all gene, tRNA, and rRNA annotations from the plastid genomes of <italic>C. caput-medusae</italic> and <italic>C. arborescens</italic> (MN202599; both Song et al., <xref ref-type="bibr" rid="B65">2020</xref>) to the assemblies of <italic>C. bakuense</italic> based on a sequence similarity threshold of 95%. Upon transfer, we conducted a manual inspection of the transferred annotations for each coding region regarding the presence of start and stop codons, the absence of internal stop codons, and their lengths as a multiple of three. Any premature stop codon that was introduced by the transfer process but not based on the nucleotide sequence was corrected; any premature stop codon based on the nucleotide sequence was recorded as an indicator of low assembly quality. The annotations of the IRs and, by extension, of the single-copy regions were inferred for each assembly using script 4 of the pipeline of Gruenstaeudl et al. (<xref ref-type="bibr" rid="B25">2018</xref>).</p>
</sec>
<sec>
<title>2.8. Evaluation of Assembly Quality</title>
<p>To assess the quality of the plastid genome assemblies of <italic>C. bakuense</italic> and, simultaneously, the performance of each assembly software, the raw output of each assembly process was evaluated with Quast v.4.6.3 (Gurevich et al., <xref ref-type="bibr" rid="B28">2013</xref>). Specifically, we assessed and compared the number, length, and contiguity of the contigs generated by each assembly software. NGA50 and LGA50 were calculated as contiguity metrics (Earl et al., <xref ref-type="bibr" rid="B20">2011</xref>; Gurevich et al., <xref ref-type="bibr" rid="B28">2013</xref>). As part of this quality assessment, we also compared the computation times of the different assembly software tools employing different read sets. All assembly statistics were calculated after the removal of contigs smaller than 100 bp, if any, to avoid the counting of mono- or di-nucleotide fragments.</p>
</sec>
<sec>
<title>2.9. Characterization of Assembly Differences</title>
<p>To compare the different plastid genome assemblies of <italic>C. bakuense</italic> as generated by different software tools, levels of sequencing coverage, seed selection, and run replication, we conducted a series of statistical evaluations based on pairwise genetic distances. As the basis for these comparisons, we generated pairwise alignments of the assemblies using MAFFT v.7.471 (Katoh and Standley, <xref ref-type="bibr" rid="B36">2013</xref>) under default settings. We then inferred the differences in sequence as well as in length of the four genome regions (i.e., LSCs, IRb, SSCs, and IRa) for each plastid genome pair. Sequence differences were calculated as the number of single nucleotide polymorphisms (SNPs) when excluding gaps but including nucleotide ambiguities using trimAl v.1.2 (Capella-Gutierrez et al., <xref ref-type="bibr" rid="B12">2009</xref>). Length differences were calculated as the absolute difference across the lengths of different genomic regions; this metric is independent of the exact number of regions per genome but may overestimate similarity, as dissenting length changes across regions may compensate each other. Upon calculation, difference values were aggregated in a pairwise genetic distance matrix. Since only the plastid genomes assembled with GetOrganelle using the capped read sets of 500x, 250x, and 100x were identical within both samples and across all tested parameters and, thus, best-supported, we designated the assembly inferred with GetOrganelle on the 500x capped read set as the "final" plastid genome sequence for each individual of <italic>C. bakuense</italic>. To verify the length and sequence differences detected, particularly between the two final plastid genomes, we visually inspected select pairwise alignments in Geneious.</p>
<p>To visualize the genetic distances among selected assemblies and both individuals, we conducted principal coordinates analyses (PCoAs) using the uncapped dataset as well as the datasets capped at 2,000x and 500x as representatives for the complete range of different levels of sequencing coverage. In our plots, we centered the projections on the final plastid genome sequences (i.e., the assemblies generated with GetOrganelle for the read set capped at 500x), scaled the first two principal coordinates to a standard range from &#x02013;1 to 1, and displayed the absolute variance (in bp) and the percentage of total variance along each axis within the plot. Since PCoAs can potentially distort pairwise distances between data points, we also plotted an overview of the genetic distances caused by changes in software (including seed selection) and coverage cap (including run replication). Moreover, we compared the pairwise genetic distances between Cb01A (set as origin) and Cb04B across software, levels of sequencing coverage, seed selection, and run replicate as a biologically meaningful standard for the assembly differences within each individual. All calculations and visualizations based on pairwise genetic distance matrices were conducted in R v.4.0.0 (R Development Core Team, <xref ref-type="bibr" rid="B56">2019</xref>).</p>
</sec>
<sec>
<title>2.10. Visualizations of Region Length, Sequencing Coverage, and SNP Location</title>
<p>Three types of visualization were employed to illustrate the structural and sequence differences between select plastid genome assemblies of <italic>C. bakuense</italic>. First, we illustrated the length differences in the LSC, the SSC, and the two IRs across assemblies through an alignment overview of the four plastid genome regions using Geneious. Second, we visualized the depth of sequencing coverage across the entire plastid genome and in relation to the four genome regions and the position of its genes with PACVr v.1.0 using a calculation window of 250 bp (Gruenstaeudl and Jenke, <xref ref-type="bibr" rid="B26">2020</xref>). Third, we determined and visualized the location of SNPs between assemblies and in relation to changes in sequencing coverage through pairwise comparisons of each assembly to the final genome sequence using MAFFT for sequence alignment and trimAl for SNP detection. We also visualized SNP locations relative to the four genome regions and the position of its genes using ShinyCircos v.29052020 (Yu et al., <xref ref-type="bibr" rid="B84">2018</xref>). Visualizations were not produced for genome assemblies generated with GetOrganelle and NOVOPlasty under the 250x and 100x coverage cap levels, as these assemblies were identical to those generated under a coverage cap of 500x.</p>
</sec>
<sec>
<title>2.11. Phylogenetic Inference</title>
<p>To test if the sequence differences among the plastid genome assemblies of <italic>C. bakuense</italic> affect the phylogenetic placement of this species within <italic>Calligonum</italic>, we inferred the phylogenetic position of all plastid genome assemblies generated in this study among a taxonomically representative set of <italic>Calligonum</italic> species. Specifically, we retrieved 21 plastid genomes of <italic>Calligonum</italic> available from NCBI GenBank as of 30-Nov-2020 as well as the plastid genome of <italic>Rheum palmatum</italic> as an outgroup (GenBank accession <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="KR816224">KR816224</ext-link>; matching the study of Song et al., <xref ref-type="bibr" rid="B65">2020</xref>) and combined these 22 genome records with the 28 genome assemblies generated here for the two individuals of <italic>C. bakuense</italic>. Then, we bioinformatically extracted 67 protein-coding regions, 17 introns, and 104 intergenic spacers from each of the 78 genome records using script 9 of Gruenstaeudl et al. (<xref ref-type="bibr" rid="B25">2018</xref>), automatically aligned the regions using MAFFT, and manually corrected the alignments where necessary. Extracting and aligning the different coding and non-coding regions individually (instead of conducting genome-wide alignments) reduces the probability of incorrect positional homology assessments during sequence alignment, especially if the input genomes differ in size (Gruenstaeudl et al., <xref ref-type="bibr" rid="B25">2018</xref>). Even under these strict conditions, a total of 48 areas of unclear homology (mostly poly-A/T microsatellites; &#x0201C;hotspots&#x0201D; in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>) were detected and removed from the alignments during manual alignment correction. The resulting alignments were concatenated to a combined matrix and their indels coded according to the simple indel coding scheme of Simmons and Ochoterena (<xref ref-type="bibr" rid="B63">2000</xref>) using 2matrix v.1.0 (Salinas and Little, <xref ref-type="bibr" rid="B60">2014</xref>). A total of eight inversions (each less than 20 bp in length) were encountered within the alignments (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table S1</xref>); to correctly include their phylogenetic information in our analyses, we coded them as presence-absence data, included this data alongside the regular indel information, and re-integrated their reverse-complemented sequences into the nucleotide alignments. The nucleotide matrix and the indel matrix were defined as separate partitions, and the best phylogenetic tree for this combined matrix was inferred under the maximum likelihood (ML) criterion using RAxML v.8.2.9 (Stamatakis, <xref ref-type="bibr" rid="B71">2014</xref>). Clade support was inferred during tree inference through 1,000 bootstrap (BS) replicates generated with the rapid BS algorithm. To infer a phylogenetic position for <italic>C. bakuense</italic> within the genus <italic>Calligonum</italic>, we also conducted a second phylogenetic reconstruction involving only the two final plastid genome sequences of <italic>C. bakuense</italic>, the 21 genome records of <italic>Calligonum</italic> from NCBI GenBank, and the plastid genome of <italic>Rheum palmatum</italic> as an outgroup. For this second reconstruction, the best ML tree (including clade support) was inferred using RAxML as described above.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3. Results</title>
<sec>
<title>3.1. Number of Sequence Reads</title>
<p>Genome skimming of the two individuals of <italic>C. bakuense</italic> resulted in a total of 151,567,745 paired raw sequence reads for Cb01A and a total of 166,362,653 paired raw sequence reads for Cb04B. Upon extraction of the plastid genome reads, we counted 5,062,912 paired reads (3.3% of raw reads) for Cb01A and 2,998,391 paired reads (1.8%) for Cb04B. Upon capping sequencing coverage, the read sets of Cb01A comprised 1,181,510 paired reads (0.78% of raw reads) under a level of sequencing coverage of 2,000x, 590,217 paired reads (0.39%) under 1,000x, 332,662 paired reads (0.22%) under 500x, 145,294 paired reads (0.10%) under 250x, 57,034 paired reads (0.04%) under 100x, and 27,913 paired reads (0.02%) under 50x. Similarly, the read sets of Cb04B comprised 1,149,375 paired reads (0.69% of raw reads) under a coverage cap of 2,000x, 571,982 paired reads (0.34%) under 1,000x, 333,839 paired reads (0.20%) under 500x, 140,425 paired reads (0.08%) under 250x, 54,922 paired reads (0.03%) under 100x, and 26,767 paired reads (0.02%) under 50x.</p>
</sec>
<sec>
<title>3.2. Impact of Software Choice</title>
<p>The choice of assembly software had a considerable effect on the number and size of the generated assembly contigs, the contiguity of the assemblies, sequence equality of the inferred IRs, and the time required to conduct each assembly (<xref ref-type="table" rid="T1">Table 1</xref>). While some software tools assembled the complete plastid genome of <italic>C. bakuense</italic> as a single contig, others did not. GetOrganelle and IOGA represented the extremes among the tested software tools: GetOrganelle succeeded in assembling the complete plastid genome as a single contig under nearly all settings, whereas IOGA failed in this task under all settings. Even under the original sequencing depth, which is representative of low-coverage nuclear genome skimming or even small nuclear genome sequencing projects, GetOrganelle successfully assembled the complete plastid genome of <italic>C. bakuense</italic> into a single, circular contig for both individuals and run replicates, precluding the need for any manual post-processing of the contigs. Similarly, NOVOPlasty succeeded in assembling the complete plastid genome of <italic>C. bakuense</italic> as a single, circular contig under the original sequencing depth for both individuals, run replicates, and seed sequences. For Cb01A, however, the assemblies generated with NOVOPlasty exhibited considerable size variability and often exceeded the length of the final plastid genome sequence; moreover, the inferred IRs were not identical in one of the assemblies. FastPlast also succeeded in assembling the complete plastid genome of <italic>C. bakuense</italic> as a single, circular contig under the original sequencing depth. However, the contigs produced for both individuals and both replicate runs lagged or exceeded the length of the final plastid genome sequences due to incomplete or duplicated sections of the IRs, ranging from 201 to 143 kb in Cb01A and from 192 kb to 175 kb in Cb04B. The smaller than expected contig lacked a section of the IRa, whereas the larger than expected contigs exhibited a duplication of sections of the LSC adjacent to the IRs, necessitating manual post-processing of the contigs and affecting the calculation of NGA50. IOGA, by contrast, did not succeed in assembling the complete plastid genome of <italic>C. bakuense</italic> as a single, complete contig under any setting. For both individuals, it generated more than 20 separate contigs, which represented only sections of the complete genome. Hence, the IOGA contigs had to be manually concatenated for both individuals and run replicates to generate circular assemblies. Moreover, the contigs assembled by IOGA for Cb01A did not imply identical IRs in one run replicate, indicating additional assembly problems. Computation times differed strongly across software tools and&#x02014;in the case of IOGA and FastPlast&#x02014;across run replicates, but were similar across different seed sequences in NOVOPlasty. Under the original sequencing depth, GetOrganelle and NOVOPlasty were typically the fastest to generate assembly contigs, whereas FastPlast and IOGA often required a multiple of their computation time. In summary, the plastid genome assemblies generated for <italic>C. bakuense</italic> with GetOrganelle and NOVOPlasty under the original sequencing depth were more consistent and required less, if any, manual post-processing than the assemblies generated with FastPlast and IOGA. Hence, we disregarded the latter two software tools during the more detailed evaluation of the impact of sequencing coverage on plastid genome assembly (<xref ref-type="table" rid="T2">Table 2</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Assembly statistics for the plastid genomes of the two individuals of <italic>C. bakuense</italic> under study regarding the impact of assembly software choice, run replication, and seed selection.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Asmb</bold>.</th>
<th valign="top" align="center"><bold>Cov</bold>.</th>
<th valign="top" align="center"><bold>Repl</bold>.</th>
<th valign="top" align="center"><bold>NOVO seed</bold></th>
<th valign="top" align="center"><bold>Contigs</bold></th>
<th valign="top" align="center"><bold>Largest contig (bp)</bold></th>
<th valign="top" align="center"><bold>NGA50 (bp)</bold></th>
<th valign="top" align="center"><bold>LGA50</bold></th>
<th valign="top" align="center"><bold>IR equal</bold>.</th>
<th valign="top" align="center"><bold>Comp. time (h, min.)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Cb01A</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">200,694</td>
<td valign="top" align="center">118,168</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">05 h 20 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">143,261</td>
<td valign="top" align="center">135,202</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">06 h 40 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,404</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 16 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">500x</td>
<td/>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">163,292</td>
<td valign="top" align="center">162,896</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">24 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">44 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">44 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">2000x</td>
<td/>
<td/>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,241</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 08 min</td>
</tr>
<tr>
<td valign="top" align="left"><bold>GetO</bold></td>
<td valign="top" align="center"><bold>500x</bold></td>
<td/>
<td/>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>162,128</bold></td>
<td valign="top" align="center"><bold>162,128</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>20 min</bold></td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">21</td>
<td valign="top" align="center">89,039</td>
<td valign="top" align="center">88,068</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">09 h 42 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">21</td>
<td valign="top" align="center">89,039</td>
<td valign="top" align="center">88,068</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">06 h 43 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td/>
<td valign="top" align="center">83</td>
<td valign="top" align="center">129,550</td>
<td valign="top" align="center">118,520</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">07 h 50 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">500x</td>
<td/>
<td/>
<td valign="top" align="center">51</td>
<td valign="top" align="center">91,976</td>
<td valign="top" align="center">89,718</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">02 h 22 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">170,093</td>
<td valign="top" align="center">131,660</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">01 h 05 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">170,099</td>
<td valign="top" align="center">170,099</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">57 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">23 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td/>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">07 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 05 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">170,106</td>
<td valign="top" align="center">170,106</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 00 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">23 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td/>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">07 min</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left"><bold>Cb04B</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">175,272</td>
<td valign="top" align="center">175,272</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">09 h 24 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">192,943</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">03 h 36 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">163,890</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 16 min</td>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="center">500x</td>
<td/>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">163,292</td>
<td valign="top" align="center">163,292</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">24 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">04 h 16 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 04 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td/>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,238</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 28 min</td>
</tr>
<tr>
<td valign="top" align="left"><bold>GetO</bold></td>
<td valign="top" align="center"><bold>500x</bold></td>
<td/>
<td/>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>162,129</bold></td>
<td valign="top" align="center"><bold>162,129</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>20 min</bold></td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td/>
<td valign="top" align="center">54</td>
<td valign="top" align="center">90,241</td>
<td valign="top" align="center">88,240</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">11 h 14 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td/>
<td valign="top" align="center">85</td>
<td valign="top" align="center">90,630</td>
<td valign="top" align="center">87,507</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">07 h 42 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td/>
<td valign="top" align="center">102</td>
<td valign="top" align="center">55,966</td>
<td valign="top" align="center">27,790</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">08 h 17 min</td>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="center">500x</td>
<td/>
<td/>
<td valign="top" align="center">40</td>
<td valign="top" align="center">75,285</td>
<td valign="top" align="center">74,394</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">03 h 18 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 23 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 24 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">20 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td/>
<td valign="top" align="center">seed1</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">06 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl1</td>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 00 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">repl2</td>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 32 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td/>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">20 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td/>
<td valign="top" align="center">seed2</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">06 min</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>The assemblies that represent the final genome sequences are highlighted in bold. The assembly software tools compared are abbreviated as &#x0201C;FaPl&#x0201D; (for FastPlast), &#x0201C;GetO&#x0201D; (for GetOrganelles), &#x0201C;IOGA,&#x0201D; and &#x0201C;NOVO&#x0201D; (for NOVOPlasty). Run replicates are abbreviated as &#x0201C;repl1&#x0201D; or &#x0201C;repl2,&#x0201D; the original sequencing depth as &#x0201C;orig.&#x0201D; Other abbreviations used: asmb., assembly; comp., computation; cov., coverage; equal., equality in sequence; repl., replicate</italic>.</p>
</table-wrap-foot>
</table-wrap>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Assembly statistics for the plastid genomes of the two individuals of <italic>C. bakuense</italic> under study regarding the impact of different levels of sequencing coverage.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Asmb</bold>.</th>
<th valign="top" align="center"><bold>Cov</bold>.</th>
<th valign="top" align="center"><bold>Compl</bold>.</th>
<th valign="top" align="center"><bold>Contigs</bold></th>
<th valign="top" align="center"><bold>Largest contig (bp)</bold></th>
<th valign="top" align="center"><bold>NGA50 (bp)</bold></th>
<th valign="top" align="center"><bold>LGA50</bold></th>
<th valign="top" align="center"><bold>IR length (bp)</bold></th>
<th valign="top" align="center"><bold>IR equal</bold>.</th>
<th valign="top" align="center"><bold>Comp. time (h, min.)</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Cb01A</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">44 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">2,000x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,241</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 08 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">1,000x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">62,295</td>
<td valign="top" align="center">n.s.d.</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">06 min</td>
</tr>
<tr>
<td valign="top" align="left"><bold>GetO</bold></td>
<td valign="top" align="center"><bold>500x</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>162,128</bold></td>
<td valign="top" align="center"><bold>162,128</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>30,526</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>20 min</bold></td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">250x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">02 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">100x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">50x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,241</td>
<td valign="top" align="center">118,220</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">28,610</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">170,093</td>
<td valign="top" align="center">131,660</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">44,559</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">01 h 05 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">23 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">1,000x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">11 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">07 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">250x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">05 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">100x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">162,128</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">02 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">50x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">117,861</td>
<td valign="top" align="center">117,849</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">09 min</td>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left"><bold>Cb04B</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">04 h 16 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">2,000x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,238</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 28 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">1000x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">67,160</td>
<td valign="top" align="center">n.s.d.</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">06 min</td>
</tr>
<tr>
<td valign="top" align="left"><bold>GetO</bold></td>
<td valign="top" align="center"><bold>500x</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>162,129</bold></td>
<td valign="top" align="center"><bold>162,129</bold></td>
<td valign="top" align="center"><bold>1</bold></td>
<td valign="top" align="center"><bold>30,526</bold></td>
<td valign="top" align="center"><bold>Yes</bold></td>
<td valign="top" align="center"><bold>20 min</bold></td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">250x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">02 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">100x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 min</td>
</tr>
<tr>
<td valign="top" align="left">GetO</td>
<td valign="top" align="center">50x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">2</td>
<td valign="top" align="center">118,236</td>
<td valign="top" align="center">118,215</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">orig.</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">01 h 23 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">2,000x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">20 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">1,000x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">15 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">500x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">06 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">250x</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">162,129</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,476</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">04 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">100x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">4</td>
<td valign="top" align="center">112,054</td>
<td valign="top" align="center">112,054</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">30,526</td>
<td valign="top" align="center">Yes</td>
<td valign="top" align="center">07 min</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="center">50x</td>
<td valign="top" align="center">No</td>
<td valign="top" align="center">1</td>
<td valign="top" align="center">75,891</td>
<td valign="top" align="center">n.s.d.</td>
<td valign="top" align="center">-</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">n.a.</td>
<td valign="top" align="center">24 min</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p><italic>For assemblies under the original sequencing depth, only the first run replicate is displayed; for all assemblies performed with NOVOPlasty, seed sequence 1 was employed. Abbreviations used: compl., complete genome assembled; n.a., not applicable; n.s.d., no similarity detected by QUAST; all other abbreviations used as in <xref ref-type="table" rid="T1">Table 1</xref>. The assemblies that represent the final genome sequences are highlighted in bold</italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>3.3. Impact of Sequencing Coverage</title>
<p>The sequencing coverage also had a considerable effect on the number and size of the generated assembly contigs, the contiguity of the assemblies, sequence equality of the inferred IRs, and the time required to conduct each assembly. We observed that GetOrganelle assembled the complete plastid genome of <italic>C. bakuense</italic> into the same circular contig under the original sequencing depth and all levels of sequencing coverage between and including 100x and 500x for both samples under study (<xref ref-type="table" rid="T2">Table 2</xref>). For sequencing coverage levels of 50x, 1,000x, and 2,000x, however, it generated two separate contigs that had to be concatenated to create a complete genome sequence. The breakpoint between these contigs was typically located at the junction site between IRb and the SSC, indicating that this non-contiguity was correlated with the quadripartite genome structure. NOVOPlasty seemed insensitive to changes in sequencing coverage across medium depth ranges, as it assembled the same circular complete plastid genome sequence under all levels between and including 500x and 2,000x for both samples (<xref ref-type="table" rid="T2">Table 2</xref>). For sequencing coverage above and below that range, however, NOVOPlasty was unable to assemble the same contig and instead generated either multiple smaller contigs, incomplete contigs, or contigs with unequal IR size. The single circular contig generated by GetOrganelle under sequence coverages of 100x&#x02013;500x and by NOVOPlasty under 500x or 2,000x was identical within each individual, and thus identical to the designated final plastid genomes of the <italic>C. bakuense</italic> individuals (i.e.,GetOrganelle under a sequencing coverage of 500x). Hence, at a sequencing coverage of 500x, both GetOrganelle and NOVOPlasty immediately and repeatably produced a complete plastid genome assembly for both individuals.</p>
<p>A strong variability in contig number, contig sequence, and contig length with regard to sequencing coverage was detected for assemblies generated with FastPlast and IOGA (<xref ref-type="table" rid="T1">Table 1</xref>). All genome assemblies generated by FastPlast under different levels of sequencing coverage exhibited different contig lengths. Moreover, the IRs of the assembled plastid genomes were found to be identical within assemblies only under the capped read sets as well as replicate run 1 of the uncapped read set in Cb04B. The assembly process of IOGA appeared to be even more sensitive to changes in sequencing coverage: for individual Cb01A, IOGA assembled 21 contigs under the original read set, 83 contigs under a coverage cap of 2,000x, and 51 contigs under a coverage cap of 500x; for Cb04B, the software generated between 54 and 85 contigs under the original read set (depending on the run replicate), 102 contigs under a coverage cap of 2,000x, and 40 contigs under a coverage cap of 500x. While at least half of the final genome sequence was encompassed within a single contig in all but one of these cases, the assembly results for each level of sequencing coverage had to be manually concatenated to generate complete plastid genomes. In addition to this high sensitivity to sequencing coverage, differences between replicate runs also indicated low reproducibility for sequence assemblies by both FastPlast and IOGA.</p>
<p>Computation time differed strongly across different assembly software and sequencing coverage and was generally correlated with the size of the input dataset: datasets with a capped sequencing coverage were typically analyzed faster than the original datasets (<xref ref-type="table" rid="T1">Tables 1</xref>, <xref ref-type="table" rid="T2">2</xref>). For a sequencing coverage of 500x, NOVOPlasty was the software that achieved a complete plastid genome assembly for <italic>C. bakuense</italic> in the shortest amount of time (7 min. and 6 min. for Cb01A and Cb04B, respectively); for lower levels of sequencing coverage, GetOrganelle was the software to achieve complete assemblies fastest.</p>
<p>In summary, we found that among the four assembly software tools tested, GetOrganelle and NOVOPlasty usually generated plastid genome assemblies that were identical in both length and sequence across run replicates and most levels of sequencing coverage. Occasional occurrences of more than two contigs generated per assembly run (e.g., GetOrganelle under a sequencing coverage of 2,000x) do not invalidate this observation, as the break point between such contigs was typically located at the junction between IRb and the SSC, which is a natural break point in a circular quadripartite genome. Overall, GetOrganelle slightly outperformed NOVOPlasty: it produced the full plastid genome in one contig already at lower sequencing coverage and had higher assembly accuracy, as some assembly results generated by NOVOPlasty contained sequence replications that extended the plastid genome sequence beyond its actual size (e.g., assembly of Cb01A under the original sequencing depth). We, therefore, considered the plastid genome sequences generated with GetOrganelle for the two individuals of <italic>C. bakuense</italic> as the best results and submitted them as official plastid genome sequences for the species to GenBank (accessions <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="MT806099">MT806099</ext-link> for Cb01A and <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="MT806098">MT806098</ext-link> for Cb04B; <xref ref-type="fig" rid="F2">Figure 2</xref>). Based on these sequences, the plastid genomes of Cb01A and Cb04B are almost identical and differ only by three nucleotides: a missing adenine in the intergenic spacer between the genes <italic>ndhF</italic> and <italic>rpl32</italic> in Cb01A, an additional thymine within a poly-T-microsatellite in the spacer between <italic>rps16</italic> and <italic>trnQ-UUG</italic> in Cb04B, and an additional thymine within a poly-T microsatellite in the spacer between <italic>pafI</italic> and <italic>trnS-GGA</italic> in Cb01A. Plastid genome diversity within <italic>C. bakuense</italic> is, thus, extremely low, but not zero.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Map of the complete plastid genome of individual Cb01A of <italic>C. bakuense</italic> as assembled by GetOrganelle under a coverage cap of 500x. This assembly represents the final plastid genome sequence for Cb01A.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0002.tif"/>
</fig>
</sec>
<sec>
<title>3.4. Characterization of Assembly Differences</title>
<p>PCoA of the number of SNPs and the length of each of the four plastid genome regions indicated the presence of a complex pattern of differences among plastid genome assemblies of different software tools and levels of sequencing coverage (<xref ref-type="fig" rid="F3">Figure 3A</xref>). Assemblies produced by different software tools were heterogeneous in both length and sequence for both individuals and differed by additional SNPs and the length of one or more plastid genome regions. In Cb04B, the first coordinate of the PCoA explained nearly the entire variance in the lengths of the four plastid genome regions, indicating the presence of one extreme or two nearly identical outlier assemblies. In Cb01A, the first coordinate of the PCoA similarly explained nearly the entire, comparatively low variance for the SSC length, but not for the lengths of the LSC and the IRs, where more diversity among a greater number of outliers was identified. For the number of SNPs, the first two PCoA coordinates together explained &#x0003E;60% of the variance in both individuals, although overall variance for Cb01A was greater than for Cb04B according to the absolute variance values.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Comparisons of the number of SNPs and the lengths of the four genome regions across the plastid genome assemblies of <italic>C. bakuense</italic> as generated by different assembly software and levels of sequencing coverage. Subplot <bold>(A)</bold> displays the results of PCoAs, subplots <bold>(B,C)</bold> the results of comparisons between a target assembly and the final plastid genome sequence, and subplot <bold>(D)</bold> the results of assembly comparisons between the two individuals of <italic>C. bakuense</italic> under study. In the PCoA plots, the percentages indicate the variance explained by the first (x-axis) and second (y-axis) principal coordinate, and the integers express the range of the data. The abbreviations for the four distance metrics are: &#x0201C;SNPcount&#x0201D; for the total number of SNPs between two assemblies; &#x0201C;LSClendif,&#x0201D; &#x0201C;SSClendif,&#x0201D; and &#x0201C;IRlendif&#x0201D; for the differences in sequence length in the LSC, SSC, and IR between two assemblies, respectively.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0003.tif"/>
</fig>
<p>The comparison of pairwise genetic distances between the assemblies of different software tools highlighted the presence of SNPs between the final genome sequences and the assemblies generated with FastPlast and IOGA (<xref ref-type="fig" rid="F3">Figure 3B</xref>). This contrasts with the presence of IR and LSC length differences between the final genome sequences and the assemblies generated with NOVOPlasty (especially in Cb01A) and FastPlast (especially in Cb04B). The overall similarity of the length difference patterns for the LSC and the IR suggests that length deviations in either region are often compensated by a corresponding change in the other region during genome assembly, rather than changes of the SSC.</p>
<p>The comparison of pairwise genetic distances between the assemblies of different levels of sequencing coverage highlighted that the observed length and sequence deviations from the final genome sequences were not constant across different levels (<xref ref-type="fig" rid="F3">Figure 3C</xref>); only the assemblies generated with GetOrganelle were found to be unaffected by alterations in sequencing coverage. For assemblies generated with IOGA, for example, the reduction of sequencing coverage had a complex but strong effect on SNP count and region length, as it correlated with a decrease of the number of SNPs and the IR/LSC length difference in Cb01A but an increase of both factors in Cb04B. A similar pattern was found for assemblies generated with FastPlast and, for Cb01A, also for NOVOPlasty. GetOrganelle was the only assembly software found to produce assemblies with the same sequence and region lengths across all evaluated assembly parameters.</p>
<p>The comparison of genetic distances between the assemblies of the two individuals of <italic>C. bakuense</italic> demonstrated that only GetOrganelle consistently and repeatedly generated the final plastid genome sequence for each individual under study (<xref ref-type="fig" rid="F3">Figure 3D</xref>). We did not find any SNPs between the assemblies produced by GetOrganelle for the two individuals except for two nucleotide differences in the LSC (which were neutral regarding the overall length difference due to their occurrence in different individuals) and one in the SSC. Under FastPlast and IOGA, by contrast, the number of SNPs detected between the two assemblies was much greater and even exceeded the threshold of 1,000 nucleotide differences in the case of IOGA. Moreover, under both FastPlast and IOGA the number of SNPs between different assemblies of the same individual did not sum up to the number of SNPs between individuals, suggesting that at least some of the SNPs were shared between the assemblies of the same individual. The differences in LSC and IR length for assemblies generated with NOVOPlasty appeared to be correlated, suggesting that a length deviation in one region was compensated for by a corresponding change in the other region rather than a change in SSC length. Furthermore, visual examination of the assemblies indicated that several assemblies generated with IOGA under higher levels of sequencing coverage deviated from the other assemblies by insertions ranging from 170 and 334 bp; these insertions often had little, if any, similarity to other regions of the plastid genome.</p>
<p>The visual comparison of the lengths of the four plastid genome regions across different genome assemblies indicated that the differences in total genome length were primarily correlated with length changes in the LSC and the IRs (<xref ref-type="fig" rid="F4">Figure 4</xref>). While the length of the SSC was virtually constant across all software tools and sequencing coverage (&#x0007E;13,400 bp; <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S2</xref>), the length of the IR was highly sensitive to the precise assembly conditions. Especially in assemblies generated with NOVOPlasty for Cb01A as well as with FastPlast for Cb04B, the IR lengths varied by a factor of 1.5 to 2, which was partially compensated for by a corresponding reduction of the LSC length, sometimes to less than half of the length displayed in other assemblies. A complete list of the lengths of the four plastid genome regions in relation to the different software tools, levels of sequencing coverage, seed sequences, and run replicates is given in <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S2</xref> for Cb01A and <xref ref-type="supplementary-material" rid="SM1">Supplementary Table S3</xref> for Cb04B (<xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref>).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Overview of the relative lengths of the LSC, the SSC, and the two IRs across the plastid genome assemblies of the individuals Cb01A <bold>(A)</bold> and Cb04B <bold>(B)</bold> of <italic>C. bakuense</italic> as generated by different assembly software and levels of sequencing coverage.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0004.tif"/>
</fig>
</sec>
<sec>
<title>3.5. Differences in Gene Content and Annotations</title>
<p>The nucleotide and length differences between the assembled plastid genomes were located in both the coding and the non-coding sections of the genomes and often manifested themselves as differences in gene content (<xref ref-type="table" rid="T3">Table 3</xref>). Specifically, the annotated sequences of several assemblies either lacked certain protein- and tRNA-coding genes due to missing genome sections or exhibited non-functional protein-coding genes due to internal stop codons caused by nucleotide polymorphisms. All assemblies generated with IOGA, for example, exhibited housekeeping genes with internal stop codons, which are indicative of an incorrect assembly. Among the assemblies generated with FastPlast, replicate runs 1 and 2 for Cb01A and replicate run 2 for Cb04B under the original sequencing depth as well as the assembly of Cb04B under a coverage cap of 500x produced gene sequences with internal stop codons. Similarly, the length differences between the four plastid genome regions across the assemblies generated with NOVOPlasty for Cb01A correlated with a lack of up to 17 different genes compared to the final genome sequence of that plant individual, even when all assembly contigs were concatenated to a super-contig; this result was observed for both seed sequences and, thus, appears to be independent of the internal start point of the genome assembly. All of the missing genome regions in the assemblies generated with NOVOPlasty were noticeably located at the 5&#x00027; end of the LSC, suggesting a potential bias in the assembly of this genome region. All plastid genome assemblies generated with GetOrganelle, by contrast, exhibited a complete gene complement and the full genome size.</p>
<table-wrap position="float" id="T3">
<label>Table 3</label>
<caption><p>Overview of incorrect or missing annotations among the plastid genome assemblies of <italic>C. bakuense</italic> as generated under different assembly software, sequencing coverage, seed sequences, and run replicates.</p></caption>
<table frame="hsides" rules="groups">
<thead><tr>
<th valign="top" align="left"><bold>Asmb</bold>.</th>
<th valign="top" align="left"><bold>Cov</bold>.</th>
<th valign="top" align="left"><bold>Repl</bold>.</th>
<th valign="top" align="left"><bold>NOVO seed</bold></th>
<th valign="top" align="left"><bold>Internal stop codons in translation</bold></th>
<th valign="top" align="left"><bold>No DNA sequence at this position</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><bold>Cb01A</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td/>
<td valign="top" align="left">rpl23<sup>a,b</sup>, rrn16<sup>a,b</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td/>
<td valign="top" align="left">rpl2<sup>a,b</sup>, ycf2<sup>a,b</sup>, rpl23<sup>a</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">500x</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td/>
<td valign="top" align="left">psbA, rpl23<sup>a,b</sup>, ycf2<sup>a,b</sup>, rrn16<sup>a,b</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td/>
<td valign="top" align="left">psbA, ycf2<sup>a,b</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td/>
<td valign="top" align="left">psbA, ycf2<sup>a,b</sup>, ycf1<sup>a,b</sup>, ndhH</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">500x</td>
<td/>
<td/>
<td valign="top" align="left">psbA, rps2<sup>a,b</sup>, ycf2<sup>a,b</sup>, ndhH</td>
<td valign="top" align="left">trnH-GUG</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td valign="top" align="left">seed1</td>
<td/>
<td valign="top" align="left">trnH-GUG, psbA, trnK-UUU, matK, rps16</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td valign="top" align="left">seed1</td>
<td/>
<td valign="top" align="left">trnH-GUG, psbA, trnK-UUU, matK, rps16, trnQ-UUG, psbK, psbI, trnS-GCU, trnG-UCC, trnR-UCU, atpA, atpF, atpH, atpI, rps2, rpoC2</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td valign="top" align="left">seed1</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">500x</td>
<td/>
<td valign="top" align="left">seed1</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td valign="top" align="left">seed2</td>
<td/>
<td valign="top" align="left">trnH-GUG, psbA, trnK-UUU, matK, rps16, trnQ-UUG, psbK, psbI, trnS-GCU, trnG-UCC, trnR-UCU, atpA, atpF, atpH, atpI, rps2, rpoC2</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td valign="top" align="left">seed2</td>
<td/>
<td valign="top" align="left">trnH-GUG, psbA, trnK-UUU, matK, rps16, trnQ-UUG, psbK, psbI, trnS-GCU, trnG-UCC, trnR-UCU, atpA, atpF, atpH, atpI, rps2, rpoC2</td>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td valign="top" align="left">seed2</td>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">NOVO</td>
<td valign="top" align="left">500x</td>
<td/>
<td valign="top" align="left">seed2</td>
<td/>
<td/>
</tr> <tr style="border-top: thin solid #000000;">
<td valign="top" align="left"><bold>Cb04B</bold></td>
<td/>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td/>
<td valign="top" align="left">rpl23<sup>a</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td/>
<td/>
<td/>
</tr>
<tr>
<td valign="top" align="left">FaPl</td>
<td valign="top" align="left">500x</td>
<td/>
<td/>
<td valign="top" align="left">ndhF</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl1</td>
<td/>
<td valign="top" align="left">psbA, rpl23<sup>b</sup>, rpl2<sup>a</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">orig.</td>
<td valign="top" align="left">repl2</td>
<td/>
<td valign="top" align="left">psbA, ndhH, rpl23<sup>a</sup>, rpl2<sup>a,b</sup></td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">2,000x</td>
<td/>
<td/>
<td valign="top" align="left">psbA, petB</td>
<td/>
</tr>
<tr>
<td valign="top" align="left">IOGA</td>
<td valign="top" align="left">500x</td>
<td/>
<td/>
<td valign="top" align="left">psbA, rps23<sup>a,b</sup></td>
<td/>
</tr> </tbody>
</table>
<table-wrap-foot>
<p><italic>All plastid genome assemblies generated with GetOrganelle for both individuals and with NOVOPlasty for Cb04B exhibited a complete gene complement and a full genome size and are, thus, not listed. The last column denotes cases of incomplete genomes despite the assembly being circular and indicated as complete by the assembly software. A location in IRa is indicated<sup>a</sup>, a location in IRb<sup>b</sup>. Abbreviations used as in <xref ref-type="table" rid="T1">Table 1</xref></italic>.</p>
</table-wrap-foot>
</table-wrap>
</sec>
<sec>
<title>3.6. Sequencing Coverage and SNP Location</title>
<p>Visualizing the location of SNPs across the plastid genome assemblies indicated a possible association of their location with regions of low sequencing coverage. The genome-wide sequencing depth based on the uncapped datasets was 8,410x for the final plastid genome of Cb01A and 5,430x for that of Cb04B. Among the assemblies generated with different assembly software, a considerable number exhibited SNPs when compared to the final genome sequence. Notably, these SNPs were often associated with regions of reduced sequencing coverage. For example, the IRs of the plastid genome assemblies of Cb01A generated with FastPlast contained two adjacent calculation windows with a sequencing coverage of 1,200x and 2,300x, respectively; these depths represent only 14% and 27% of the genome-wide sequencing depth (<xref ref-type="fig" rid="F5">Figure 5</xref>). The two windows were located between the tRNA genes <italic>trnV-GAC</italic> and <italic>trnI-GAU</italic> and covered parts of the gene coding for the 16S rRNA subunit (<italic>rrn16</italic>). Compared to the final genome sequence of Cb01A, the assemblies of both replicate runs exhibited a high density of SNPs in the very same region (<xref ref-type="fig" rid="F5">Figure 5</xref>, circles A and B); SNPs outside this particular region also existed but were clustered less densely, if at all. Similarly, a high density of SNPs was found in replicate run 2 at the replication origin of the genome, which also exhibits a considerably reduced sequencing coverage (<xref ref-type="fig" rid="F5">Figure 5</xref>, circle B); however, the reduced sequencing coverage at the replication origin represents an artifact introduced by the mapping software during the extraction of plastid genome reads from the raw read set and should, thus, not be seen as a region with naturally reduced sequencing coverage. The assemblies generated with FastPlast under the capped read sets, by contrast, did not exhibit SNPs compared to the final genome sequence (<xref ref-type="fig" rid="F5">Figure 5</xref>, circles C and D). A similar interdependence between the location of SNPs and regions with reduced sequencing coverage was observed for the assemblies of Cb01A generated with IOGA (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S1</xref>); the amount and the distribution of SNPs in comparison to the final genome sequence were, however, greater than in the assemblies with FastPlast and neither restricted to the IRs nor any particular read set. By comparison, the plastid genome assemblies generated with GetOrganelle or NOVOPlasty did not display any SNPs in comparison to the final genome sequence, irrespective of a cap on sequencing coverage.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Visualization of the sequencing coverage across the plastid genome of individual Cb01A as generated with FastPlast and the location of SNPs of assemblies generated under different levels of sequencing coverage. Red bars in the visualization of sequencing coverage indicate calculation windows with a depth equal to, or less than, 50% of genome-wide sequencing depth. The four rings beneath the coverage visualization indicate the location of SNPs relative to the final genome sequence for the following assemblies: replicate run 1 (A) and 2 (B) under the original sequencing depth; a coverage cap of 2,000x (C); a coverage cap of 500x (D). Black bars within each ring represent the occurrence of three SNPs per 100 bp.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0005.tif"/>
</fig>
</sec>
<sec>
<title>3.7. Phylogenetic Inference</title>
<p>The results of our phylogenetic tree reconstructions on the combined set of all plastid genome assemblies of <italic>C. bakuense</italic> plus the 21 plastid genome records of other species of <italic>Calligonum</italic> and the outgroup did not indicate that the sequence variability across the assemblies generated in this study was large enough to affect the phylogenetic placement of <italic>C. bakuense</italic> within <italic>Calligonum</italic> (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figures S2</xref>, <xref ref-type="supplementary-material" rid="SM1">S3</xref>). While the different genome assemblies of <italic>C. bakuense</italic> did not cluster by assembly software or level of sequencing coverage, they did exhibit a noticeable clustering by plant individual. Specifically, a strong clustering by plant individual was observed when sequence insertions and deletions (indels) of the underlying matrix were coded and included in the phylogenetic reconstruction (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S3</xref>), whereas no such clustering was observed without the coding of indels (<xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S2</xref>). Moreover, we found that the nucleotide differences between the majority of our assemblies were not or only minimally phylogenetically informative and, thus, did not result in the identification of specific clades among the assembly sequences. The observed sequence differences among the assemblies may nonetheless be large enough to influence intra-specific evolutionary analyses of <italic>C. bakuense</italic>.</p>
<p>The results of our phylogenetic tree reconstruction to infer the phylogenetic position of <italic>C. bakuense</italic> among other species of <italic>Calligonum</italic> recovered the final plastid genomes of <italic>C. bakuense</italic> as sister to <italic>C. caput-medusae</italic> (<xref ref-type="fig" rid="F6">Figure 6</xref> and <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S2</xref>). The sister relationship between <italic>C. bakuense</italic> and <italic>C. caput-medusae</italic> was weakly supported (BS 66%) but both taxa were recovered as part of a fully-supported clade alongside <italic>C. arborescens</italic>. Overall, the reconstruction recovered the same phylogenetic relationships as reported by Song et al. (<xref ref-type="bibr" rid="B65">2020</xref>), indicating that the inclusion of <italic>C. bakuense</italic> did not alter the tree reconstruction of the genus.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Phylogenetic position of <italic>C. bakuense</italic> among other species of <italic>Calligonum</italic>. <italic>C. bakuense</italic> is represented by the final plastid genomes of individuals Cb01A and Cb04B, which are highlighted in bold. The displayed phylogenetic tree represents the best tree inferred under ML, visualized as <bold>(A)</bold> cladogram with bootstrap node support (given above branches) and <bold>(B)</bold> the corresponding phylogram with exact branch lengths.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-13-779830-g0006.tif"/>
</fig>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4. Discussion</title>
<sec>
<title>4.1. Phylogenetic Position of <italic>C. bakuense</italic> Based on Complete Plastid Genomes</title>
<p>This investigation is the first to report the complete plastid genome of <italic>C. bakuense</italic> and, thus, advances our understanding of <italic>Calligonum</italic>, as knowledge of the plastid genome of this Caucasian endemic supports research on the evolutionary diversification of the genus. For example, our analyses underscore the potential of complete plastid genome sequences for resolving species-level relationships in <italic>Calligonum</italic> (e.g., Song et al., <xref ref-type="bibr" rid="B65">2020</xref>), whereas individual regions of the plastid genome appear to yield insufficient phylogenetic information (e.g., Tavakkoli et al., <xref ref-type="bibr" rid="B72">2010</xref>). The results of our phylogenetic analyses (<xref ref-type="fig" rid="F6">Figure 6</xref>) only partially agree with the current taxonomic classification of <italic>Calligonum</italic>. <italic>Calligonum bakuense</italic> is considered a member of sect. <italic>Calligonum</italic>, yet was recovered as part of a clade formed by individuals of <italic>C. arborescens</italic> and <italic>C. caput-medusae</italic>, both of which are members of sect. <italic>Medusa</italic> SOSK. ET ALEXANDR (Soskov, <xref ref-type="bibr" rid="B68">2011</xref>). The current sectional classification of <italic>Calligonum</italic> is primarily based on differences in fruit morphology and probably not natural, as suggested by Song et al. (<xref ref-type="bibr" rid="B65">2020</xref>); our results provide further evidence for this interpretation. The phylogenetic position of <italic>C. bakuense</italic> in a clade with <italic>C. arborescens</italic> and <italic>C. caput-medusae</italic> may indicate that <italic>C. bakuense</italic> represents an isolated lineage endemic to Azerbaijan in a Caucasian-central Asian clade. <italic>Calligonum bakuense</italic> occurs on the west coast of the Caspian Sea, whereas <italic>C. arborescens</italic> and <italic>C. caput-medusae</italic> both grow in steppe habitats east of the Caspian Sea, ranging from Turkmenistan to China.</p>
<p>Due to similarities in fruit morphology and its tetraploid nature (2n = 36; Bolkhovskikh et al., <xref ref-type="bibr" rid="B8">1969</xref>), <italic>C. bakuense</italic> was hypothesized to be an allotetraploid that arose from ancestors of <italic>C. polygonoides</italic> L. and <italic>C. acanthopterum</italic> I.G. BORSHCH (Soskov and Akhmed-Zade, <xref ref-type="bibr" rid="B67">1974</xref>). While <italic>C. polygonoides</italic> is widespread and also occurs in Azerbaijan (Karjagin, <xref ref-type="bibr" rid="B35">1952</xref>), <italic>C. acanthopterum</italic> is known only from Kazakhstan and Turkmenistan. By contrast, the widespread species <italic>C. aphyllum</italic>, which is distributed from North Africa to the Caucasus (including Azerbaijan) and China, is morphologically distinct from <italic>C. bakuense</italic> (e.g., winged fruits that lack bristles) and probably not a close relative to <italic>C. bakuense</italic>. Future phylogenetic investigations should, thus, increase both the taxon sampling and, where possible, the geographic representation of the more widespread taxa of <italic>Calligonum</italic> such as <italic>C. polygonoides</italic>. Since the relationships among <italic>C. bakuense, C. arborescens</italic>, and <italic>C. caput-medusae</italic> were unsupported when only the coding sections of the plastid genome were used for phylogenetic reconstruction (trees not shown), our results corroborate the observation that the inclusion of the non-coding sections of the plastid genome (i.e., introns and intergenic spacers) in a genus-wide plastid phylogenomic analysis represents an important aspect in clarifying the phylogenetic history of angiosperm genera with low genetic distances among species (e.g., <italic>Gynoxys</italic>; Escobari et al., <xref ref-type="bibr" rid="B21">2021</xref>). The inclusion of phylogenetic information from the nuclear genome in future investigations will likely assist in clarifying possible reticulate speciation events within <italic>Calligonum</italic>.</p>
<p>The three nucleotide differences detected between the plastid genomes of the two individuals of <italic>C. bakuense</italic> are comparatively few but could be in the same range as those of other narrow endemic plant species. While intra-specific comparisons of complete plastid genomes are still rare (e.g., Jiang et al., <xref ref-type="bibr" rid="B32">2017</xref>; Teshome et al., <xref ref-type="bibr" rid="B73">2020</xref>), published studies of endemics often report only a handful of SNPs between plant individuals. The narrow endemic <italic>Pinus torreyana</italic>, for example, had five SNPs between the plastid genomes of two individuals from both parts of its disjunct distribution range (Whittall et al., <xref ref-type="bibr" rid="B78">2010</xref>, indels not reported). Similarly, at least two SNPs and one indel were found between two plastid genomes of <italic>Fagus multinervis</italic>, which is endemic to Ulleung Island near the South Korean coast (Yang et al., <xref ref-type="bibr" rid="B82">2020</xref>). While plastid genome sequences of <italic>C. bakuense</italic> provide valuable background on the evolutionary history of this species, further analysis of its nuclear genomic diversity remains necessary for a sound conservation genetic assessment.</p>
</sec>
<sec>
<title>4.2. Impact of Software Choice on Plastid Genome Assembly</title>
<p>By comparing the assembly contigs of <italic>C. bakuense</italic> that were generated with four different software tools, we found that assembly software choice can have an inordinate influence on the inferred plastid genome sequences and that the results of some tools need to be treated with caution. Among the differences across the assemblies were the presence of SNPs and indels (compared to the final genome sequences), the incorrect absence of entire genes or loss of their functionality, and the expansion and contraction of the IRs (as well as compensatory length changes in the LSC). Such occurrences have been occasionally interpreted in an evolutionary context (e.g., Mohanta et al., <xref ref-type="bibr" rid="B49">2020</xref>), but it stands to reason that at least some of the differences between the plastid genomes of closely related species may have a more technical origin, as recently demonstrated by Freudenthal et al. (<xref ref-type="bibr" rid="B22">2020</xref>). The results of this investigation support the hypothesis that differences among plastid genome assemblies may also be technical in nature. We found that the plastid genome assemblies of different software tools exhibited considerably different genome sequences despite employing the same input data and that some of the assembly contigs could not be replicated in different runs of the same software (<xref ref-type="table" rid="T1">Table 1</xref>). Only the software GetOrganelle was found to generate consistent and repeatable results for both datasets. The software FastPlast, by contrast, was found to be prone to the introduction of SNPs and, in some cases, also structural deviation among the assemblies. NOVOPlasty introduced few, if any, SNPs compared to the final genome sequences but exhibited a tendency for generating structural deviations, which even occurred when the same assembly was conducted with different seed sequences. The assemblies generated with IOGA were fragmentary in all cases and exhibited numerous SNPs and structural deviations compared to the final genome sequences. Worse still, we found that many of these sequence deviations generated by IOGA would result in incorrect conclusions about gene content and functionality when compared to the final genome sequences (<xref ref-type="table" rid="T3">Table 3</xref>). We, therefore, concur with Freudenthal et al. (<xref ref-type="bibr" rid="B22">2020</xref>) that users should abstain from employing the software IOGA (which is no longer maintained) for plastid genome assembly and that the assemblers FastPlast and NOVOPlasty should be employed with caution. We also concur with the suggestion that the replication of assembly results across different software runs and seed sequences (where applicable) are beneficial precautions in the generation of trustworthy plastid genome sequences.</p>
<p>Our results do not imply that the assemblies generated with GetOrganelle necessarily represent true plastid genome sequences for <italic>C. bakuense</italic>. It is possible for a software tool to consistently and repeatably produce incorrect results, and we also cannot rule out the presence of more than one unique plastid genome per plant individual (Scarcelli et al., <xref ref-type="bibr" rid="B61">2016</xref>; Wang and Lanfear, <xref ref-type="bibr" rid="B77">2019</xref>). However, the software tools FastPlast and NOVOPlasty produced the same genome sequence as identified through GetOrganelle under some of the evaluated settings. We, therefore, considered the plastid genome assemblies generated with GetOrganelle under the read sets capped at a sequencing coverage of 500x as the most likely genome sequences for the two individuals of <italic>C. bakuense</italic> and employed them as the final plastid genomes. Aside from the idiosyncrasies introduced by different assembly software, the observed differences among the plastid genome assemblies may also be the result of nucleotide polymorphism among the input reads (Scarcelli et al., <xref ref-type="bibr" rid="B61">2016</xref>). Such polymorphism within the read set could represent genuinely different variants of the plastid genome (i.e., heteroplasmy; Walker et al., <xref ref-type="bibr" rid="B76">2015</xref>; Wang and Lanfear, <xref ref-type="bibr" rid="B77">2019</xref>), genomic transfers of sections of the plastid to the nuclear or the mitochondrial genome, followed by a pseudogenization of the transferred regions (Ruhlman and Jansen, <xref ref-type="bibr" rid="B58">2014</xref>), or sequencing errors during data generation (Nakamura et al., <xref ref-type="bibr" rid="B52">2011</xref>), and may be decoded differently by different assembly software.</p>
</sec>
<sec>
<title>4.3. Impact of Sequencing Coverage on Plastid Genome Assembly</title>
<p>By comparing the assembly contigs of <italic>C. bakuense</italic> generated under different levels of sequencing coverage, we found that sequencing coverage can also have an impact on plastid genome assembly. Specifically, we found that the capping of sequencing coverage prior to genome assembly had a measurable effect on the number of assembly contigs constructed, the nucleotide sequences of these contigs, the length of the different plastid genome regions (particularly the IRs), and the number of valid gene annotations. The effects of capping sequencing coverage were measurable in both samples and suggested the trend that a sequencing depth between 100x and 500x rendered the assemblies relatively consistent in sequence and length (<xref ref-type="table" rid="T2">Table 2</xref>). Specifically, a sequencing depth between 100x and 500x appeared to ensure replicability of the genome assemblies with GetOrganelle and NOVOPlasty, whereas levels of sequencing coverage above and below that range did not enable a complete plastid genome assembly. A similar albeit slightly lower range of optimal sequencing depth for the assembly of plastid genomes has been reported for PacBio sequencing data (i.e., 50&#x02013;200x; Soorni et al., <xref ref-type="bibr" rid="B66">2017</xref>) and is in line with observations on the absolute minimum sequencing coverage for the reliable plastid genome assembly (i.e., 30&#x02013;50x; Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>; Sharpe et al., <xref ref-type="bibr" rid="B62">2020</xref>). In practice, an amount of approximately two to 10 million Illumina read pairs of 150 bp length per read, generated from DNA fragments with an average length of 300 bp, can cover a plastid genome of approximately 160,000 bp with a sequencing coverage of 100x to 500x. This assumes that an average of 2.5% of all reads represent the plastid genome, which is a common value in genome skimming experiments (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>; McKain et al., <xref ref-type="bibr" rid="B47">2018</xref>). Although we cannot exclude that the optimal plastid genome coverage, and with it the raw sequence data needed, differs across species and datasets, we found the same result for data from two different individuals and two different assembly pipelines, indicating a potential pattern.</p>
<p>The results of this investigation indicate that the evenness of sequencing coverage may be an important but as of yet insufficiently recognized factor in the successful assembly of plastid genomes. Both the original and several of the capped read sets analyzed here vastly exceed the recommended level of sequencing coverage for plastid genome assembly (Twyford and Ness, <xref ref-type="bibr" rid="B74">2017</xref>; McKain et al., <xref ref-type="bibr" rid="B47">2018</xref>). When only considering the plastid genome reads of this uncapped read set, a sequencing depth of 8,410x and a minimum sequencing coverage of more than 1,000x in any genome position exists, indicating that the original read set of Cb01A comprises more than enough sequence information to completely assemble the plastid genome. The failure of some of the tested software tools to assemble the plastid genome is, thus, more likely associated with the unevenness than the depth of sequencing coverage. A medium but comparatively even level of sequencing coverage may be the best strategy for a successful plastid genome assembly with the tested software tools.</p>
<p>Our results are congruent with the findings of other investigations that report an impact of sequencing coverage on the genome assembly process (Stadermann et al., <xref ref-type="bibr" rid="B70">2015</xref>; Pedersen et al., <xref ref-type="bibr" rid="B54">2017</xref>) or a correlation between local extremes in sequencing coverage and assembly contig deviations (Kim et al., <xref ref-type="bibr" rid="B38">2015</xref>). In general, the level of sequencing coverage is indicative for a reliable identification of sequence rearrangements and other structural variants (Sims et al., <xref ref-type="bibr" rid="B64">2014</xref>; Izan et al., <xref ref-type="bibr" rid="B31">2017</xref>), but the relationship between sequencing coverage and assembly reliability is not straightforward. While greater sequencing coverage typically increases the chance that rearrangement endpoints are captured and confirmed by multiple reads (Chen et al., <xref ref-type="bibr" rid="B15">2009</xref>), genomic regions with exceptionally high depth of sequencing coverage have also been reported as problematic for the identification of SNPs (Li, <xref ref-type="bibr" rid="B40">2014</xref>).</p>
</sec>
<sec>
<title>4.4. Impact of Assembly Differences on Phylogenetic Placement</title>
<p>The results of this investigation illustrate that a correct plastid genome assembly cannot be taken for granted without a subsequent evaluation of the assembly, even when employing dedicated software tools. Incorrect genome assemblies have the potential to affect downstream biological interpretations, such as analyses of evolutionary relationships or genetic diversity. Even if the assembly differences observed in this study only marginally affected the inferred phylogenetic position of <italic>C. bakuense</italic> within <italic>Calligonum</italic> (<xref ref-type="fig" rid="F6">Figure 6</xref> and <xref ref-type="supplementary-material" rid="SM1">Supplementary Figure S3</xref>), we cannot exclude the possibility that errors introduced during the assembly can lead to incorrect phylogenetic reconstructions. Plant lineages with low genetic distances between species are likely particularly sensitive to this problem (e.g., Escobari et al., <xref ref-type="bibr" rid="B21">2021</xref>).</p>
</sec>
<sec>
<title>4.5. Recommendations for Future Studies</title>
<p>Given the results of this investigation, we propose three recommendations for the application of <italic>de novo</italic> plastid genome assembly. First, we recommend comparing the assembly results of different software tools and multiple software runs before accepting any assembly as the final genome sequence. As demonstrated here, results from different assembly software tools may vary considerably in their accuracy and repeatability. We, therefore, recommend considering only such results for subsequent analyses that are reproducible across different tools and replicate runs. This is not restricted to the four software tools tested in this investigation; there are various software applications for <italic>de novo</italic> genome assembly from genome skimming data, including tools specialized in circular genomes (such as plastid genomes) and general short-read assemblers. We tested three such general assemblers on the complete, unfiltered sequence dataset of <italic>C. bakuense</italic> in a preliminary investigation: SOAPdenovo2 (Luo et al., <xref ref-type="bibr" rid="B43">2012</xref>), Platanus (Kajitani et al., <xref ref-type="bibr" rid="B34">2014</xref>), and Meraculous (Chapman et al., <xref ref-type="bibr" rid="B14">2011</xref>) and found that only Platanus generated assembly contigs that collectively represented either the complete plastid genome of <italic>C. bakuense</italic> (Cb01A) or sections of it (Cb04B). This strongly suggests that even in sequencing projects primarily targeting the nuclear genome, a separate assembly of the plastid genome with dedicated software may be required to produce reliable results. Second, we recommend capping the sequencing coverage of the input read data to an approximately even distribution along the whole genome sequence while keeping the sequencing depth within a range of 500x to 100x when conducting plastid genome assembly. While the exact relationship between sequence accuracy and both sequencing coverage and evenness is poorly understood for the assembly of plastid genomes, the results of similar investigations on bacterial genomes indicate a considerable impact of both factors (Magoc et al., <xref ref-type="bibr" rid="B44">2013</xref>; Pedersen et al., <xref ref-type="bibr" rid="B54">2017</xref>). More research is needed to determine the optimal balance between the depth and the evenness of sequencing coverage for reliable plastid genome assembly. Third, we recommend the release of detailed assembly and annotation information during the publication of new plastid genomes. Only by sharing a precise description of the type and succession of the software tools employed are assembly results genuinely reproducible and, ultimately, reliable (Gruening et al., <xref ref-type="bibr" rid="B23">2018</xref>; Gruenstaeudl et al., <xref ref-type="bibr" rid="B25">2018</xref>). The provisioning of detailed assembly and annotation information is also essential if researchers wish to re-analyze the data with new and improved methods (e.g., Gruenstaeudl, <xref ref-type="bibr" rid="B24">2019</xref>). Expressly for this purpose, we release the raw sequence reads, the read datasets capped at different levels of sequencing coverage, and the raw assembly results as <xref ref-type="supplementary-material" rid="SM1">Supplementary Material</xref> to this investigation.</p>
</sec>
</sec>
<sec sec-type="data-availability" id="s5">
<title>Data Availability Statement</title>
<p>The datasets presented in this study can be found in online repositories. The names of the repository/repositories and accession number(s) can be found below: <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/record/6577786">https://zenodo.org/record/6577786</ext-link>, Zenodo record 6577786; <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/sra">https://www.ncbi.nlm.nih.gov/sra</ext-link>, NCBI SRA records <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="SRX9433946">SRX9433946</ext-link> and <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="SRX9433941">SRX9433941</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>The study was devised by MG and KR, with participation from TB and VK. The distribution data was assessed by VK and visualized by KR. Preliminary analyses were conducted by CC and MG, and final analyzes by EG, and MG. EG performed all post-assembly finishing steps. MG performed the sequence comparisons and the phylogenetic reconstructions. KR calculated the PCoAs. The writing of the manuscript was led by MG and KR, with additional input from EG, and TB. The revision of the manuscript was organized by MG, with additional input from KR, and TB. All authors have read and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="funding-information" id="s7">
<title>Funding</title>
<p>This study was partially funded by the Volkswagen Foundation, Grant No. AZ 89 950 Developing tools for conserving the plant diversity of the South Caucasus.</p>
</sec>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x00027;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec> </body>
<back>
<ack><p>We thank Gerald Parolly, Nadja Korotkova, and Tural Qasimov for assistance with field collection, Bettina Giesicke for DNA extraction, Halil Atis for Illumina library preparation, and Cathrin Schierenbeck for assistance with sequence data archiving. The authors acknowledge the Berlin Center of Genomics in Biodiversity Research for providing lab assistance and the high-performance computing service of the ZEDAT of the Freie Universit&#x000E4;t Berlin for providing allocations of computing time. Several of the analyses presented here represent part of a thesis by EG toward a master of science degree.</p>
</ack>
<sec sec-type="supplementary-material" id="s9">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2022.779830/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpls.2022.779830/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abdellaoui</surname> <given-names>R.</given-names></name> <name><surname>Gouja</surname> <given-names>H.</given-names></name> <name><surname>Sayah</surname> <given-names>A.</given-names></name> <name><surname>Neffati</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>An efficient DNA extraction method for desert <italic>Calligonum</italic> species</article-title>. <source>Biochem. Genet</source>. <volume>49</volume>, <fpage>695</fpage>&#x02013;<lpage>703</lpage>. <pub-id pub-id-type="doi">10.1007/s10528-011-9443-7</pub-id><pub-id pub-id-type="pmid">21681578</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ankenbrand</surname> <given-names>M.</given-names></name> <name><surname>Pfaff</surname> <given-names>S.</given-names></name> <name><surname>Terhoeven</surname> <given-names>N.</given-names></name> <name><surname>Qureischi</surname> <given-names>M.</given-names></name> <name><surname>G&#x000FC;ndel</surname> <given-names>M.</given-names></name> <name><surname>Weiss</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>chloroExtractor: extraction and assembly of the chloroplast genome from whole genome shotgun data</article-title>. <source>J. Open Source Softw</source>. <volume>3</volume>, 464. <pub-id pub-id-type="doi">10.21105/joss.00464</pub-id></citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Atamov</surname> <given-names>V.</given-names></name></person-group> (<year>2008</year>). <article-title>Phytosociological characteristics the vegetation of the Caspians shores in Azerbaijan</article-title>. <source>Int. J. Bot</source>. <volume>4</volume>, <fpage>1</fpage>&#x02013;<lpage>13</lpage>. <pub-id pub-id-type="doi">10.3923/ijb.2008.1.13</pub-id></citation>
</ref>
<ref id="B4">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Baillie</surname> <given-names>J.</given-names></name> <name><surname>Hilton-Taylor</surname> <given-names>C.</given-names></name> <name><surname>Stuart</surname> <given-names>S.</given-names></name></person-group> (<year>2004</year>). <source>2004 IUCN Red List of Threatened Species: A Global Species Assessment</source>. <publisher-loc>Gland</publisher-loc>: <publisher-name>IUCN Conservation Centre</publisher-name>.</citation>
</ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bakker</surname> <given-names>F.</given-names></name></person-group> (<year>2017</year>). <article-title>Herbarium genomics: skimming and plastomics from archival specimens</article-title>. <source>Webbia</source> <volume>72</volume>, <fpage>35</fpage>&#x02013;<lpage>45</lpage>. <pub-id pub-id-type="doi">10.1080/00837792.2017.1313383</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bakker</surname> <given-names>F.</given-names></name> <name><surname>Lei</surname> <given-names>D.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name> <name><surname>Mohammadin</surname> <given-names>S.</given-names></name> <name><surname>Wei</surname> <given-names>Z.</given-names></name> <name><surname>van de Kerke</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Herbarium genomics: plastome sequence assembly from a range of herbarium specimens using an iterative organelle genome assembly pipeline</article-title>. <source>Biol. J. Linn. Soc</source>. <volume>117</volume>, <fpage>33</fpage>&#x02013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1111/bij.12642</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bankevich</surname> <given-names>A.</given-names></name> <name><surname>Nurk</surname> <given-names>S.</given-names></name> <name><surname>Antipov</surname> <given-names>D.</given-names></name> <name><surname>Gurevich</surname> <given-names>A.</given-names></name> <name><surname>Dvorkin</surname> <given-names>M.</given-names></name> <name><surname>Kulikov</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing</article-title>. <source>J. Comput. Biol</source>. <volume>19</volume>, <fpage>455</fpage>&#x02013;<lpage>477</lpage>. <pub-id pub-id-type="doi">10.1089/cmb.2012.0021</pub-id><pub-id pub-id-type="pmid">22506599</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Bolkhovskikh</surname> <given-names>Z.</given-names></name> <name><surname>Grif</surname> <given-names>V.</given-names></name> <name><surname>Zakharieva</surname> <given-names>O.</given-names></name> <name><surname>Matveeva</surname> <given-names>T.</given-names></name></person-group> (<year>1969</year>). <source>Chromosome Numbers of Flowering Plants.</source> <publisher-loc>Moscow</publisher-loc>: <publisher-name>USSR</publisher-name>. p. <fpage>926</fpage>.</citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Borsch</surname> <given-names>T.</given-names></name> <name><surname>Hilu</surname> <given-names>K.</given-names></name> <name><surname>Quandt</surname> <given-names>D.</given-names></name> <name><surname>Wilde</surname> <given-names>V.</given-names></name> <name><surname>Neinhuis</surname> <given-names>C.</given-names></name> <name><surname>Barthlott</surname> <given-names>W.</given-names></name></person-group> (<year>2003</year>). <article-title>Noncoding plastid trnT-trnF sequences reveal a well resolved phylogeny of basal angiosperms</article-title>. <source>J. Evol. Biol</source>. <volume>16</volume>, <fpage>558</fpage>&#x02013;<lpage>576</lpage>. <pub-id pub-id-type="doi">10.1046/j.1420-9101.2003.00577.x</pub-id><pub-id pub-id-type="pmid">14632220</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Brandbyge</surname> <given-names>J.</given-names></name></person-group> (<year>1993</year>). <article-title>&#x0201C;The families and genera of vascular plants,&#x0201D;</article-title> in <source>Polygonaceae</source>, eds K. Kubitzki, J. Rohwer, and V. Bittrich (<publisher-loc>Verlag; Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>531</fpage>&#x02013;<lpage>544</lpage>.</citation>
</ref>
<ref id="B11">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Bushnell</surname> <given-names>B.</given-names></name></person-group> (<year>2015</year>). <source>BBTools Software Package v.33.89</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://sourceforge.net/projects/bbmap/">https://sourceforge.net/projects/bbmap/</ext-link><pub-id pub-id-type="pmid">28505226</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Capella-Gutierrez</surname> <given-names>S.</given-names></name> <name><surname>Silla-Martinez</surname> <given-names>J.</given-names></name> <name><surname>Gabaldon</surname> <given-names>T.</given-names></name></person-group> (<year>2009</year>). <article-title>trimAl: a tool for automated alignment trimming in large-scale phylogenetic analyses</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>1972</fpage>&#x02013;<lpage>1973</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp348</pub-id><pub-id pub-id-type="pmid">19505945</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carrion</surname> <given-names>A.</given-names></name> <name><surname>Hinsinger</surname> <given-names>D.</given-names></name> <name><surname>Strijk</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>ECuADOR-easy curation of angiosperm duplicated organellar regions, a tool for cleaning and curating plastomes assembled from next generation sequencing pipelines</article-title>. <source>PeerJ</source>. <volume>8</volume>, e8699. <pub-id pub-id-type="doi">10.7717/peerj.8699</pub-id><pub-id pub-id-type="pmid">32292644</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chapman</surname> <given-names>J.</given-names></name> <name><surname>Ho</surname> <given-names>I.</given-names></name> <name><surname>Sunkara</surname> <given-names>S.</given-names></name> <name><surname>Luo</surname> <given-names>S.</given-names></name> <name><surname>Schroth</surname> <given-names>G.</given-names></name> <name><surname>Rokhsar</surname> <given-names>D.</given-names></name></person-group> (<year>2011</year>). <article-title>Meraculous: de novo genome assembly with short paired-end reads</article-title>. <source>PLoS ONE</source> <volume>6</volume>, <fpage>e23501</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0023501</pub-id><pub-id pub-id-type="pmid">21876754</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>K.</given-names></name> <name><surname>Wallis</surname> <given-names>J.</given-names></name> <name><surname>Mclellan</surname> <given-names>M.</given-names></name> <name><surname>Larson</surname> <given-names>D.</given-names></name> <name><surname>Kalicki</surname> <given-names>J.</given-names></name> <name><surname>Pohl</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>BreakDancer: an algorithm for high-resolution mapping of genomic structural variation</article-title>. <source>Nat. Methods</source> <volume>6</volume>, <fpage>677</fpage>&#x02013;<lpage>681</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.1363</pub-id><pub-id pub-id-type="pmid">19668202</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>Coissac</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <source>Org.Asm: The Genome ORGanelle ASseMbler v.1.0.3</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://pypi.org/project/ORG.asm/">https://pypi.org/project/ORG.asm/</ext-link></citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>del Valle</surname> <given-names>J.</given-names></name> <name><surname>Casimiro-Soriguer</surname> <given-names>I.</given-names></name> <name><surname>Buide</surname> <given-names>M.</given-names></name> <name><surname>Narbona</surname> <given-names>E.</given-names></name> <name><surname>Whittall</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>Whole plastome sequencing within <italic>Silene</italic> section <italic>Psammophilae</italic> reveals mainland hybridization and divergence with the balearic island populations</article-title>. <source>Front. Plant Sci</source>. <volume>10</volume>, 1466. <pub-id pub-id-type="doi">10.3389/fpls.2019.01466</pub-id><pub-id pub-id-type="pmid">31803208</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dierckxsens</surname> <given-names>N.</given-names></name> <name><surname>Mardulyn</surname> <given-names>P.</given-names></name> <name><surname>Smits</surname> <given-names>G.</given-names></name></person-group> (<year>2017</year>). <article-title>NOVOPlasty: De novo assembly of organelle genomes from whole genome data</article-title>. <source>Nucleic Acids Res</source>. <volume>45</volume>, e<fpage>18</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw955</pub-id><pub-id pub-id-type="pmid">28204566</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Doorduin</surname> <given-names>L.</given-names></name> <name><surname>Gravendeel</surname> <given-names>B.</given-names></name> <name><surname>Lammers</surname> <given-names>Y.</given-names></name> <name><surname>Ariyurek</surname> <given-names>Y.</given-names></name> <name><surname>Chin-A-Woeng</surname> <given-names>T.</given-names></name> <name><surname>Vrieling</surname> <given-names>K.</given-names></name></person-group> (<year>2011</year>). <article-title>The complete chloroplast genome of 17 individuals of pest species <italic>Jacobaea vulgaris</italic>: SNPs, microsatellites and barcoding markers for population and phylogenetic studies</article-title>. <source>DNA Res</source>. <volume>18</volume>, <fpage>93</fpage>&#x02013;<lpage>105</lpage>. <pub-id pub-id-type="doi">10.1093/dnares/dsr002</pub-id><pub-id pub-id-type="pmid">21444340</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Earl</surname> <given-names>D.</given-names></name> <name><surname>Bradnam</surname> <given-names>K.</given-names></name> <name><surname>John</surname> <given-names>J.</given-names></name> <name><surname>Darling</surname> <given-names>A.</given-names></name> <name><surname>Lin</surname> <given-names>D.</given-names></name> <name><surname>Fass</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Assemblathon 1: a competitive assessment of <italic>de novo</italic> short read assembly methods</article-title>. <source>Genome Res</source>. <volume>21</volume>, <fpage>2224</fpage>&#x02013;<lpage>2241</lpage>. <pub-id pub-id-type="doi">10.1101/gr.126599.111</pub-id><pub-id pub-id-type="pmid">21926179</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Escobari</surname> <given-names>B.</given-names></name> <name><surname>Borsch</surname> <given-names>T.</given-names></name> <name><surname>Quedensley</surname> <given-names>T.</given-names></name> <name><surname>Gruenstaeudl</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Plastid phylogenomics of the Gynoxoid group (Senecioneae, Asteraceae) highlights the importance of motif-based sequence alignment amid low genetic distances</article-title>. <source>Am. J. Bot</source>. <volume>108</volume>, <fpage>2235</fpage>&#x02013;<lpage>2256</lpage>. <pub-id pub-id-type="doi">10.1002/ajb2.1775</pub-id><pub-id pub-id-type="pmid">34636417</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Freudenthal</surname> <given-names>J.</given-names></name> <name><surname>Pfaff</surname> <given-names>S.</given-names></name> <name><surname>Terhoeven</surname> <given-names>N.</given-names></name> <name><surname>Korte</surname> <given-names>A.</given-names></name> <name><surname>Ankenbrand</surname> <given-names>M.</given-names></name> <name><surname>Foerster</surname> <given-names>F.</given-names></name></person-group> (<year>2020</year>). <article-title>A systematic comparison of chloroplast genome assembly tools</article-title>. <source>Genome Biol</source>. <volume>21</volume>, <fpage>254</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-02153-6</pub-id><pub-id pub-id-type="pmid">32988404</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gruening</surname> <given-names>B.</given-names></name> <name><surname>Chilton</surname> <given-names>J.</given-names></name> <name><surname>Koester</surname> <given-names>J.</given-names></name> <name><surname>Dale</surname> <given-names>R.</given-names></name> <name><surname>Soranzo</surname> <given-names>N.</given-names></name> <name><surname>van den Beek</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>Practical computational reproducibility in the life sciences</article-title>. <source>Cell Syst</source>. <volume>6</volume>, <fpage>631</fpage>&#x02013;<lpage>635</lpage>. <pub-id pub-id-type="doi">10.1016/j.cels.2018.03.014</pub-id><pub-id pub-id-type="pmid">29953862</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gruenstaeudl</surname> <given-names>M.</given-names></name></person-group> (<year>2019</year>). <article-title>Why the monophyly of Nymphaeaceae currently remains indeterminate: an assessment based on gene-wise plastid phylogenomics</article-title>. <source>Plant Syst. Evolut</source>. <volume>305</volume>, <fpage>827</fpage>&#x02013;<lpage>836</lpage>. <pub-id pub-id-type="doi">10.1007/s00606-019-01610-5</pub-id></citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gruenstaeudl</surname> <given-names>M.</given-names></name> <name><surname>Gerschler</surname> <given-names>N.</given-names></name> <name><surname>Borsch</surname> <given-names>T.</given-names></name></person-group> (<year>2018</year>). <article-title>Bioinformatic workflows for generating complete plastid genome sequences-an example from <italic>Cabomba</italic> (Cabombaceae) in the context of the phylogenomic analysis of the water-lily clade</article-title>. <source>Life</source> <volume>8</volume>, <fpage>25</fpage>. <pub-id pub-id-type="doi">10.3390/life8030025</pub-id><pub-id pub-id-type="pmid">29933597</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gruenstaeudl</surname> <given-names>M.</given-names></name> <name><surname>Jenke</surname> <given-names>N.</given-names></name></person-group> (<year>2020</year>). <article-title>PACVr: plastome assembly coverage visualization in R</article-title>. <source>BMC Bioinform</source>. <volume>21</volume>, <fpage>207</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-020-3475-0</pub-id><pub-id pub-id-type="pmid">32448146</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>C.</given-names></name> <name><surname>Tembrock</surname> <given-names>L.</given-names></name> <name><surname>Johnson</surname> <given-names>N.</given-names></name> <name><surname>Simmons</surname> <given-names>M.</given-names></name> <name><surname>Wu</surname> <given-names>Z.</given-names></name></person-group> (<year>2016</year>). <article-title>The complete plastid genome of <italic>Lagerstroemia</italic> fauriei and loss of rpl2 intron from <italic>Lagerstroemia</italic> (Lythraceae)</article-title>. <source>PLoS ONE</source> <volume>11</volume>:e0150752. <pub-id pub-id-type="doi">10.1371/journal.pone.0150752</pub-id><pub-id pub-id-type="pmid">26950701</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gurevich</surname> <given-names>A.</given-names></name> <name><surname>Saveliev</surname> <given-names>V.</given-names></name> <name><surname>Vyahhi</surname> <given-names>N.</given-names></name> <name><surname>Tesler</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>QUAST: Quality assessment tool for genome assemblies</article-title>. <source>Bioinformatics</source> <volume>29</volume>, <fpage>1072</fpage>&#x02013;<lpage>1075</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt086</pub-id><pub-id pub-id-type="pmid">23422339</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huang</surname> <given-names>B.</given-names></name> <name><surname>Ruess</surname> <given-names>H.</given-names></name> <name><surname>Liang</surname> <given-names>Q.</given-names></name> <name><surname>Colleoni</surname> <given-names>C.</given-names></name> <name><surname>Spooner</surname> <given-names>D.</given-names></name></person-group> (<year>2019</year>). <article-title>Analyses of 202 plastid genomes elucidate the phylogeny of solanum section petota</article-title>. <source>Sci. Rep</source>. <volume>9</volume>, <fpage>7</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-019-40790-5</pub-id><pub-id pub-id-type="pmid">30872631</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hubisz</surname> <given-names>M.</given-names></name> <name><surname>Lin</surname> <given-names>M.</given-names></name> <name><surname>Kellis</surname> <given-names>M.</given-names></name> <name><surname>Siepel</surname> <given-names>A.</given-names></name></person-group> (<year>2011</year>). <article-title>Error and error mitigation in low-coverage genome assemblies</article-title>. <source>PLoS ONE</source> <volume>6</volume>, <fpage>e17034</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0017034</pub-id><pub-id pub-id-type="pmid">21340033</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Izan</surname> <given-names>S.</given-names></name> <name><surname>Esselink</surname> <given-names>D.</given-names></name> <name><surname>Visser</surname> <given-names>R.</given-names></name> <name><surname>Smulders</surname> <given-names>M.</given-names></name> <name><surname>Borm</surname> <given-names>T.</given-names></name></person-group> (<year>2017</year>). <article-title>De novo assembly of complete chloroplast genomes from non-model species based on a k-mer frequency-based selection of chloroplast reads from total DNA sequences</article-title>. <source>Front. Plant Sci</source>. <volume>8</volume>, 1271. <pub-id pub-id-type="doi">10.3389/fpls.2017.01271</pub-id><pub-id pub-id-type="pmid">28824658</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiang</surname> <given-names>D.</given-names></name> <name><surname>Zhao</surname> <given-names>Z.</given-names></name> <name><surname>Zhang</surname> <given-names>T.</given-names></name> <name><surname>Zhong</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Yuan</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2017</year>). <article-title>The chloroplast genome sequence of <italic>Scutellaria baicalensis</italic> provides insight into intraspecific and interspecific chloroplast genome diversity in <italic>Scutellaria</italic></article-title>. <source>Genes</source> <fpage>8</fpage>, 227. <pub-id pub-id-type="doi">10.3390/genes8090227</pub-id><pub-id pub-id-type="pmid">28902130</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jin</surname> <given-names>J.-J.</given-names></name> <name><surname>Yu</surname> <given-names>W.-B.</given-names></name> <name><surname>Yang</surname> <given-names>J.-B.</given-names></name> <name><surname>Song</surname> <given-names>Y.</given-names></name> <name><surname>dePamphilis</surname> <given-names>C. W.</given-names></name> <name><surname>Yi</surname> <given-names>T.-S.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>GetOrganelle: a fast and versatile toolkit for accurate de novo assembly of organelle genomes</article-title>. <source>Genome Biol</source>. <volume>21</volume>, <fpage>1</fpage>&#x02013;<lpage>31</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-02154-5</pub-id><pub-id pub-id-type="pmid">32912315</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kajitani</surname> <given-names>R.</given-names></name> <name><surname>Toshimoto</surname> <given-names>K.</given-names></name> <name><surname>Noguchi</surname> <given-names>H.</given-names></name> <name><surname>Toyoda</surname> <given-names>A.</given-names></name> <name><surname>Ogura</surname> <given-names>Y.</given-names></name> <name><surname>Okuno</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Efficient de novo assembly of highly heterozygous genomes from whole-genome shotgun short reads</article-title>. <source>Genome Res</source>. <volume>24</volume>, <fpage>1384</fpage>&#x02013;<lpage>1395</lpage>. <pub-id pub-id-type="doi">10.1101/gr.170720.113</pub-id><pub-id pub-id-type="pmid">24755901</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Karjagin</surname> <given-names>I.</given-names></name></person-group> (<year>1952</year>). <article-title>&#x0201C;<italic>Calligonum</italic>,&#x0201D;</article-title> in <source>Flora Azerbajd&#x0017D;ana, Vol. 3</source>, ed I. E. A. E. Karjagin (<publisher-loc>Baku</publisher-loc>: <publisher-name>Izdatelstvo Akademii nauk Azerbaidzhanskoi SSR</publisher-name>), <fpage>165</fpage>&#x02013;<lpage>166</lpage>.</citation>
</ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Katoh</surname> <given-names>K.</given-names></name> <name><surname>Standley</surname> <given-names>D.</given-names></name></person-group> (<year>2013</year>). <article-title>MAFFT multiple sequence alignment software version 7: improvements in performance and usability</article-title>. <source>Mol. Biol. Evol</source>. <volume>30</volume>, <fpage>772</fpage>&#x02013;<lpage>780</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst010</pub-id><pub-id pub-id-type="pmid">23329690</pub-id></citation></ref>
<ref id="B37">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kearse</surname> <given-names>M.</given-names></name> <name><surname>Moir</surname> <given-names>R.</given-names></name> <name><surname>Wilson</surname> <given-names>A.</given-names></name> <name><surname>Stones-Havas</surname> <given-names>S.</given-names></name> <name><surname>Sturrock</surname> <given-names>S.</given-names></name> <name><surname>Buxton</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Geneious Basic: An integrated and extendable desktop software platform for the organization and analysis of sequence data</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>1647</fpage>&#x02013;<lpage>1649</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts199</pub-id><pub-id pub-id-type="pmid">22543367</pub-id></citation></ref>
<ref id="B38">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K.</given-names></name> <name><surname>Lee</surname> <given-names>S.-C.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Yang</surname> <given-names>K.</given-names></name> <name><surname>Choi</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Complete chloroplast and ribosomal sequences for 30 accessions elucidate evolution of <italic>Oryza</italic> AA genome species</article-title>. <source>Sci. Rep</source>. <volume>5</volume>, 15655. <pub-id pub-id-type="doi">10.1038/srep15655</pub-id><pub-id pub-id-type="pmid">26506948</pub-id></citation></ref>
<ref id="B39">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Koehler</surname> <given-names>M.</given-names></name> <name><surname>Reginato</surname> <given-names>M.</given-names></name> <name><surname>Souza-Chies</surname> <given-names>T.</given-names></name> <name><surname>Majure</surname> <given-names>L.</given-names></name></person-group> (<year>2020</year>). <article-title>Insights into chloroplast genome evolution across <italic>Opuntioideae</italic> (Cactaceae) reveals robust yet sometimes conflicting phylogenetic topologies</article-title>. <source>Front. Plant Sci</source>. <volume>11</volume>, 729. <pub-id pub-id-type="doi">10.3389/fpls.2020.00729</pub-id><pub-id pub-id-type="pmid">32636853</pub-id></citation></ref>
<ref id="B40">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2014</year>). <article-title>Toward better understanding of artifacts in variant calling from high-coverage samples</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>2843</fpage>&#x02013;<lpage>2851</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu356</pub-id><pub-id pub-id-type="pmid">24974202</pub-id></citation></ref>
<ref id="B41">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liao</surname> <given-names>Y.</given-names></name> <name><surname>Lin</surname> <given-names>S.</given-names></name> <name><surname>Lin</surname> <given-names>H.</given-names></name></person-group> (<year>2015</year>). <article-title>Completing bacterial genome assemblies: strategy and performance comparisons</article-title>. <source>Sci. Rep</source>. <volume>5</volume>, 8747. <pub-id pub-id-type="doi">10.1038/srep08747</pub-id><pub-id pub-id-type="pmid">25735824</pub-id></citation></ref>
<ref id="B42">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lim</surname> <given-names>C.</given-names></name> <name><surname>Kim</surname> <given-names>G.-B.</given-names></name> <name><surname>Ryu</surname> <given-names>S.-A.</given-names></name> <name><surname>Yu</surname> <given-names>H.-J.</given-names></name> <name><surname>Mun</surname> <given-names>J.-H.</given-names></name></person-group> (<year>2018</year>). <article-title>The complete chloroplast genome of <italic>Artemisia hallaisanensis</italic> nakai (asteraceae), an endemic medicinal herb in korea</article-title>. <source>Mitochondrial DNA B</source> <volume>3</volume>, <fpage>359</fpage>&#x02013;<lpage>360</lpage>. <pub-id pub-id-type="doi">10.1080/23802359.2018.1450680</pub-id><pub-id pub-id-type="pmid">33474169</pub-id></citation></ref>
<ref id="B43">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Luo</surname> <given-names>R.</given-names></name> <name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Xie</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>Z.</given-names></name> <name><surname>Huang</surname> <given-names>W.</given-names></name> <name><surname>Yuan</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>SOAPdenovo2: an empirically improved memory-efficient short-read de novo assembler</article-title>. <source>Gigascience</source> <volume>1</volume>, <fpage>18</fpage>. <pub-id pub-id-type="doi">10.1186/2047-217X-1-18</pub-id><pub-id pub-id-type="pmid">26161257</pub-id></citation></ref>
<ref id="B44">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Magoc</surname> <given-names>T.</given-names></name> <name><surname>Pabinger</surname> <given-names>S.</given-names></name> <name><surname>Canzar</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Su</surname> <given-names>Q.</given-names></name> <name><surname>Puiu</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>GAGE-B: an evaluation of genome assemblers for bacterial organisms</article-title>. <source>Bioinformatics</source> <volume>29</volume>, <fpage>1718</fpage>&#x02013;<lpage>1725</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btt273</pub-id><pub-id pub-id-type="pmid">23665771</pub-id></citation></ref>
<ref id="B45">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Martin</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>Cutadapt removes adapter sequences from high-throughput sequencing reads</article-title>. <source>EMBnet.J</source>. <volume>17</volume>, <fpage>10</fpage>&#x02013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.14806/ej.17.1.200</pub-id></citation>
</ref>
<ref id="B46">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McCorrison</surname> <given-names>J.</given-names></name> <name><surname>Venepally</surname> <given-names>P.</given-names></name> <name><surname>Singh</surname> <given-names>I.</given-names></name> <name><surname>Fouts</surname> <given-names>D.</given-names></name> <name><surname>Lasken</surname> <given-names>R.</given-names></name> <name><surname>Methe</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>NeatFreq: reference-free data reduction and coverage normalization for de-novo sequence assembly</article-title>. <source>BMC Bioinf</source>. <volume>15</volume>, <fpage>357</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-014-0357-3</pub-id><pub-id pub-id-type="pmid">25407910</pub-id></citation></ref>
<ref id="B47">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKain</surname> <given-names>M.</given-names></name> <name><surname>Johnson</surname> <given-names>M.</given-names></name> <name><surname>Uribe-Convers</surname> <given-names>S.</given-names></name> <name><surname>Eaton</surname> <given-names>D.</given-names></name> <name><surname>Yang</surname> <given-names>Y.</given-names></name></person-group> (<year>2018</year>). <article-title>Practical considerations for plant phylogenomics</article-title>. <source>Appl. Plant Sci</source>. <volume>6</volume>, e1038. <pub-id pub-id-type="doi">10.1002/aps3.1038</pub-id><pub-id pub-id-type="pmid">29732268</pub-id></citation></ref>
<ref id="B48">
<citation citation-type="web"><person-group person-group-type="author"><name><surname>McKain</surname> <given-names>M.</given-names></name> <name><surname>Wilson</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <source>Fast-Plast v.1.2.6</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="https://github.com/mrmckain/Fast-Plast">https://github.com/mrmckain/Fast-Plast</ext-link></citation>
</ref>
<ref id="B49">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mohanta</surname> <given-names>T.</given-names></name> <name><surname>Mishra</surname> <given-names>A.</given-names></name> <name><surname>Khan</surname> <given-names>A.</given-names></name> <name><surname>Hashem</surname> <given-names>A.</given-names></name> <name><surname>Abdallah</surname> <given-names>E.</given-names></name> <name><surname>Al-Harrasi</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). <article-title>Gene loss and evolution of the plastome</article-title>. <source>Genes</source> <volume>11</volume>, <fpage>1133</fpage>. <pub-id pub-id-type="doi">10.3390/genes11101133</pub-id><pub-id pub-id-type="pmid">32992972</pub-id></citation></ref>
<ref id="B50">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Moner</surname> <given-names>A.</given-names></name> <name><surname>Furtado</surname> <given-names>A.</given-names></name> <name><surname>Henry</surname> <given-names>R.</given-names></name></person-group> (<year>2018</year>). <article-title>Chloroplast phylogeography of AA genome rice species</article-title>. <source>Mol. Phylogenet. Evol</source>. <volume>127</volume>, <fpage>475</fpage>&#x02013;<lpage>487</lpage>. <pub-id pub-id-type="doi">10.1016/j.ympev.2018.05.002</pub-id><pub-id pub-id-type="pmid">29753711</pub-id></citation></ref>
<ref id="B51">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Morrison</surname> <given-names>S.</given-names></name> <name><surname>Pyzh</surname> <given-names>R.</given-names></name> <name><surname>Jeon</surname> <given-names>M.</given-names></name> <name><surname>Amaro</surname> <given-names>C.</given-names></name> <name><surname>Roig</surname> <given-names>F.</given-names></name> <name><surname>Baker-Austin</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Impact of analytic provenance in genome analysis</article-title>. <source>BMC Genomics</source> <volume>15</volume>, <fpage>S1</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-15-S8-S1</pub-id><pub-id pub-id-type="pmid">25435180</pub-id></citation></ref>
<ref id="B52">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakamura</surname> <given-names>K.</given-names></name> <name><surname>Oshima</surname> <given-names>T.</given-names></name> <name><surname>Morimoto</surname> <given-names>T.</given-names></name> <name><surname>Ikeda</surname> <given-names>S.</given-names></name> <name><surname>Yoshikawa</surname> <given-names>H.</given-names></name> <name><surname>Shiwa</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>Sequence-specific error profile of Illumina sequencers</article-title>. <source>Nucleic Acids Res</source>. <volume>39</volume>, e<fpage>90</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr344</pub-id><pub-id pub-id-type="pmid">21576222</pub-id></citation></ref>
<ref id="B53">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Olson</surname> <given-names>N.</given-names></name> <name><surname>Treangen</surname> <given-names>T.</given-names></name> <name><surname>Hill</surname> <given-names>C.</given-names></name> <name><surname>Cepeda-Espinoza</surname> <given-names>V.</given-names></name> <name><surname>Ghurye</surname> <given-names>J.</given-names></name> <name><surname>Koren</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Metagenomic assembly through the lens of validation: recent advances in assessing and improving the quality of genomes assembled from metagenomes</article-title>. <source>Brief Bioinform</source>. <volume>20</volume>, <fpage>1140</fpage>&#x02013;<lpage>1150</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbx098</pub-id><pub-id pub-id-type="pmid">28968737</pub-id></citation></ref>
<ref id="B54">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pedersen</surname> <given-names>B.</given-names></name> <name><surname>Collins</surname> <given-names>R.</given-names></name> <name><surname>Talkowski</surname> <given-names>M.</given-names></name> <name><surname>Quinlan</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Indexcov: fast coverage quality control for whole-genome sequencing</article-title>. <source>Gigascience</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>6</lpage>. <pub-id pub-id-type="doi">10.1093/gigascience/gix090</pub-id><pub-id pub-id-type="pmid">29048539</pub-id></citation></ref>
<ref id="B55">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Peng</surname> <given-names>Y.</given-names></name> <name><surname>Leung</surname> <given-names>H.</given-names></name> <name><surname>Yiu</surname> <given-names>S.</given-names></name> <name><surname>Chin</surname> <given-names>F.</given-names></name></person-group> (<year>2012</year>). <article-title>IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>1420</fpage>&#x02013;<lpage>1428</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts.174</pub-id><pub-id pub-id-type="pmid">22495754</pub-id></citation></ref>
<ref id="B56">
<citation citation-type="web"><person-group person-group-type="author"><collab>R Development Core Team</collab></person-group> (<year>2019</year>). <source>R: A Language and Environment for Statistical Computing. Vienna: Computing, R Foundation for Statistical</source>. Available online at: <ext-link ext-link-type="uri" xlink:href="http://www.r-project.org/">http://www.r-project.org/</ext-link></citation>
</ref>
<ref id="B57">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rogalski</surname> <given-names>M.</given-names></name> <name><surname>Nascimento Vieira</surname> <given-names>L.</given-names></name> <name><surname>Fraga</surname> <given-names>H.</given-names></name> <name><surname>Guerra</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>Plastid genomics in horticultural species: importance and applications for plant population genetics, evolution, and biotechnology</article-title>. <source>Front. Plant Sci</source>. <volume>6</volume>, 586. <pub-id pub-id-type="doi">10.3389/fpls.2015.00586</pub-id><pub-id pub-id-type="pmid">26284102</pub-id></citation></ref>
<ref id="B58">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Ruhlman</surname> <given-names>T.</given-names></name> <name><surname>Jansen</surname> <given-names>R.</given-names></name></person-group> (<year>2014</year>). <article-title>&#x0201C;The plastid genomes of flowering plants,&#x0201D;</article-title> in <source>Chloroplast Biotechnology, volume 1132 of Methods in Molecular Biology (Methods and Protocols)</source>, ed P. Maliga (<publisher-loc>Totowa, NJ</publisher-loc>: <publisher-name>Humana Press</publisher-name>), <fpage>3</fpage>&#x02013;<lpage>38</lpage>.<pub-id pub-id-type="pmid">34028761</pub-id></citation></ref>
<ref id="B59">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Saarela</surname> <given-names>J.</given-names></name> <name><surname>Burke</surname> <given-names>S.</given-names></name> <name><surname>Wysocki</surname> <given-names>W.</given-names></name> <name><surname>Barrett</surname> <given-names>M.</given-names></name> <name><surname>Clark</surname> <given-names>L.</given-names></name> <name><surname>Craine</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>A 250 plastome phylogeny of the grass family (Poaceae): topological support under different data partitions</article-title>. <source>PeerJ</source>. <volume>6</volume>, e4299. <pub-id pub-id-type="doi">10.7717/peerj.4299</pub-id><pub-id pub-id-type="pmid">29416954</pub-id></citation></ref>
<ref id="B60">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Salinas</surname> <given-names>N.</given-names></name> <name><surname>Little</surname> <given-names>D.</given-names></name></person-group> (<year>2014</year>). <article-title>2matrix: a utility for indel coding and phylogenetic matrix concatenation</article-title>. <source>Appl. Plant. Sci</source>. <volume>2</volume>, apps.1300083. <pub-id pub-id-type="doi">10.3732/apps.1300083</pub-id><pub-id pub-id-type="pmid">25202595</pub-id></citation></ref>
<ref id="B61">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Scarcelli</surname> <given-names>N.</given-names></name> <name><surname>Mariac</surname> <given-names>C.</given-names></name> <name><surname>Couvreur</surname> <given-names>T. L. P.</given-names></name> <name><surname>Faye</surname> <given-names>A.</given-names></name> <name><surname>Richard</surname> <given-names>D.</given-names></name> <name><surname>Sabot</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Intra-individual polymorphism in chloroplasts from NGS data: where does it come from and how to handle it?</article-title> <source>Mol. Ecol. Resour</source>. <volume>16</volume>, <fpage>434</fpage>&#x02013;<lpage>445</lpage>. <pub-id pub-id-type="doi">10.1111/1755-0998.12462</pub-id><pub-id pub-id-type="pmid">26388536</pub-id></citation></ref>
<ref id="B62">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharpe</surname> <given-names>R.</given-names></name> <name><surname>Williamson-Benavides</surname> <given-names>B.</given-names></name> <name><surname>Edwards</surname> <given-names>G.</given-names></name> <name><surname>Dhingra</surname> <given-names>A.</given-names></name></person-group> (<year>2020</year>). Methods of analysis of chloroplast genomes of C3, Kranz type C4 and single cell C4 photosynthetic members of <italic>Chenopodiaceae. Plant Methods</italic>. <volume>16</volume>, <fpage>119</fpage>. <pub-id pub-id-type="doi">10.1186/s13007-020-00662-w</pub-id><pub-id pub-id-type="pmid">32874195</pub-id></citation></ref>
<ref id="B63">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Simmons</surname> <given-names>M.</given-names></name> <name><surname>Ochoterena</surname> <given-names>H.</given-names></name></person-group> (<year>2000</year>). <article-title>Gaps as characters in sequence-based phylogenetic analyses</article-title>. <source>Syst. Biol</source>. <volume>49</volume>, <fpage>369</fpage>&#x02013;<lpage>381</lpage>. <pub-id pub-id-type="doi">10.1093/sysbio/49.2.369</pub-id><pub-id pub-id-type="pmid">12118412</pub-id></citation></ref>
<ref id="B64">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sims</surname> <given-names>D.</given-names></name> <name><surname>Sudbery</surname> <given-names>I.</given-names></name> <name><surname>Ilott</surname> <given-names>N.</given-names></name> <name><surname>Heger</surname> <given-names>A.</given-names></name> <name><surname>Ponting</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Sequencing depth and coverage: key considerations in genomic analyses</article-title>. <source>Nat. Rev. Genet</source>. <volume>15</volume>, <fpage>121</fpage>&#x02013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3642</pub-id><pub-id pub-id-type="pmid">24434847</pub-id></citation></ref>
<ref id="B65">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Song</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>T.</given-names></name> <name><surname>Burgess</surname> <given-names>K.</given-names></name> <name><surname>Feng</surname> <given-names>Y.</given-names></name> <name><surname>Ge</surname> <given-names>X.-J.</given-names></name></person-group> (<year>2020</year>). <article-title>Complete plastome sequencing resolves taxonomic relationships among species of <italic>Calligonum</italic> L.(Polygonaceae) in China</article-title>. <source>BMC Plant Biol</source>. <volume>20</volume>, <fpage>1</fpage>&#x02013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1186/s12870-020-02466-5</pub-id><pub-id pub-id-type="pmid">32513105</pub-id></citation></ref>
<ref id="B66">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soorni</surname> <given-names>A.</given-names></name> <name><surname>Haak</surname> <given-names>D.</given-names></name> <name><surname>Zaitlin</surname> <given-names>D.</given-names></name> <name><surname>Bombarely</surname> <given-names>A.</given-names></name></person-group> (<year>2017</year>). <article-title>Organelle_PBA, a pipeline for assembling chloroplast and mitochondrial genomes from PacBio DNA sequencing data</article-title>. <source>BMC Genomics</source> <volume>18</volume>, <fpage>49</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-016-3412-9</pub-id><pub-id pub-id-type="pmid">28061749</pub-id></citation></ref>
<ref id="B67">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Soskov</surname> <given-names>Y.</given-names></name> <name><surname>Akhmed-Zade</surname> <given-names>F.</given-names></name></person-group> (<year>1974</year>). <article-title>Characteristics of habitats and polymorphism of the Azerbaijan endemic <italic>Calligonum</italic> bakuense Litv</article-title>. <source>Bull. Moscow Soc. Natur. Biol. Ser</source>. <volume>59</volume>, <fpage>109</fpage>&#x02013;<lpage>114</lpage>.</citation>
</ref>
<ref id="B68">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Soskov</surname> <given-names>Y.</given-names></name></person-group> (<year>2011</year>). <source>The Genus Calligonum L.: Taxonomy, Distribution, Evolution, Introduction</source>. <publisher-loc>Novosibirsk</publisher-loc>: <publisher-name>Russian Academy of Agricultural Sciences</publisher-name>. p. <fpage>361</fpage>.</citation>
</ref>
<ref id="B69">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Souvorov</surname> <given-names>A.</given-names></name> <name><surname>Agarwala</surname> <given-names>R.</given-names></name> <name><surname>Lipman</surname> <given-names>D.</given-names></name></person-group> (<year>2018</year>). <article-title>SKESA: strategic k-mer extension for scrupulous assemblies</article-title>. <source>Genome Biol</source>. <volume>19</volume>, <fpage>153</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-018-1540-z</pub-id><pub-id pub-id-type="pmid">30286803</pub-id></citation></ref>
<ref id="B70">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stadermann</surname> <given-names>K.</given-names></name> <name><surname>Weisshaar</surname> <given-names>B.</given-names></name> <name><surname>Holtgr&#x000E4;we</surname> <given-names>D.</given-names></name></person-group> (<year>2015</year>). <article-title>SMRT sequencing only de novo assembly of the sugar beet (Beta vulgaris) chloroplast genome</article-title>. <source>BMC Bioinform</source>. <volume>16</volume>, <fpage>295</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-015-0726-6</pub-id><pub-id pub-id-type="pmid">26377912</pub-id></citation></ref>
<ref id="B71">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Stamatakis</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>RAxML version 8: a tool for phylogenetic analysis and post-analysis of large phylogenies</article-title>. <source>Bioinformatics</source> <volume>30</volume>, <fpage>1312</fpage>&#x02013;<lpage>1313</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu033</pub-id><pub-id pub-id-type="pmid">24451623</pub-id></citation></ref>
<ref id="B72">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tavakkoli</surname> <given-names>S.</given-names></name> <name><surname>Osaloo</surname> <given-names>S. K.</given-names></name> <name><surname>Maassoumi</surname> <given-names>A.</given-names></name></person-group> (<year>2010</year>). <article-title>The phylogeny of <italic>Calligonum</italic> and <italic>Pteropyrum</italic> (Polygonaceae) based on nuclear ribosomal DNA ITS and chloroplast trnL-F sequences</article-title>. <source>Iran J. Biotechnol</source>. <volume>8</volume>, <fpage>7</fpage>&#x02013;<lpage>15</lpage>.</citation>
</ref>
<ref id="B73">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Teshome</surname> <given-names>G.</given-names></name> <name><surname>Mekbib</surname> <given-names>Y.</given-names></name> <name><surname>Hu</surname> <given-names>G.</given-names></name> <name><surname>Li</surname> <given-names>Z.-Z.</given-names></name> <name><surname>Chen</surname> <given-names>J.</given-names></name></person-group> (<year>2020</year>). <article-title>Comparative analyses of 32 complete plastomes of Tef (<italic>Eragrostis tef</italic> ) accessions from Ethiopia: phylogenetic relationships and mutational hotspots</article-title>. <source>PeerJ</source>. <volume>8</volume>, e9314. <pub-id pub-id-type="doi">10.7717/peerj.9314</pub-id><pub-id pub-id-type="pmid">32596045</pub-id></citation></ref>
<ref id="B74">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Twyford</surname> <given-names>A.</given-names></name> <name><surname>Ness</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>Strategies for complete plastid genome sequencing</article-title>. <source>Mol. Ecol. Resour</source>. <volume>17</volume>, <fpage>858</fpage>&#x02013;<lpage>868</lpage>. <pub-id pub-id-type="doi">10.1111/1755-0998.12626</pub-id><pub-id pub-id-type="pmid">27790830</pub-id></citation></ref>
<ref id="B75">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Walker</surname> <given-names>B.</given-names></name> <name><surname>Abeel</surname> <given-names>T.</given-names></name> <name><surname>Shea</surname> <given-names>T.</given-names></name> <name><surname>Priest</surname> <given-names>M.</given-names></name> <name><surname>Abouelliel</surname> <given-names>A.</given-names></name> <name><surname>Sakthikumar</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Pilon: an integrated tool for comprehensive microbial variant detection and genome assembly improvement</article-title>. <source>PLoS ONE</source> <volume>9</volume>, <fpage>e112963</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0112963</pub-id><pub-id pub-id-type="pmid">25409509</pub-id></citation></ref>
<ref id="B76">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Walker</surname> <given-names>J.</given-names></name> <name><surname>Jansen</surname> <given-names>R.</given-names></name> <name><surname>Zanis</surname> <given-names>M.</given-names></name> <name><surname>Emery</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Sources of inversion variation in the small single copy (SSC) region of chloroplast genomes</article-title>. <source>Am. J. Bot</source>. <volume>102</volume>, <fpage>1751</fpage>&#x02013;<lpage>1752</lpage>. <pub-id pub-id-type="doi">10.3732/ajb.1500299</pub-id><pub-id pub-id-type="pmid">26546126</pub-id></citation></ref>
<ref id="B77">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>W.</given-names></name> <name><surname>Lanfear</surname> <given-names>R.</given-names></name></person-group> (<year>2019</year>). <article-title>Long-reads reveal that the chloroplast genome exists in two distinct versions in most plants</article-title>. <source>Genome Biol. Evol</source>. <volume>11</volume>, <fpage>3372</fpage>&#x02013;<lpage>3381</lpage>. <pub-id pub-id-type="doi">10.1093/gbe/evz256</pub-id><pub-id pub-id-type="pmid">31750905</pub-id></citation></ref>
<ref id="B78">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Whittall</surname> <given-names>J.</given-names></name> <name><surname>Syring</surname> <given-names>J.</given-names></name> <name><surname>Parks</surname> <given-names>M.</given-names></name> <name><surname>Buenrostro</surname> <given-names>J.</given-names></name> <name><surname>Dick</surname> <given-names>C.</given-names></name> <name><surname>Liston</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Finding a (pine) needle in a haystack: chloroplast genome sequence divergence in rare and widespread pines</article-title>. <source>Mol. Ecol</source>. <volume>19</volume>, <fpage>100</fpage>&#x02013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-294X.2009.04474.x</pub-id><pub-id pub-id-type="pmid">20331774</pub-id></citation></ref>
<ref id="B79">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>P.</given-names></name> <name><surname>Chen</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Zhang</surname> <given-names>X.-C.</given-names></name> <name><surname>Zhou</surname> <given-names>S.-L.</given-names></name></person-group> (<year>2021</year>). <article-title>NOVOWrap: an automated solution for plastid genome assembly and structure standardization</article-title>. <source>Mol. Ecol. Resour</source>. <volume>21</volume>, <fpage>2177</fpage>&#x02013;<lpage>2186</lpage>. <pub-id pub-id-type="doi">10.1111/1755-0998.13410</pub-id><pub-id pub-id-type="pmid">33934526</pub-id></citation></ref>
<ref id="B80">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z.</given-names></name> <name><surname>Tembrock</surname> <given-names>L.</given-names></name> <name><surname>Ge</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>Are differences in genomic data sets due to true biological variants or errors in genome assembly: an example from two chloroplast genomes</article-title>. <source>PLoS ONE</source> <volume>10</volume>, <fpage>e0118019</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0118019</pub-id><pub-id pub-id-type="pmid">25658309</pub-id></citation></ref>
<ref id="B81">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>L.-S.</given-names></name> <name><surname>Herrando-Moraira</surname> <given-names>S.</given-names></name> <name><surname>Susanna</surname> <given-names>A.</given-names></name> <name><surname>Galbany-Casals</surname> <given-names>M.</given-names></name> <name><surname>Chen</surname> <given-names>Y.-S.</given-names></name></person-group> (<year>2019</year>). <article-title>Phylogeny, origin and dispersal of <italic>Saussurea</italic> (Asteraceae) based on chloroplast genome data</article-title>. <source>Mol. Phylogenet. Evol</source>. <volume>141</volume>, 106613. <pub-id pub-id-type="doi">10.1016/j.ympev.2019.106613</pub-id><pub-id pub-id-type="pmid">31525421</pub-id></citation></ref>
<ref id="B82">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Takayama</surname> <given-names>K.</given-names></name> <name><surname>Youn</surname> <given-names>J.-S.</given-names></name> <name><surname>Pak</surname> <given-names>J.-H.</given-names></name> <name><surname>Kim</surname> <given-names>S.-C.</given-names></name></person-group> (<year>2020</year>). <article-title>Plastome characterization and phylogenomics of east asian beeches with a special emphasis on <italic>Fagus multinervis</italic> on ulleung island, korea</article-title>. <source>Genes</source> <volume>11</volume>, <fpage>1338</fpage>. <pub-id pub-id-type="doi">10.3390/genes11111338</pub-id><pub-id pub-id-type="pmid">33198274</pub-id></citation></ref>
<ref id="B83">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.-B.</given-names></name> <name><surname>Tang</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>H.-T.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.-R.</given-names></name> <name><surname>Li</surname> <given-names>D.-Z.</given-names></name></person-group> (<year>2013</year>). <article-title>Complete chloroplast genome of the genus <italic>Cymbidium</italic>: lights into the species identification, phylogenetic implications and population genetic analyses</article-title>. <source>BMC Evol. Biol</source>. <volume>13</volume>, 84. <pub-id pub-id-type="doi">10.1186/1471-2148-13-84</pub-id><pub-id pub-id-type="pmid">23597078</pub-id></citation></ref>
<ref id="B84">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yu</surname> <given-names>Y.</given-names></name> <name><surname>Ouyang</surname> <given-names>Y.</given-names></name> <name><surname>Yao</surname> <given-names>W.</given-names></name></person-group> (<year>2018</year>). <article-title>shinyCircos: an R/Shiny application for interactive creation of Circos plot</article-title>. <source>Bioinformatics</source> <volume>34</volume>, <fpage>1229</fpage>&#x02013;<lpage>1231</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btx763</pub-id><pub-id pub-id-type="pmid">29186362</pub-id></citation></ref>
</ref-list> 
</back>
</article> 