<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2018.00006</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Method for the Identification of Taxon-Specific <italic>k</italic>-mers from Chloroplast Genome: A Case Study on Tomato Plant (<italic>Solanum lycopersicum</italic>)</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Raime</surname> <given-names>Kairi</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/469149/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Remm</surname> <given-names>Maido</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/243037/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><institution>Department of Bioinformatics, Institute of Molecular and Cell Biology, University of Tartu</institution>, <addr-line>Tartu</addr-line>, <country>Estonia</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Gustavo Glusman, Institute for Systems Biology, United States</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Thiruvarangan Ramaraj, National Center for Genome Resources, United States; Juan Caballero, Universidad Aut&#x00F3;noma de Quer&#x00E9;taro, Mexico</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>Kairi Raime, <email>kairi.raime@ut.ee</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Bioinformatics and Computational Biology, a section of the journal Frontiers in Plant Science</p></fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>01</month>
<year>2018</year>
</pub-date>
<pub-date pub-type="collection">
<year>2018</year>
</pub-date>
<volume>9</volume>
<elocation-id>6</elocation-id>
<history>
<date date-type="received">
<day>24</day>
<month>08</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>03</day>
<month>01</month>
<year>2018</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2018 Raime and Remm.</copyright-statement>
<copyright-year>2018</copyright-year>
<copyright-holder>Raime and Remm</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Polymerase chain reaction and different barcoding methods commonly used for plant identification from metagenomics samples are based on the amplification of a limited number of pre-selected barcoding regions. These methods are often inapplicable due to DNA degradation, low amplification success or low species discriminative power of selected genomic regions. Here we introduce a method for the rapid identification of plant taxon-specific <italic>k</italic>-mers, that is applicable for the fast detection of plant taxa directly from raw sequencing reads without aligning, mapping or assembling the reads. We identified more than 800 <italic>Solanum lycopersicum</italic> specific <italic>k</italic>-mers (32 nucleotides in length) from 42 different chloroplast genome regions using the developed method. We demonstrated that identified <italic>k</italic>-mers are also detectable in whole genome sequencing raw reads from <italic>S. lycopersicum</italic>. Also, we demonstrated the usability of taxon-specific <italic>k</italic>-mers in artificial mixtures of sequences from closely related species. Developed method offers a novel strategy for fast identification of taxon-specific genome regions and offers new perspectives for detection of plant taxa directly from sequencing raw reads.</p>
</abstract>
<kwd-group>
<kwd><italic>k</italic>-mer-based method</kwd>
<kwd>taxon-specific <italic>k</italic>-mers</kwd>
<kwd>plant taxon identification</kwd>
<kwd>raw sequencing reads</kwd>
<kwd><italic>Solanum lycopersicum</italic></kwd>
</kwd-group>
<counts>
<fig-count count="7"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="65"/>
<page-count count="12"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>The molecular identification of plant species is one possible approach when processed material has to be analyzed and the visual inspection of morphology is not possible. Using DNA-based tests for authentication plays an important role in the detection of mislabelled species and inadvertent or intentional species substitutions in food products (<xref ref-type="bibr" rid="B15">Galimberti et al., 2013</xref>; <xref ref-type="bibr" rid="B13">Ferri et al., 2015</xref>), detecting species and their origin in forensic samples in investigating of crimes (e.g., illegal trade of flora and fauna) (<xref ref-type="bibr" rid="B19">Iyengar, 2014</xref>), the identification of unlabelled herbal ingredients in herbal drugs (<xref ref-type="bibr" rid="B37">Newmaster et al., 2013</xref>) and the analysis of the composition of decayed plant material from soil, e.g., root or leaf litter samples (<xref ref-type="bibr" rid="B56">Wallinger et al., 2012</xref>).</p>
<p>DNA barcoding is a widely used method for the identification of plant, animal or fungi taxa by sequencing a standardized short DNA fragment. Barcoding methods identify taxa very efficiently within metagenomics samples of different origin (<xref ref-type="bibr" rid="B6">Coghlan et al., 2012</xref>; <xref ref-type="bibr" rid="B37">Newmaster et al., 2013</xref>; <xref ref-type="bibr" rid="B55">Tillmar et al., 2013</xref>; <xref ref-type="bibr" rid="B65">Zhou et al., 2013</xref>), but the success of the identification or detection of species depends on the selected loci. DNA barcodes usually constitute only a very small part of the genome (usually less than 1000 bp), and due to different limiting factors (e.g., low PCR efficiency, gene deletions, and insufficient variation of selected barcode regions), no single-locus barcode has been identified as a universal DNA barcode region for identification of all plants. Using multi-locus combinations of two or three loci has been suggested (<xref ref-type="bibr" rid="B36">Newmaster et al., 2006</xref>; <xref ref-type="bibr" rid="B27">Kress and Erickson, 2007</xref>; <xref ref-type="bibr" rid="B1">Cbol Plant Working, 2009</xref>; <xref ref-type="bibr" rid="B34">Mishra et al., 2016</xref>). Using only a few barcoding regions may still not be enough for differentiate phylogenetically close plant species (<xref ref-type="bibr" rid="B5">Clement and Donoghue, 2012</xref>; <xref ref-type="bibr" rid="B63">Zhang et al., 2012</xref>), but designing primers for the amplification of hundreds or thousands of different barcode regions is very time-consuming and costly.</p>
<p>DNA metabarcoding method bases on next-generation sequencing of pre-amplified DNA barcodes for identification of multiple taxa simultaneously from a single metagenomics sample and has also been applied for the identification of plant and animal species from the various metagenomics samples (<xref ref-type="bibr" rid="B52">Staats et al., 2016</xref>). However, the barcoding of samples containing multiple taxa can be complicated in samples with degraded DNA. Varied polymerase chain reaction (PCR) success for selected genes, gene copy numbers and PCR bias caused by primer-template mismatches across species may cause species to be missed and reduce the quantitative potential of the method (<xref ref-type="bibr" rid="B11">Elbrecht and Leese, 2015</xref>; <xref ref-type="bibr" rid="B46">Pi&#x00F1;ol et al., 2015</xref>).</p>
<p>DNA mini-barcodes with significantly reduced length barcode sequences (100&#x2013;250 bp) have been introduced to improve PCR amplification success of samples with degraded DNA (<xref ref-type="bibr" rid="B8">Dong et al., 2014</xref>; <xref ref-type="bibr" rid="B51">Shokralla et al., 2015</xref>). Mini-barcodes from the chloroplast genome display very low inter-populational variations and are found to be almost free of intra-populational variations and, therefore, sequence divergence is predominantly between species. Taxa mini-barcodes with high resolution are not easily found (<xref ref-type="bibr" rid="B8">Dong et al., 2014</xref>). All PCR-based detection methods (including all barcoding methods) are able to detect only DNA from species to which PCR primers bind efficiently. Therefore, discriminating power of different barcoding methods is directly dependent on the selected barcode markers and reference database composition (<xref ref-type="bibr" rid="B14">Ficetola et al., 2010</xref>; <xref ref-type="bibr" rid="B18">Hollingsworth et al., 2011</xref>; <xref ref-type="bibr" rid="B8">Dong et al., 2014</xref>).</p>
<p>More innovative approaches for the identification of species from metagenomics samples include the deep sequencing of total genomic DNA from samples with various taxon composition. These potentially avoid some of the problems associated with targeted PCR-based methods. Therefore, these has also been suggested as a valuable tool for species identification and quantification in food testing (<xref ref-type="bibr" rid="B47">Ripp et al., 2014</xref>).</p>
<p>However, the taxonomic classification of metagenome sequencing reads can be challenging, especially when analyzing short reads derived from next generation sequencing. One approach for the identification of taxa is comparing the pre-assembled reads to the reference genomes. However, metagenomics assemblies and the quantification of taxa from contigs is computationally challenging and assembly free methods that are based on sequence alignment are too slow to cope with the increasing amount of available genomic data (<xref ref-type="bibr" rid="B32">Menzel et al., 2016</xref>).</p>
<p>Thus, the algorithms have been developed that are not using traditional alignment methods for the taxonomic classification of individual sequencing reads, but are based on the hash-based index structures created from a set of reference sequences. For the taxonomic assignment of the reads, all the <italic>k</italic>-mers (short exact-matching substrings of a fixed-length <italic>k</italic>) contained in the reference genomes are stored in an index for fast lookup and the <italic>k</italic>-mers in each sequencing read are searched in this index. The read is assigned to a taxon based on the matching genomes (<xref ref-type="bibr" rid="B32">Menzel et al., 2016</xref>). Recent programs following this approach are Kraken (<xref ref-type="bibr" rid="B61">Wood and Salzberg, 2014</xref>), Clark (<xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>), GenomeTester4 (<xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>), and Centrifuge (<xref ref-type="bibr" rid="B24">Kim et al., 2016</xref>). These programs do not require a genome assembly or mapping of the data to a reference genome. The analysis can be performed directly on sequencing reads and therefore has the potential to be less error-prone and faster than traditional methods (<xref ref-type="bibr" rid="B45">Patro et al., 2014</xref>; <xref ref-type="bibr" rid="B61">Wood and Salzberg, 2014</xref>; <xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>; <xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>). There are many examples for applications of <italic>k</italic>-mer-based methods in the detection of bacterial taxa in metagenomics samples (<xref ref-type="bibr" rid="B61">Wood and Salzberg, 2014</xref>; <xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>; <xref ref-type="bibr" rid="B24">Kim et al., 2016</xref>; <xref ref-type="bibr" rid="B48">Roosaare et al., 2017</xref>), but most of these are not developed or tested for the identification or detection of plant taxa in metagenomics samples. As a result of analysis of the sequencing reads of fruit shake containing more than a dozen plant species using Centrifuge (<xref ref-type="bibr" rid="B24">Kim et al., 2016</xref>) about half of the plant species were identified. Many plant species remained unidentified, probably because their genome sequences were substantially different compared to those in the database used for analysis or because the abundance of those plants was very low in the sample. There were also some problems with discriminating close species (e.g., apple and pear) (<xref ref-type="bibr" rid="B24">Kim et al., 2016</xref>).</p>
<p>More than 100 nuclear plant genomes have been sequenced and many more are expected in the years to come (<xref ref-type="bibr" rid="B58">Weigel and Mott, 2009</xref>; <xref ref-type="bibr" rid="B33">Michael and Jackson, 2013</xref>; <xref ref-type="bibr" rid="B39">Nystedt et al., 2013</xref>). However, unlike bacteria, the number of sequenced nuclear plant genomes is currently not sufficient for finding taxon-specific <italic>k</italic>-mers. Mitochondrial genomes from plants are a poor choice for finding taxon-specific DNA regions, because plant mitochondrial DNA evolves slowly in sequence (<xref ref-type="bibr" rid="B43">Palmer and Herbon, 1988</xref>; <xref ref-type="bibr" rid="B10">Drouin et al., 2008</xref>) and because of intra-individual variability (caused by heteroplasmy) (<xref ref-type="bibr" rid="B26">Kmiec et al., 2006</xref>).</p>
<p>The low variation in the plastid genes analyzed for DNA barcoding and the availability of more than 2500 sequenced chloroplast genomes from a variety of land plants has led to the idea of using entire plastid genome sequences to improve the resolution in resolving evolutionary relationships at lower taxonomic levels with limited sequence variation (<xref ref-type="bibr" rid="B44">Parks et al., 2009</xref>; <xref ref-type="bibr" rid="B59">Whittall et al., 2010</xref>; <xref ref-type="bibr" rid="B38">Nock et al., 2011</xref>; <xref ref-type="bibr" rid="B64">Zhang et al., 2011</xref>; <xref ref-type="bibr" rid="B22">Kane et al., 2012</xref>; <xref ref-type="bibr" rid="B62">Yang et al., 2013</xref>; <xref ref-type="bibr" rid="B49">Ruhsam et al., 2015</xref>). The size of most chloroplast genomes in higher plants ranges between 115 and 165 kb and show a high similarity in their structure and gene organization (<xref ref-type="bibr" rid="B20">Jansen et al., 2005</xref>). The chloroplast genome contains a large single-copy (LSC) and a small single-copy (SSC) region, which are separated by two copies of an inverted repeat (IR) (<xref ref-type="bibr" rid="B53">Sugiura et al., 1998</xref>). It is also known that IR-s are highly conserved among plants. The mutation rate and sequence variability in the single-copy regions of the plastome is higher than in the IR-s (<xref ref-type="bibr" rid="B60">Wolfe et al., 1987</xref>; <xref ref-type="bibr" rid="B31">Maier et al., 1995</xref>; <xref ref-type="bibr" rid="B21">Kahlau et al., 2006</xref>; <xref ref-type="bibr" rid="B9">Dong et al., 2012</xref>).</p>
<p>The advantage of using taxon-specific DNA <italic>k</italic>-mers found in the chloroplast genome for the identification of plant taxa from metagenomics samples is that the chloroplast genome is endemic to plants and may help to bypass DNA contamination from organisms without chloroplasts (e.g., animals and fungi) (<xref ref-type="bibr" rid="B8">Dong et al., 2014</xref>). Chloroplast genome-derived markers are reliable for the identification of plant species, as the chloroplast DNA is high copy and has small and generally stable, mechanical breakdown resistant, circular form compared to nuclear DNA (<xref ref-type="bibr" rid="B25">Kim et al., 2015</xref>).</p>
<p>In this work, we introduce a fast pipeline for the identification plant taxon-specific <italic>k</italic>-mers from the chloroplast genome, that can be used for the qualitative detection of plant taxa directly from raw sequencing data. We used the plant species <italic>Solanum lycopersicum</italic> (tomato plant) as an example for finding species-specific <italic>k</italic>-mers. We also analyzed the presence and number of taxon-specific <italic>k</italic>-mers (identified from chloroplast genome) in whole genome sequencing raw reads of different <italic>Solanaceae</italic> species.</p>
</sec>
<sec><title>Results</title>
<sec><title>Pipeline for Selecting Taxon-Specific <italic>k</italic>-mers</title>
<p>The first step for selecting the taxon-specific <italic>k</italic>-mers is creating the <italic>k</italic>-mer lists of every target sequence and <italic>k</italic>-mer list of all non-targets (<bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>). Target sequences and non-target sequences used as inputs can be assembled chloroplast genome sequences or any other genome regions for target and non-target taxa.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>The workflow for identifying taxon-specific <italic>k</italic>-mers. The process of finding a set of taxon-specific <italic>k</italic>-mers starts with creating <italic>k</italic>-mer lists for each sequence of target taxon and all sequences of non-target taxa (from FASTA or FASTQ files). Next the set of <italic>k</italic>-mers that are not present in a specified number of target sequences or are also present in any non-target sequences are removed to identify the set of taxon-specific <italic>k</italic>-mers.</p></caption>
<graphic xlink:href="fpls-09-00006-g001.tif"/>
</fig>
<p>The choice of the most appropriate <italic>k</italic> value (the length of oligomers) depends on the specific task. Reducing the length of <italic>k</italic>-mers when selecting taxon-specific <italic>k</italic>-mers would increase the number of <italic>k</italic>-mers that are present in the different target-taxon sequences, but with the cost of increasing the probability of finding those <italic>k</italic>-mers in sequences from non-target taxa (<xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>).</p>
<p>The next step is finding a union of unique target <italic>k</italic>-mers that are present in the specified number of target sequences. The final step is comparing the list of target <italic>k</italic>-mers to the list of non-targets <italic>k</italic>-mers to identify the final list of target taxon specific <italic>k</italic>-mers that are present in the specified numbers of target sequences but are not present in any non-target sequences. GenomeTester4 programs (<xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>) have been used for creating and comparing the <italic>k</italic>-mer lists.</p>
</sec>
<sec><title><italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers</title>
<p>To test whether species-specific <italic>k</italic>-mers can be selected for real species, we tested our pipeline for the detection of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers using available and assembled chloroplast genome sequences.</p>
<p>Although the cultivated tomato plant, <italic>S. lycopersicum</italic>, has been a subject of extensive breeding programs, the nucleotide sequences of plastid genomes in different tomato varieties show very low sequence variation (<xref ref-type="bibr" rid="B21">Kahlau et al., 2006</xref>). The overall chloroplast genomic structures and sequences from different species in the <italic>Solanaceae</italic> family have been found to be quite similar, though an analysis of complete alignment of the chloroplast genome sequences identified a number of genetic variations in the intergenic spacer regions and protein-coding genes between different <italic>Solanaceae</italic> species (<xref ref-type="bibr" rid="B3">Chung et al., 2006</xref>). Complete chloroplast genome sequences have been rarely used to discriminate <italic>Solanum</italic> species (<xref ref-type="bibr" rid="B16">Gargano et al., 2012</xref>; <xref ref-type="bibr" rid="B2">Cho and Park, 2016</xref>).</p>
<p>To get an overview of the available data for chloroplast genome sequences, we constructed a tree containing all 1,719 chloroplast genome sequences downloaded from the GenBank database. According to the tree, all five <italic>S. lycopersicum</italic> plastid genome sequences are in the same branch of the tree, and the <italic>S. lycopersicum</italic> sequences are the most closely related to <italic>S. pimpinellifolium</italic> (wild species of tomato plant) sequences (<bold>Figure <xref ref-type="fig" rid="F2">2</xref></bold>).</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>The <italic>Solanaceae</italic> family subtree in the tree containing 1,719 plant chloroplast genome sequences. The genome name contains the NCBI GenBank name and the accession number. The sequences <italic>Solanum lycopersicum</italic> species sequences are highlighted in blue and other sequences (non-target species) are indicated in black. The neighbor-joining tree was constructed using MEGA version 6 (<xref ref-type="bibr" rid="B54">Tamura et al., 2013</xref>) and is based on distance matrix created with andi (<xref ref-type="bibr" rid="B17">Haubold et al., 2015</xref>).</p></caption>
<graphic xlink:href="fpls-09-00006-g002.tif"/>
</fig>
<p>Using our pipeline for the identification of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers from assembled chloroplast sequences in five <italic>S. lycopersicum</italic> as the target taxon sequences and 1,714 other plant species as the non-target taxa sequences, we detected 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers (32 nucleotides in length). The tomato plant chloroplast genome is approximately 155,000 bp and there were 129,812 <italic>k</italic>-mers that were present in at least two chloroplast genome sequences from <italic>S. lycopersicum</italic>. After removing the <italic>k</italic>-mers that were not present in at least 2 <italic>S. lycopersicum</italic> sequences (individual-specific <italic>k</italic>-mers) and the unspecific <italic>k</italic>-mers that were present in the assembled chloroplast genome sequences of non-target species or whole genome sequencing raw reads for <italic>S. pimpinellifolium</italic> and <italic>S. tuberosum</italic>, we detected 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers (<bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold>). We used <italic>k</italic>-mer length 32 nt, which gave us the maximum number of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers (<bold>Figure <xref ref-type="fig" rid="F4">4</xref></bold>) for the identification of <italic>S. lycopersicum</italic> in sequencing raw reads.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>The number of <italic>S. lycopersicum k</italic>-mers in every step of the pipeline. Each FASTA file contains assembled chloroplast genome sequences from different <italic>S. lycopersicum</italic> samples. There were 129,799 &#x2013; 129,840 unique <italic>k</italic>-mers (32 nucleotides in length) in each chloroplast genome sequence, 130,110 unique <italic>k</italic>-mers in the union of <italic>k</italic>-mers from all the sequences and 129,812 <italic>k</italic>-mers that were present in at least 2 <italic>S. lycopersicum</italic> chloroplast genome sequences. After removing the <italic>k</italic>-mers that were not present in at least two target sequences and unspecific <italic>k</italic>-mers that were present in assembled chloroplast genome sequences of non-target taxa or whole genome sequencing raw reads from <italic>Solanum pimpinellifolium</italic> and <italic>Solanum tuberosum</italic>, we detected 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers.</p></caption>
<graphic xlink:href="fpls-09-00006-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p>The number of identified <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers with different lengths (8, 12, 16, 20, 24, 28, and 32 nucleotides).</p></caption>
<graphic xlink:href="fpls-09-00006-g004.tif"/>
</fig>
</sec>
<sec><title>The Number of <italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers Detected from Genomic Reads</title>
<p>To test whether <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers (detected from plastid genome sequences) are detectable also in whole genome sequencing raw data, we analyzed the number of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers detected in whole genome sequencing raw reads from different <italic>Solanaceae</italic> species (<italic>S. lycopersicum</italic>, <italic>S. pimpinellifolium, S. tuberosum</italic>, <italic>S. melongena</italic>, and <italic>Capsicum annuum</italic>). The reads were downloaded from NCBI SRA database (details in the see section &#x201C;Materials and Methods&#x201D;). To also consider the influence of the number of sequencing reads on <italic>k</italic>-mers detection, we simulated new FASTQ files with different numbers of sequencing reads (10<sup>3</sup>&#x2013;10<sup>8</sup>) from original FASTQ files for all the samples.</p>
<p>The results showed that <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers, that originated from the chloroplast genome, are also detectable in whole genome sequencing raw reads from <italic>S. lycopersicum.</italic> The number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers from <italic>S. lycopersicum</italic> sequencing data increased with the number of sequencing reads. Almost all 882 <italic>k</italic>-mers from the pre-selected set of taxon-specific <italic>k</italic>-mers were detected, when the number of sequencing reads was at least 10<sup>5</sup> (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Assuming that the sequencing coverage of different genomic regions is almost equal, 10<sup>5</sup> sequencing reads provides coverage for <italic>S. lycopersicum</italic> nuclear genome of approximately 0,01&#x00D7;. The chloroplast genome coverage is difficult to estimate because of the variable copy number of plastome in individual plant cells.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>The number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in the whole genome sequencing raw data from <italic>S. lycopersicum, S. pimpinellifolium</italic>,<italic>S. tuberosum, Solanum melongena</italic>, and <italic>Capsicum annuum</italic> with variable number of sequencing reads (10<sup>2</sup>&#x2013;10<sup>8</sup>). The set of taxon-specific <italic>k</italic>-mers contains 882 <italic>k</italic>-mers that were present in at least 2 <italic>S. lycopersicum</italic> chloroplast sequence. The samples from the <italic>S. lycopersicum</italic> target species are marked with a red color, and the non-target species are in different colors.</p></caption>
<graphic xlink:href="fpls-09-00006-g005.tif"/>
</fig>
<p>In addition to the <italic>S. lycopersicum</italic> specific <italic>k</italic>-mer list containing <italic>k</italic>-mers that were present in at least 2 <italic>S. lycopersicum</italic> chloroplast genome sequences, we performed similar experiments with the <italic>k</italic>-mer lists that contained <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers that were present in at least 1 or at least 5 (all) <italic>S. lycopersicum</italic> chloroplast sequences to analyze the influence of universality cut-off value on <italic>k</italic>-mers detectability. The results for the different sets were similar (Supplementary Figure <xref ref-type="supplementary-material" rid="SM2">1</xref>).</p>
<p>To analyze the specificity of the <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers, we analyzed the number of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in the whole genome sequencing raw reads from phylogenetically close non-target taxa (<italic>S. pimpinellifolium</italic>, <italic>S. tuberosum</italic>, <italic>S. melongena</italic>, and <italic>C. annuum</italic>). The results showed that the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in whole genome sequencing raw data from phylogenetically close species started to increase as the read number exceeded 10<sup>6</sup>, it reached to 548 in the whole sequencing raw data from <italic>S. pimpinellifolium</italic> when the sequencing read number was 10<sup>8</sup> (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Therefore, for a metagenomics sample with an unknown proportion of <italic>S. lycopersicum</italic> and other taxa (e.g., <italic>S. pimpinellifolium</italic> or <italic>S. tuberosum</italic>), detecting less than 550 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers will generate complications for determine if these <italic>k</italic>-mers originate from <italic>S. lycopersicum</italic> or from other <italic>Solanaceae</italic> species.</p>
<p>We also used our method to identify species-specific <italic>k</italic>-mers for <italic>Oryza sativa</italic> and <italic>Zea mays</italic> and analyzed the number of <italic>O. sativa</italic> and <italic>Z. mays</italic> specific <italic>k</italic>-mers detected in whole genome sequencing raw reads from different samples of target taxa and also phylogenetically close non-target taxa (Supplementary Figure <xref ref-type="supplementary-material" rid="SM3">2</xref>).</p>
</sec>
<sec><title>Using the Frequency of <italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers Detected in Genomic Reads to Increase Specificity</title>
<p>In the detection of taxon-specific <italic>k</italic>-mers from whole genome sequencing reads, in addition to the detection of the presence or absence of these <italic>k</italic>-mers (binary YES/NO detection, <italic>k</italic>-mer was detected if it was represented in the sample with frequency at least 1), it would be helpful, if we could also consider the frequency of every <italic>k</italic>-mer in the sequencing reads. Although, the copy numbers of chloroplast genome may vary in different samples, we assumed that all the <italic>k</italic>-mers from the <italic>S. lycopersicum</italic> chloroplast genomes are represented in raw sequencing reads from one sample with similar frequency, which reflects the copy numbers of chloroplast genomes in the sample. This could help to distinguish the <italic>k</italic>-mers truly from <italic>S. lycopersicum</italic> (should be with similar frequency) and <italic>k</italic>-mers from close non-target species (caused by sequencing errors etc.).</p>
<p>Assuming that <italic>k</italic>-mers with very low frequency can be the result of sequencing errors, we tried to increase the <italic>k</italic>-mer frequency cut-off value (from 1 to 10) to increase the specificity. The increased cut-off value decreased the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in <italic>S. tuberosum</italic> and <italic>S. pimpinellifolium</italic>, even when the read number was more than 10<sup>7</sup> (specificity increased). However, at in the same time we were not able to detect <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in some of the <italic>S. lycopersicum</italic> samples when the sequencing read number was 2.5<sup>&#x2217;</sup>10<sup>5</sup> or less (sensitivity of <italic>S. lycopersicum</italic> detection decreased) (<bold>Figure <xref ref-type="fig" rid="F6">6</xref></bold>). Therefore, it is possible to increase the specificity but with the price of sensitivity. Thus the appropriate choice of the frequency cut-off value depends on the specific case.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>The numbers of detected <italic>S. lycopersicum</italic> specific k-mers with frequencies of at least <bold>(A)</bold> 1, <bold>(B)</bold> 2, <bold>(C)</bold> 5, and <bold>(D)</bold> 10 in the whole genome sequencing raw data from <italic>S. lycopersicum, S. pimpinellifolium, S. tuberosum, S. melongena</italic>, and <italic>C. annuum</italic> with variable numbers of sequencing reads (10<sup>3</sup>&#x2013;10<sup>8</sup>). The detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers were present in at least 2 <italic>S. lycopersicum</italic> chloroplast sequences.</p></caption>
<graphic xlink:href="fpls-09-00006-g006.tif"/>
</fig>
</sec>
<sec><title>The Location of the <italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers in the Chloroplast Genome</title>
<p>The locations of all 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers identified in this work using our pipeline are described (in <bold>Figure <xref ref-type="fig" rid="F7">7</xref></bold>). The <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers were located along the entire chloroplast genome as 42 clustered taxon-specific regions along the single copy regions in the <italic>S. lycopersicum</italic> chloroplast genome (most of them were located in intergenic spacer regions or introns). The <italic>k</italic>-mers that were located adjacent to each other and differ by only one nucleotide belonged in the same cluster. The sequences in the <italic>k</italic>-mers clusters contain at least 1 mismatch per 32 bp of <italic>S. lycopersicum</italic> sequence compared to all non-target sequences.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>The location of the <italic>S. lycopersicum</italic> specific k-mer clusters in the chloroplast genome (in sequence NC_007898.3). The outer circle illustrates the location of the chloroplast genes and the middle circle indicates the location of the inverted repeats (IR<sub>A</sub> and IR<sub>B</sub>) and single copy regions (SSC and LSC). The location of the 42 different clusters of <italic>S. lycopersicum</italic> specific k-mers are shown as small black strokes in the inner circle in the figure.</p></caption>
<graphic xlink:href="fpls-09-00006-g007.tif"/>
</fig>
<p>The lengths of the 42 detected taxon-specific regions in the <italic>S. lycopersicum</italic> chloroplast genome varied from 33 to 62 bp (average 53.3 bp, median 57.5 bp), meaning that there was 2&#x2013;31 <italic>k</italic>-mers in one cluster. The sequence lengths bp) between two adjacent cluster regions (region between the end of the last <italic>k</italic>-mer from one cluster and the start of the first <italic>k</italic>-mer sequence from the next cluster) varied from 30 to 39,507 bp.</p>
</sec>
</sec>
<sec><title>Discussion</title>
<p>Current developments in the identification plant taxa in degraded metagenomics samples are moving toward using a combination of many different DNA barcodes with reduced lengths (e.g., mini-barcodes) (<xref ref-type="bibr" rid="B8">Dong et al., 2014</xref>; <xref ref-type="bibr" rid="B51">Shokralla et al., 2015</xref>; <xref ref-type="bibr" rid="B57">Wang et al., 2016</xref>) and deep sequencing of total genomic DNA from samples with various taxon composition, followed by the identification of the taxonomical origin of the sequencing reads (<xref ref-type="bibr" rid="B47">Ripp et al., 2014</xref>). <italic>k</italic>-mer based methods, that do not require prior primer design and the pre-amplification of specific regions, genome assembly and mapping of sequencing reads to a reference genome, have already been applied for the identification of microbial species or strains in metagenomics samples (<xref ref-type="bibr" rid="B61">Wood and Salzberg, 2014</xref>; <xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>; <xref ref-type="bibr" rid="B24">Kim et al., 2016</xref>; <xref ref-type="bibr" rid="B48">Roosaare et al., 2017</xref>).</p>
<p>Here we introduce a <italic>k</italic>-mers-based approach to detect plant taxa from raw sequencing reads. We showed that it is possible to identify a set of plant taxon-specific <italic>k</italic>-mers from the chloroplast genome, that are also detectable from whole genome sequencing raw reads of target taxon. Different pieces of available softwares enabled fast identification and counting of taxon-specific <italic>k</italic>-mers from genome sequences from different taxa (e.g., <xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>).</p>
<p>To find a correct set of taxon-specific <italic>k</italic>-mers, it is recommended to consider also taxonomic ambiguity and use sequences from a reliable database, where the reference specimen has been correctly identified, because several gaps and false sequences could influence the results. The availability of many reliable genome sequences for a target taxon as well as many sequences for phylogenetically close non-target taxa are good presumptions for finding target taxon-specific <italic>k</italic>-mers. The number of available plant whole genome sequences is still not sufficient for finding taxon-specific <italic>k</italic>-mers, but there are many advantages to use chloroplast genome sequences for finding taxon-specific <italic>k</italic>-mers for detection of pre-selected plant taxon from metagenomics sample (e.g., high copy number, interspecific variability, low intraspecific variability, consistently increasing number of available sequences). To get overview of the available data of target and non-target sequences we constructed a tree containing more than 1,700 chloroplast genome sequences.</p>
<p>Although the variability of chloroplast genome sequences between very close plant species can be quite low and finding barcoding regions for amplification has been restricted (<xref ref-type="bibr" rid="B12">Fazekas et al., 2008</xref>; <xref ref-type="bibr" rid="B18">Hollingsworth et al., 2011</xref>), we identified a set of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers from the chloroplast genome. The set of taxon-specific <italic>k</italic>-mers can be identified from the chloroplast genome using available software in only a few minutes, and the set of taxon-specific <italic>k</italic>-mers can be easily updated if additional genomic sequences are available in biological databases for either target or non-target species.</p>
<p>Using assembled chloroplast sequences from the GenBank database, we detected 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers that were 32 nucleotides in length (excluding individual-specific <italic>k</italic>-mers). Our results showed that the identified <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers were located in 42 different regions along the <italic>S. lycopersicum</italic> chloroplast genome single copy regions, which is in accordance with the previous findings that the mutation rate in the plastome single copy regions is much higher than in the IR regions (<xref ref-type="bibr" rid="B60">Wolfe et al., 1987</xref>; <xref ref-type="bibr" rid="B31">Maier et al., 1995</xref>; <xref ref-type="bibr" rid="B21">Kahlau et al., 2006</xref>). None of the <italic>k</italic>-mers were in <italic>S. lycopersicum</italic> chloroplast genome IR regions, which is probably due to the low mutation rate and therefore very few differences between the sequences in the IR regions of very close plant species. The detected <italic>k</italic>-mers were most frequently in intergenic spacer regions or introns. Previous studies have identified the following thirteen most variable plastid marker regions in <italic>Solanum</italic> species: atpB-rbcL, clpP-psbB, ndhF, ndhF-rpl32, petL-psaJ (including petL-petG-trnW and trnP-psaJ), petN-psbM, rpl32-trnL, rpoC1-rpoB, trnA-trnI, trnK-rps16, and ycf1 (parts 1&#x2013;3) (<xref ref-type="bibr" rid="B50">S&#x00E4;rkinen and George, 2013</xref>). Most of these are also regions that contain <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers identified in our study.</p>
<p>The detection of pre-selected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in the whole genome sequencing raw reads from <italic>S. lycopersicum</italic> showed that <italic>k</italic>-mers from chloroplast genomes are also detectable from whole genome sequencing raw reads, and the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers increases with the increase in the number of sequencing reads. Almost all 882 k-mers were detected when read number was at least 100,000.</p>
<p>Most of the <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers differed from <italic>S. pimpinellifolium</italic> or <italic>S. tuberosum</italic> chloroplast genome sequences by only one nucleotide. Therefore, probably due to sequencing errors (<xref ref-type="bibr" rid="B35">Nakamura et al., 2011</xref>; <xref ref-type="bibr" rid="B30">Liu et al., 2012</xref>), the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers also increased in the samples from other <italic>Solanaceae</italic> species when the read number was greater than 1,000,000. This could be a particular challenge when detecting tomato plant metagenomics samples containing large amounts of sequencing reads from phylogenetically close non-target taxa (like currant tomato or potato) and small amount of reads from <italic>S. lycopersicum</italic>. However, it would be possible to detect <italic>S. lycopersicum</italic> from raw sequencing reads using taxon-specific <italic>k</italic>-mers from the chloroplast genome, even from metagenomics samples containing tomato, currant tomato, potato, eggplant, and pepper, if the read number from <italic>S. lycopersicum</italic> is at least 100,000 and approximately 600 <italic>k</italic>-mers (from 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers) are detected.</p>
<p>It is difficult to take into account the frequency of detected <italic>k</italic>-mers to increase the specificity or to estimate the amount of <italic>S. lycopersicum</italic> in a sample, as the frequency of <italic>k</italic>-mers from the chloroplast genome is influenced on many factors (organization of chloroplast genome, the copy number of chloroplasts and chloroplast genomes in different cells, in different parts of plants, depending on plant species, growth conditions, etc.) (<xref ref-type="bibr" rid="B40">Oldenburg and Bendich, 2004</xref>; <xref ref-type="bibr" rid="B29">Liere and B&#x00F6;rner, 2013</xref>).</p>
<p>To find ways to increase the detection specificity while also considering the frequency of detected <italic>k</italic>-mers, we analyzed the frequency distribution of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in different <italic>S. lycopersicum</italic> and other <italic>Solanaceae</italic> species samples. Although the average frequency varied between different samples, the overall frequency of detected <italic>S. lycopersicum k</italic>-mers in <italic>S. lycopersicum</italic> samples was substantially higher than in other (non-target) <italic>Solanaceae</italic> samples. Using increased <italic>k</italic>-mer frequency cut-off values increased the specificity (the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in <italic>S. tuberosum</italic> and <italic>S. pimpinellifolium</italic> was much lower, even when the read number was more than 10<sup>7</sup>), though the sensitivity of <italic>S. lycopersicum</italic> detection was decreased (we were not able to detect <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in some of the <italic>S. lycopersicum</italic> samples if the sequencing read number was 2.5<sup>&#x2217;</sup>10<sup>5</sup> or less). Therefore, the appropriate choice for the frequency cut-off value depends on the specific case.</p>
<p>The increasing number of available plant whole genome sequences probably in the future gives opportunity to identify additional plant taxa specific <italic>k</italic>-mers from nuclear genomes which would be useful, when discriminating very close species, subspecies or cultivars with very little or lack of chloroplast genome sequence variability.</p>
<p>Using hundreds of taxon-specific short <italic>k</italic>-mers from all over the genome for the identification or qualitative detection of plant taxa would give improved resolution at the species level as well as aid in analyzing complex and degraded samples that sequence barcoding or other traditional methods fail to resolve. However, to apply the identified <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers to detect <italic>S. lycopersicum</italic> (tomato plant) from real metagenomics samples (food, medicines, environmental samples, etc.), additional testing with real metagenomics samples is required.</p>
</sec>
<sec><title>Conclusion</title>
<p>Accurate and fast methods for the identification of plant(s) from degraded metagenomics samples play an important role in identifying the composition of complex mixtures of processed biological materials, including food, herbal products, gut contents, environmental samples, etc., PCR and different barcoding methods commonly used for plant identification are based on the amplification of a limited number of pre-selected barcoding regions. These methods are often inapplicable due to the degree of DNA degradation, low PCR amplification success or low species discriminative power of the selected barcoding regions.</p>
<p>Our method enables rapid identification of plant taxon-specific <italic>k</italic>-mers from the chloroplast genome. By applying our method to select a set of <italic>S. lycopersicum</italic> (tomato plant) specific <italic>k</italic>-mers, we identified 882 <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers that were present in at least two chloroplast genome sequences from <italic>S. lycopersicum</italic> and none of the 1,714 chloroplast sequences from non-target species. <italic>In silico</italic> experiments using raw sequencing data from the <italic>S. lycopersicum, S. tuberosum</italic> (potato), <italic>S. pimpinellifolium</italic> (currant tomato), <italic>S. melongena</italic> (eggplant), and <italic>C. annuum</italic> (pepper) whole genomes showed that <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers found in the chloroplast genome can also be detected from <italic>S. lycopersicum</italic> whole genome sequencing raw reads as well as in other <italic>Solanaceae</italic> species. If sequencing data from metagenomic samples contain at least 10<sup>5</sup> reads from the <italic>S. lycopersicum</italic> genome, it is possible to discriminate <italic>k</italic>-mers from <italic>S. lycopersicum</italic> and <italic>k</italic>-mers from other <italic>Solanaceae</italic> species.</p>
<p>We are providing a valuable method here for the identification of plant taxon-specific <italic>k</italic>-mers that could be used in the future for developing diagnostic tests for the fast detection of different plant taxa in raw sequencing reads from metagenomics samples.</p>
</sec>
<sec id="s1" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec><title>Identification of Taxon-Specific <italic>k</italic>-mers</title>
<p>By combining different programs (GListMaker, GListCompare, and GListQuery) from the GenomeTester4 software package (<xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>) and in-house scripts, we constructed <italic>k</italic>-mer analysis pipeline for finding a set of taxon-specific <italic>k</italic>-mers. GListMaker converts assembled chloroplast sequences or raw sequencing reads (FASTA or FASTQ) from the sample to <italic>k</italic>-mer lists and GListCompare compares different lists of <italic>k</italic>-mers to find intersections, differences or unions between the lists for detecting a set of taxon-specific <italic>k</italic>-mers.</p>
</sec>
<sec><title>Identification of <italic>Solanum lycopersicum</italic> (Tomato Plant) Specific <italic>k</italic>-mers</title>
<p>To get an overview of the available data from assembled chloroplast genome sequences, we constructed a neighbor-joining tree consisting of 1,719 chloroplast genome sequences from plants and some green algae including 17 sequences from <italic>Solanum</italic> species, and five sequences from <italic>S. lycopersicum</italic>, which were downloaded from the GenBank database (<xref ref-type="bibr" rid="B4">Clark et al., 2016</xref>) (the accession numbers of complete chloroplast genome sequences we used to find <italic>k</italic>-mers are in Supplementary Data <xref ref-type="supplementary-material" rid="SM1">1</xref>). The tree was constructed using the computer program MEGA version 6 (<xref ref-type="bibr" rid="B54">Tamura et al., 2013</xref>) and was based on a distance matrix constructed using the computer program andi (<xref ref-type="bibr" rid="B17">Haubold et al., 2015</xref>) to get an overview of the available data. The phylogenetic tree of <italic>Solanaceae</italic> is a subtree from the constructed tree.</p>
<p>To identify <italic>S. lycopersicum</italic> specific k-mers, we used different GenomeTester4 programs and in-house scripts (pipeline is available in the public repository Github). We created a <italic>k</italic>-mer list for <italic>S. lycopersicum</italic> and a <italic>k</italic>-mer list for other (non-target) species. The final set of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers containing 882 <italic>k</italic>-mers was identified using two steps for removing non-specific <italic>k</italic>-mers from the <italic>S. lycopersicum k</italic>-mer list. The first step used assembled plastome sequences from non-target taxa to remove non-specific <italic>k</italic>-mers that were represented in the chloroplast genome of non-target taxa. The second step used whole genome sequencing raw data from the two phylogenetically close non-target species <italic>S. pimpinellifolium</italic> and <italic>S. tuberosum</italic> with a frequency at least 10 to remove additional non-specific <italic>k</italic>-mers that may be represented in the nuclear or mitochondrial genomes of non-target taxa <italic>S. pimpinellifolium</italic> and <italic>S. tuberosum</italic>. A cut-off value of 10 removed most of the non-specific <italic>k</italic>-mers present in the nuclear, mitochondrial or chloroplast genome regions of the non-target taxa, but did not include <italic>k</italic>-mers caused by sequencing errors in the sequencing reads. The raw sequencing reads from 1 <italic>S. pimpinellifolium</italic> and 3 <italic>S. tuberosum</italic> (ERR418080, SRR1608100, SRR2069941, and SRR1481624) were downloaded from the NCBI SRA database (<xref ref-type="bibr" rid="B28">Leinonen et al., 2011</xref>).</p>
<p>In our study, we made analyses with different <italic>k</italic>-mer lengths to select the <italic>k</italic> value that gives the maximum number <italic>of S. lycopersicum</italic> specific <italic>k</italic>-mers, that can next be used for the identification of <italic>S. lycopersicum</italic> in sequencing raw reads. We used <italic>k</italic> values 8, 12, 16, 20, 24, 28, and 32 for finding taxon-specific <italic>k</italic>-mers for <italic>S. lycopersicum</italic>. The appropriate length of the <italic>k</italic>-mer depends on the taxon or the taxonomic level we were trying to find specific <italic>k</italic>-mers for and on the variability of the used genomic sequences. Smaller <italic>k</italic>-mer length leads to an increased number of <italic>k</italic>-mers that are represented in all sequences from a target taxon, but to the lower specificity of those <italic>k</italic>-mers (<xref ref-type="bibr" rid="B41">Ounit et al., 2015</xref>). Working with <italic>k</italic>-mers up to <italic>k</italic> = 32 is more efficient on 64-bit computers than using longer <italic>k</italic>-mers. GenomeTester4 programs used in our pipeline are summarizing the frequencies of the <italic>k</italic>-mer and its reverse complement <italic>k</italic>-mer when creating <italic>k</italic>-mers&#x2019; lists, i.e., one <italic>k</italic>-mer in a stored list represent both a <italic>k</italic>-mer and its reverse complement (<xref ref-type="bibr" rid="B23">Kaplinski et al., 2015</xref>). Therefore, it is not important to distinguish <italic>k</italic>-mers with odd and even length or prefer only odd <italic>k</italic>-mers.</p>
<p>We created the following three lists of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers: (1) <italic>k</italic>-mers that were present in all five <italic>S. lycopersicum</italic> chloroplast sequences, (2) <italic>k</italic>-mers that were present at least in 2 <italic>S. lycopersicum</italic> chloroplast sequences, and (3) <italic>k</italic>-mers that were present at least 1 <italic>S. lycopersicum</italic> chloroplast sequence. We selected one of these lists of <italic>k</italic>-mers for the subsequent analysis.</p>
</sec>
<sec><title><italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers in Whole Genome Sequencing Raw Reads from Different <italic>Solanaceae</italic> Species</title>
<p>We used gmer_counter (<xref ref-type="bibr" rid="B42">Pajuste et al., 2017</xref>) and different Python scripts to detect of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in the raw sequencing reads from whole genome samples of the following different <italic>Solanaceae</italic> species: (1) 16 <italic>S. lycopersicum</italic> (tomato), (2) 10 <italic>S. tuberosum</italic> (potato), (3) 4 <italic>C. annuum</italic> (pepper), (4) 2 <italic>S. pimpinellifolium</italic> (currant tomato), and (5) 1 <italic>S. melongena</italic> (eggplant) (ERR539598, SRR1572263, ERR418062, ERR418047, DRR000741, DRR040154, DRR040149, DRR040095, ERR418079, ERR418069, SRR1572661, SRR1572551, SRR404081, DRR022703, DRR022708, ERR964441, SRR307673, SRR2069932, SRR2070065, SRR1501269, SRR2069942, SRR1608091, SRR1 607674, SRR1525230, SRR1594256, ERR023052, SRR2752033, SRR2752003, SRR653457, SRR653499, ERR418082, SRR074949, and DRR014074). All the datasets downloaded from NCBI SRA database (<xref ref-type="bibr" rid="B28">Leinonen et al., 2011</xref>) contained whole genome sequencing raw reads. The sequencing read lengths in these datasets were from about 80 to 500 bp, but predominantly 200 bp, DNA was extracted mostly from the plant leaves or from tuber (potato).</p>
<p>To analyze the relationship between the number of detected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers and the number of next-generation sequencing reads, we simulated new FASTQ files with different number of reads (10<sup>3</sup>, 10<sup>4</sup>, 10<sup>5</sup>, 2.5<sup>&#x2217;</sup>10<sup>5</sup>, 5<sup>&#x2217;</sup>10<sup>5</sup>, 10<sup>6</sup>, 10<sup>7</sup>, and 10<sup>8</sup>) from the raw sequencing data files for all the samples. The reads were evenly selected from an original FASTQ file [downloaded from NCBI SRA database (<xref ref-type="bibr" rid="B28">Leinonen et al., 2011</xref>), i.e., when creating the new FASTQ file with 100,000 reads, every 1000. read from the original FASTQ file with 100,000,000 reads goes to the new FASTQ filean in-house developed Python script was used for this]. Therefore, different new FASTQ files for the same sample with different numbers of sequencing reads may not contain exactly the same set of reads.</p>
<p>We used four different cut-off values (1, 2, 5, and 10) for the frequency of the detected <italic>k</italic>-mers to analyse the influence of increased cut-off values on the specificity and sensitivity of <italic>S. lycopersicum</italic> detection from raw sequencing reads using pre-selected <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers.</p>
<p>We also analyzed the frequency of individual <italic>k</italic>-mers from the final set of <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers in raw sequencing reads from 16 different <italic>S. lycopersicum</italic> and 17 other plant species whole genomes.</p>
</sec>
<sec><title>The Location of <italic>Solanum lycopersicum</italic> Specific <italic>k</italic>-mers in the Chloroplast Genome</title>
<p>The locations of the taxon-specific <italic>k</italic>-mers in the chloroplast genome were found using the <italic>S. lycopersicum</italic> chloroplast sequence NC_007898.3 using Python scripts. A physical map of the tomato plant plastome was drawn using the GenomeVX software (<xref ref-type="bibr" rid="B7">Conant and Wolfe, 2008</xref>).</p>
</sec>
</sec>
<sec><title>Availability of Data and Material</title>
<p>The datasets analyzed in the current study are available in the NCBI SRA (Sequence Read Archive), <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/sra">https://www.ncbi.nlm.nih.gov/sra</ext-link>, and in GenBank, <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/Genbank">https://www.ncbi.nlm.nih.gov/Genbank</ext-link>. The full list of accession numbers for the used sequences are given in Supplementary Data <xref ref-type="supplementary-material" rid="SM1">1</xref>. The sequences of the <italic>S. lycopersicum</italic> specific <italic>k</italic>-mers identified in the current study are available from the corresponding author upon reasonable request. The pipeline, in-house scripts, including used parameters, are available in the public repository Github: <ext-link ext-link-type="uri" xlink:href="https://github.com/bioinfo-ut/PlantTaxSeeker">https://github.com/bioinfo-ut/PlantTaxSeeker</ext-link>.</p>
</sec>
<sec><title>Author Contributions</title>
<p>KR constructed the <italic>k</italic>-mer detection pipelines, wrote in-house Python scripts, analyzed the sequence data and wrote the manuscript. KR and MR designed the experiments and interpreted the results of all analysis. MR edited the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was funded by institutional grant IUT34-11 from the Estonian Ministry of Education and Research and the EU ERDF grant No. 2014-2020.4.01.15-0012 (Estonian Center of Excellence in Genomics and Translational Medicine).</p>
</fn>
</fn-group>
<ack>
<p>The authors are grateful to Triinu K&#x00F5;ressaar, Joachim Mattias Gerhold, and Reedik M&#x00E4;gi for their invaluable suggestions toward improvement of the manuscript.</p>
</ack>
<sec sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2018.00006/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpls.2018.00006/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.XLSX" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
<supplementary-material xlink:href="Supplementary_Figure_1.docx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
<supplementary-material xlink:href="Supplementary_Figure_2.DOCX" id="SM3" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document" xmlns:xlink="http://www.w3.org/1999/xlink">
</supplementary-material>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><collab>Cbol Plant Working</collab>. (<year>2009</year>). <article-title>A DNA barcode for land plants.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>106</volume> <fpage>12794</fpage>&#x2013;<lpage>12797</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.0905845106</pub-id> <pub-id pub-id-type="pmid">19666622</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cho</surname> <given-names>K.-S.</given-names></name> <name><surname>Park</surname> <given-names>T.-H.</given-names></name></person-group> (<year>2016</year>). <article-title>Complete chloroplast genome sequence of <italic>Solanum nigrum</italic> and development of markers for the discrimination of <italic>S. nigrum</italic>.</article-title> <source><italic>Hortic. Environ. Biotechnol.</italic></source> <volume>57</volume> <fpage>69</fpage>&#x2013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1007/s13580-016-0003-2</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chung</surname> <given-names>H.-J.</given-names></name> <name><surname>Jung</surname> <given-names>J. D.</given-names></name> <name><surname>Park</surname> <given-names>H.-W.</given-names></name> <name><surname>Kim</surname> <given-names>J.-H.</given-names></name> <name><surname>Cha</surname> <given-names>H. W.</given-names></name> <name><surname>Min</surname> <given-names>S. R.</given-names></name><etal/></person-group> (<year>2006</year>). <article-title>The complete chloroplast genome sequences of <italic>Solanum tuberosum</italic> and comparative analysis with Solanaceae species identified the presence of a 241-bp deletion in cultivated potato chloroplast DNA sequence.</article-title> <source><italic>Plant Cell Rep.</italic></source> <volume>25</volume> <fpage>1369</fpage>&#x2013;<lpage>1379</lpage>. <pub-id pub-id-type="doi">10.1007/s00299-006-0196-4</pub-id> <pub-id pub-id-type="pmid">16835751</pub-id></citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clark</surname> <given-names>K.</given-names></name> <name><surname>Karsch-Mizrachi</surname> <given-names>I.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name> <name><surname>Ostell</surname> <given-names>J.</given-names></name> <name><surname>Sayers</surname> <given-names>E. W.</given-names></name></person-group> (<year>2016</year>). <article-title>GenBank.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>44</volume> <fpage>D67</fpage>&#x2013;<lpage>D72</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkv1276</pub-id> <pub-id pub-id-type="pmid">26590407</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Clement</surname> <given-names>W. L.</given-names></name> <name><surname>Donoghue</surname> <given-names>M. J.</given-names></name></person-group> (<year>2012</year>). <article-title>Barcoding success as a function of phylogenetic relatedness in <italic>Viburnum</italic>, a clade of woody angiosperms.</article-title> <source><italic>BMC Evol. Biol.</italic></source> <volume>12</volume>:<issue>73</issue>. <pub-id pub-id-type="doi">10.1186/1471-2148-12-73</pub-id> <pub-id pub-id-type="pmid">22646220</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Coghlan</surname> <given-names>M. L.</given-names></name> <name><surname>Haile</surname> <given-names>J.</given-names></name> <name><surname>Houston</surname> <given-names>J.</given-names></name> <name><surname>Murray</surname> <given-names>D. C.</given-names></name> <name><surname>White</surname> <given-names>N. E.</given-names></name> <name><surname>Moolhuijzen</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Deep sequencing of plant and animal DNA contained within traditional Chinese medicines reveals legality issues and health safety concerns.</article-title> <source><italic>PLOS Genet.</italic></source> <volume>8</volume>:<issue>e1002657</issue>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1002657</pub-id> <pub-id pub-id-type="pmid">22511890</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conant</surname> <given-names>G. C.</given-names></name> <name><surname>Wolfe</surname> <given-names>K. H.</given-names></name></person-group> (<year>2008</year>). <article-title>GenomeVx: simple web-based creation of editable circular chromosome maps.</article-title> <source><italic>Bioinformatics</italic></source> <volume>24</volume> <fpage>861</fpage>&#x2013;<lpage>862</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btm598</pub-id> <pub-id pub-id-type="pmid">18227121</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>C.</given-names></name> <name><surname>Zuo</surname> <given-names>Y.</given-names></name> <name><surname>Chen</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>S.</given-names></name></person-group> (<year>2014</year>). <article-title>A chloroplast genomic strategy for designing taxon specific DNA mini-barcodes: a case study on ginsengs.</article-title> <source><italic>BMC Genet.</italic></source> <volume>15</volume>:<issue>138</issue>. <pub-id pub-id-type="doi">10.1186/s12863-014-0138-z</pub-id> <pub-id pub-id-type="pmid">25526752</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>W.</given-names></name> <name><surname>Liu</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>J.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>S.</given-names></name></person-group> (<year>2012</year>). <article-title>Highly variable chloroplast markers for evaluating plant phylogeny at low taxonomic levels and for DNA barcoding.</article-title> <source><italic>PLOS ONE</italic></source> <volume>7</volume>:<issue>e35071</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0035071</pub-id> <pub-id pub-id-type="pmid">22511980</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Drouin</surname> <given-names>G.</given-names></name> <name><surname>Daoud</surname> <given-names>H.</given-names></name> <name><surname>Xia</surname> <given-names>J.</given-names></name></person-group> (<year>2008</year>). <article-title>Relative rates of synonymous substitutions in the mitochondrial, chloroplast and nuclear genomes of seed plants.</article-title> <source><italic>Mol. Phylogenet. Evol.</italic></source> <volume>49</volume> <fpage>827</fpage>&#x2013;<lpage>831</lpage>. <pub-id pub-id-type="doi">10.1016/j.ympev.2008.09.009</pub-id> <pub-id pub-id-type="pmid">18838124</pub-id></citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Elbrecht</surname> <given-names>V.</given-names></name> <name><surname>Leese</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>Can DNA-based ecosystem assessments quantify species abundance? Testing primer bias and biomass&#x2014;sequence relationships with an innovative metabarcoding protocol.</article-title> <source><italic>PLOS ONE</italic></source> <volume>10</volume>:<issue>e0130324</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0130324</pub-id> <pub-id pub-id-type="pmid">26154168</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fazekas</surname> <given-names>A. J.</given-names></name> <name><surname>Burgess</surname> <given-names>K. S.</given-names></name> <name><surname>Kesanakurti</surname> <given-names>P. R.</given-names></name> <name><surname>Graham</surname> <given-names>S. W.</given-names></name> <name><surname>Newmaster</surname> <given-names>S. G.</given-names></name> <name><surname>Husband</surname> <given-names>B. C.</given-names></name><etal/></person-group> (<year>2008</year>). <article-title>Multiple multilocus DNA barcodes from the plastid genome discriminate plant species equally well.</article-title> <source><italic>PLOS ONE</italic></source> <volume>3</volume>:<issue>e2802</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0002802</pub-id> <pub-id pub-id-type="pmid">18665273</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ferri</surname> <given-names>E.</given-names></name> <name><surname>Galimberti</surname> <given-names>A.</given-names></name> <name><surname>Casiraghi</surname> <given-names>M.</given-names></name> <name><surname>Airoldi</surname> <given-names>C.</given-names></name> <name><surname>Ciaramelli</surname> <given-names>C.</given-names></name> <name><surname>Palmioli</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Towards a universal approach based on omics technologies for the quality control of food.</article-title> <source><italic>BioMed. Res. Int.</italic></source> <volume>2015</volume> <fpage>1</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1155/2015/365794</pub-id> <pub-id pub-id-type="pmid">26783518</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ficetola</surname> <given-names>G. F.</given-names></name> <name><surname>Coissac</surname> <given-names>E.</given-names></name> <name><surname>Zundel</surname> <given-names>S.</given-names></name> <name><surname>Riaz</surname> <given-names>T.</given-names></name> <name><surname>Shehzad</surname> <given-names>W.</given-names></name> <name><surname>Bessi&#x00E8;re</surname> <given-names>J.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>An <italic>in silico</italic> approach for the evaluation of DNA barcodes.</article-title> <source><italic>BMC Genomics</italic></source> <volume>11</volume>:<issue>434</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-11-434</pub-id> <pub-id pub-id-type="pmid">20637073</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galimberti</surname> <given-names>A.</given-names></name> <name><surname>De Mattia</surname> <given-names>F.</given-names></name> <name><surname>Losa</surname> <given-names>A.</given-names></name> <name><surname>Bruni</surname> <given-names>I.</given-names></name> <name><surname>Federici</surname> <given-names>S.</given-names></name> <name><surname>Casiraghi</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>DNA barcoding as a new tool for food traceability.</article-title> <source><italic>Food Res. Int.</italic></source> <volume>50</volume> <fpage>55</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1016/j.foodres.2012.09.036</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gargano</surname> <given-names>D.</given-names></name> <name><surname>Scotti</surname> <given-names>N.</given-names></name> <name><surname>Vezzi</surname> <given-names>A.</given-names></name> <name><surname>Bilardi</surname> <given-names>A.</given-names></name> <name><surname>Valle</surname> <given-names>G.</given-names></name> <name><surname>Grillo</surname> <given-names>S.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Genome-wide analysis of plastome sequence variation and development of plastidial CAPS markers in common potato and related <italic>Solanum species</italic>.</article-title> <source><italic>Genet. Resour. Crop Evol.</italic></source> <volume>59</volume> <fpage>419</fpage>&#x2013;<lpage>430</lpage>.</citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Haubold</surname> <given-names>B.</given-names></name> <name><surname>Kl&#x00F6;tzl</surname> <given-names>F.</given-names></name> <name><surname>Pfaffelhuber</surname> <given-names>P.</given-names></name></person-group> (<year>2015</year>). <article-title>Andi: fast and accurate estimation of evolutionary distances between closely related genomes.</article-title> <source><italic>Bioinformatics</italic></source> <volume>31</volume> <fpage>1169</fpage>&#x2013;<lpage>1175</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu815</pub-id> <pub-id pub-id-type="pmid">25504847</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hollingsworth</surname> <given-names>P. M.</given-names></name> <name><surname>Graham</surname> <given-names>S. W.</given-names></name> <name><surname>Little</surname> <given-names>D. P.</given-names></name></person-group> (<year>2011</year>). <article-title>Choosing and using a plant DNA barcode.</article-title> <source><italic>PLOS ONE</italic></source> <volume>6</volume>:<issue>e19254</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0019254</pub-id> <pub-id pub-id-type="pmid">21637336</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Iyengar</surname> <given-names>A.</given-names></name></person-group> (<year>2014</year>). <article-title>Forensic DNA analysis for animal protection and biodiversity conservation: a review.</article-title> <source><italic>J. Nat. Conserv.</italic></source> <volume>22</volume> <fpage>195</fpage>&#x2013;<lpage>205</lpage>.</citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jansen</surname> <given-names>R. K.</given-names></name> <name><surname>Raubeson</surname> <given-names>L. A.</given-names></name> <name><surname>Boore</surname> <given-names>J. L.</given-names></name> <name><surname>dePamphilis</surname> <given-names>C. W.</given-names></name> <name><surname>Chumley</surname> <given-names>T. W.</given-names></name> <name><surname>Haberle</surname> <given-names>R. C.</given-names></name><etal/></person-group> (<year>2005</year>). <article-title>&#x201C;Methods for obtaining and analyzing whole chloroplast genome sequences,&#x201D; in</article-title> <source><italic>Methods in Enzymology</italic></source> <volume>Vol. 395</volume> <role>eds</role> <person-group person-group-type="editor"><name><surname>Zimmer</surname> <given-names>E. A.</given-names></name> <name><surname>Roalson</surname> <given-names>E. H.</given-names></name></person-group> (<publisher-loc>London</publisher-loc>: <publisher-name>Academic Press</publisher-name>), <fpage>348</fpage>&#x2013;<lpage>384</lpage>.</citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kahlau</surname> <given-names>S.</given-names></name> <name><surname>Aspinall</surname> <given-names>S.</given-names></name> <name><surname>Gray</surname> <given-names>J. C.</given-names></name> <name><surname>Bock</surname> <given-names>R.</given-names></name></person-group> (<year>2006</year>). <article-title>Sequence of the tomato chloroplast DNA and evolutionary comparison of Solanaceous plastid genomes.</article-title> <source><italic>J. Mol. Evol.</italic></source> <volume>63</volume> <fpage>194</fpage>&#x2013;<lpage>207</lpage>. <pub-id pub-id-type="pmid">16830097</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kane</surname> <given-names>N.</given-names></name> <name><surname>Sveinsson</surname> <given-names>S.</given-names></name> <name><surname>Dempewolf</surname> <given-names>H.</given-names></name> <name><surname>Yang</surname> <given-names>J. Y.</given-names></name> <name><surname>Zhang</surname> <given-names>D.</given-names></name> <name><surname>Engels</surname> <given-names>J. M. M.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Ultra-barcoding in cacao (<italic>Theobroma</italic> spp.; Malvaceae) using whole chloroplast genomes and nuclear ribosomal DNA.</article-title> <source><italic>Am. J. Bot.</italic></source> <volume>99</volume> <fpage>320</fpage>&#x2013;<lpage>329</lpage>. <pub-id pub-id-type="doi">10.3732/ajb.1100570</pub-id> <pub-id pub-id-type="pmid">22301895</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kaplinski</surname> <given-names>L.</given-names></name> <name><surname>Lepamets</surname> <given-names>M.</given-names></name> <name><surname>Remm</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>GenomeTester4: a toolkit for performing basic set operations - union, intersection and complement on k-mer lists.</article-title> <source><italic>GigaScience</italic></source> <volume>4</volume>:<issue>58</issue>. <pub-id pub-id-type="doi">10.1186/s13742-015-0097-y</pub-id> <pub-id pub-id-type="pmid">26640690</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Song</surname> <given-names>L.</given-names></name> <name><surname>Breitwieser</surname> <given-names>F. P.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2016</year>). <article-title>Centrifuge: rapid and sensitive classification of metagenomic sequences.</article-title> <source><italic>Genome Res.</italic></source> <volume>26</volume> <fpage>1721</fpage>&#x2013;<lpage>1729</lpage>. <pub-id pub-id-type="pmid">27852649</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>K.</given-names></name> <name><surname>Lee</surname> <given-names>S.-C.</given-names></name> <name><surname>Lee</surname> <given-names>J.</given-names></name> <name><surname>Lee</surname> <given-names>H. O.</given-names></name> <name><surname>Joh</surname> <given-names>H. J.</given-names></name> <name><surname>Kim</surname> <given-names>N.-H.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Comprehensive survey of genetic diversity in chloroplast genomes and 45S nrDNAs within <italic>Panax ginseng</italic> species.</article-title> <source><italic>PLOS ONE</italic></source> <volume>10</volume>:<issue>e0117159</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0117159</pub-id> <pub-id pub-id-type="pmid">26061692</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kmiec</surname> <given-names>B.</given-names></name> <name><surname>Woloszynska</surname> <given-names>M.</given-names></name> <name><surname>Janska</surname> <given-names>H.</given-names></name></person-group> (<year>2006</year>). <article-title>Heteroplasmy as a common state of mitochondrial genetic information in plants and animals.</article-title> <source><italic>Curr. Genet.</italic></source> <volume>50</volume> <fpage>149</fpage>&#x2013;<lpage>159</lpage>. <pub-id pub-id-type="doi">10.1007/s00294-006-0082-1</pub-id> <pub-id pub-id-type="pmid">16763846</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kress</surname> <given-names>W. J.</given-names></name> <name><surname>Erickson</surname> <given-names>D. L.</given-names></name></person-group> (<year>2007</year>). <article-title>A two-locus global DNA barcode for land plants: the coding <italic>rbcL</italic> gene complements the non-coding <italic>trnH-psbA</italic> spacer region.</article-title> <source><italic>PLOS ONE</italic></source> <volume>2</volume>:<issue>e508</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0000508</pub-id> <pub-id pub-id-type="pmid">17551588</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leinonen</surname> <given-names>R.</given-names></name> <name><surname>Sugawara</surname> <given-names>H.</given-names></name> <name><surname>Shumway</surname> <given-names>M.</given-names></name></person-group> (<year>2011</year>). <article-title>The sequence read archive.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>39</volume> <fpage>D19</fpage>&#x2013;<lpage>D21</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkq1019</pub-id> <pub-id pub-id-type="pmid">21062823</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liere</surname> <given-names>K.</given-names></name> <name><surname>B&#x00F6;rner</surname> <given-names>T.</given-names></name></person-group> (<year>2013</year>). <article-title>&#x201C;Development-dependent changes in the amount and structural organization of plastid DNA,&#x201D; in</article-title> <source><italic>Plastid Development in Leaves During Growth and Senescence, Advances in Photosynthesis and Respiration</italic></source> <volume>Vol. 36</volume> <role>eds</role> <person-group person-group-type="editor"><name><surname>Biswal</surname> <given-names>B.</given-names></name> <name><surname>Krupinska</surname> <given-names>K.</given-names></name> <name><surname>Biswal</surname> <given-names>U. C.</given-names></name></person-group> (<publisher-loc>Dordrecht</publisher-loc>: <publisher-name>Springer</publisher-name>), <fpage>215</fpage>&#x2013;<lpage>237</lpage>.</citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>N.</given-names></name> <name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Pong</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Comparison of next-generation sequencing systems.</article-title> <source><italic>J. BioMed. Biotechnol.</italic></source> <volume>2012</volume>:<issue>251364</issue>.</citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Maier</surname> <given-names>R. M.</given-names></name> <name><surname>Neckermann</surname> <given-names>K.</given-names></name> <name><surname>Igloi</surname> <given-names>G. L.</given-names></name> <name><surname>K&#x00F6;ssel</surname> <given-names>H.</given-names></name></person-group> (<year>1995</year>). <article-title>Complete sequence of the maize chloroplast genome: gene content, hotspots of divergence and fine tuning of genetic information by transcript editing.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>251</volume> <fpage>614</fpage>&#x2013;<lpage>628</lpage>. <pub-id pub-id-type="doi">10.1006/jmbi.1995.0460</pub-id> <pub-id pub-id-type="pmid">7666415</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Menzel</surname> <given-names>P.</given-names></name> <name><surname>Ng</surname> <given-names>K. L.</given-names></name> <name><surname>Krogh</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>Fast and sensitive taxonomic classification for metagenomics with Kaiju.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>7</volume>:<issue>11257</issue>. <pub-id pub-id-type="doi">10.1038/ncomms11257</pub-id> <pub-id pub-id-type="pmid">27071849</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Michael</surname> <given-names>T. P.</given-names></name> <name><surname>Jackson</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>The First 50 Plant Genomes.</article-title> <source><italic>Plant Genome</italic></source> <volume>6</volume> <fpage>1</fpage>&#x2013;<lpage>7</lpage>.</citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mishra</surname> <given-names>P.</given-names></name> <name><surname>Kumar</surname> <given-names>A.</given-names></name> <name><surname>Nagireddy</surname> <given-names>A.</given-names></name> <name><surname>Mani</surname> <given-names>D. N.</given-names></name> <name><surname>Shukla</surname> <given-names>A. K.</given-names></name> <name><surname>Tiwari</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>DNA barcoding: an efficient tool to overcome authentication challenges in the herbal market.</article-title> <source><italic>Plant Biotechnol. J.</italic></source> <volume>14</volume> <fpage>8</fpage>&#x2013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1111/pbi.12419</pub-id> <pub-id pub-id-type="pmid">26079154</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakamura</surname> <given-names>K.</given-names></name> <name><surname>Oshima</surname> <given-names>T.</given-names></name> <name><surname>Morimoto</surname> <given-names>T.</given-names></name> <name><surname>Ikeda</surname> <given-names>S.</given-names></name> <name><surname>Yoshikawa</surname> <given-names>H.</given-names></name> <name><surname>Shiwa</surname> <given-names>Y.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Sequence-specific error profile of Illumina sequencers.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>39</volume>:<issue>e90</issue>. <pub-id pub-id-type="doi">10.1093/nar/gkr344</pub-id> <pub-id pub-id-type="pmid">21576222</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newmaster</surname> <given-names>S. G.</given-names></name> <name><surname>Fazekas</surname> <given-names>A. J.</given-names></name> <name><surname>Ragupathy</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>DNA barcoding in land plants: evaluation of rbcL in a multigene tiered approach.</article-title> <source><italic>Can. J. Bot.</italic></source> <volume>84</volume> <fpage>335</fpage>&#x2013;<lpage>341</lpage>.</citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Newmaster</surname> <given-names>S. G.</given-names></name> <name><surname>Grguric</surname> <given-names>M.</given-names></name> <name><surname>Shanmughanandhan</surname> <given-names>D.</given-names></name> <name><surname>Ramalingam</surname> <given-names>S.</given-names></name> <name><surname>Ragupathy</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>DNA barcoding detects contamination and substitution in North American herbal products.</article-title> <source><italic>BMC Med.</italic></source> <volume>11</volume>:<issue>222</issue>. <pub-id pub-id-type="doi">10.1186/1741-7015-11-222</pub-id> <pub-id pub-id-type="pmid">24120035</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nock</surname> <given-names>C. J.</given-names></name> <name><surname>Waters</surname> <given-names>D. L. E.</given-names></name> <name><surname>Edwards</surname> <given-names>M. A.</given-names></name> <name><surname>Bowen</surname> <given-names>S. G.</given-names></name> <name><surname>Rice</surname> <given-names>N.</given-names></name> <name><surname>Cordeiro</surname> <given-names>G. M.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Chloroplast genome sequences from total DNA for plant identification.</article-title> <source><italic>Plant Biotechnol. J.</italic></source> <volume>9</volume> <fpage>328</fpage>&#x2013;<lpage>333</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-7652.2010.00558.x</pub-id> <pub-id pub-id-type="pmid">20796245</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nystedt</surname> <given-names>B.</given-names></name> <name><surname>Street</surname> <given-names>N. R.</given-names></name> <name><surname>Wetterbom</surname> <given-names>A.</given-names></name> <name><surname>Zuccolo</surname> <given-names>A.</given-names></name> <name><surname>Lin</surname> <given-names>Y.-C.</given-names></name> <name><surname>Scofield</surname> <given-names>D. G.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>The Norway spruce genome sequence and conifer genome evolution.</article-title> <source><italic>Nature</italic></source> <volume>497</volume> <fpage>579</fpage>&#x2013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1038/nature12211</pub-id> <pub-id pub-id-type="pmid">23698360</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Oldenburg</surname> <given-names>D. J.</given-names></name> <name><surname>Bendich</surname> <given-names>A. J.</given-names></name></person-group> (<year>2004</year>). <article-title>Most chloroplast DNA of maize seedlings in linear molecules with defined ends and branched forms.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>335</volume> <fpage>953</fpage>&#x2013;<lpage>970</lpage>. <pub-id pub-id-type="pmid">14698291</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ounit</surname> <given-names>R.</given-names></name> <name><surname>Wanamaker</surname> <given-names>S.</given-names></name> <name><surname>Close</surname> <given-names>T. J.</given-names></name> <name><surname>Lonardi</surname> <given-names>S.</given-names></name></person-group> (<year>2015</year>). <article-title>CLARK: fast and accurate classification of metagenomic and genomic sequences using discriminative <italic>k</italic>-mers.</article-title> <source><italic>BMC Genomics</italic></source> <volume>16</volume>:<issue>236</issue>. <pub-id pub-id-type="doi">10.1186/s12864-015-1419-2</pub-id> <pub-id pub-id-type="pmid">25879410</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pajuste</surname> <given-names>F.-D.</given-names></name> <name><surname>Kaplinski</surname> <given-names>L.</given-names></name> <name><surname>M&#x00F6;ls</surname> <given-names>M.</given-names></name> <name><surname>Puurand</surname> <given-names>T.</given-names></name> <name><surname>Lepamets</surname> <given-names>M.</given-names></name> <name><surname>Remm</surname> <given-names>M.</given-names></name></person-group> (<year>2017</year>). <article-title>FastGT: an alignment-free method for calling common SNVs directly from raw sequencing reads.</article-title> <source><italic>Sci. Rep.</italic></source> <fpage>7</fpage>:<lpage>2537</lpage>. <pub-id pub-id-type="doi">10.1038/s41598-017-02487-5</pub-id> <pub-id pub-id-type="pmid">28566690</pub-id></citation></ref> 
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Palmer</surname> <given-names>J. D.</given-names></name> <name><surname>Herbon</surname> <given-names>L. A.</given-names></name></person-group> (<year>1988</year>). <article-title>Plant mitochondrial DNA evolved rapidly in structure, but slowly in sequence.</article-title> <source><italic>J. Mol. Evol.</italic></source> <volume>28</volume> <fpage>87</fpage>&#x2013;<lpage>97</lpage>.</citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Parks</surname> <given-names>M.</given-names></name> <name><surname>Cronn</surname> <given-names>R.</given-names></name> <name><surname>Liston</surname> <given-names>A.</given-names></name></person-group> (<year>2009</year>). <article-title>Increasing phylogenetic resolution at low taxonomic levels using massively parallel sequencing of chloroplast genomes.</article-title> <source><italic>BMC Biol.</italic></source> <volume>7</volume>:<issue>84</issue>. <pub-id pub-id-type="doi">10.1186/1741-7007-7-84</pub-id> <pub-id pub-id-type="pmid">19954512</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Patro</surname> <given-names>R.</given-names></name> <name><surname>Mount</surname> <given-names>S. M.</given-names></name> <name><surname>Kingsford</surname> <given-names>C.</given-names></name></person-group> (<year>2014</year>). <article-title>Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms.</article-title> <source><italic>Nat. Biotechnol.</italic></source> <volume>32</volume> <fpage>462</fpage>&#x2013;<lpage>464</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.2862</pub-id> <pub-id pub-id-type="pmid">24752080</pub-id></citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pi&#x00F1;ol</surname> <given-names>J.</given-names></name> <name><surname>Mir</surname> <given-names>G.</given-names></name> <name><surname>Gomez-Polo</surname> <given-names>P.</given-names></name> <name><surname>Agust&#x00ED;</surname> <given-names>N.</given-names></name></person-group> (<year>2015</year>). <article-title>Universal and blocking primer mismatches limit the use of high-throughput DNA sequencing for the quantitative metabarcoding of arthropods.</article-title> <source><italic>Mol. Ecol. Resour.</italic></source> <volume>15</volume> <fpage>819</fpage>&#x2013;<lpage>830</lpage>. <pub-id pub-id-type="doi">10.1111/1755-0998.12355</pub-id> <pub-id pub-id-type="pmid">25454249</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ripp</surname> <given-names>F.</given-names></name> <name><surname>Krombholz</surname> <given-names>C. F.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Weber</surname> <given-names>M.</given-names></name> <name><surname>Sch&#x00E4;fer</surname> <given-names>A.</given-names></name> <name><surname>Schmidt</surname> <given-names>B.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>All-Food-Seq (AFS): a quantifiable screen for species in biological samples by deep DNA sequencing.</article-title> <source><italic>BMC Genomics</italic></source> <volume>15</volume>:<issue>639</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-15-639</pub-id> <pub-id pub-id-type="pmid">25081296</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Roosaare</surname> <given-names>M.</given-names></name> <name><surname>Vaher</surname> <given-names>M.</given-names></name> <name><surname>Kaplinski</surname> <given-names>L.</given-names></name> <name><surname> M&#x00F6;ls</surname> <given-names>M.</given-names></name> <name><surname>Andreson</surname> <given-names>R.</given-names></name> <name><surname>Lepamets</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2017</year>). <article-title>StrainSeeker: Fast Identification of Bacterial Strains from raw sequencing reads using user-provided guide trees.</article-title> <source><italic>PeerJ</italic></source> <fpage>5</fpage>:<lpage>e3353</lpage>. <pub-id pub-id-type="doi">10.7717/peerj.3353</pub-id> <pub-id pub-id-type="pmid">28533988</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ruhsam</surname> <given-names>M.</given-names></name> <name><surname>Rai</surname> <given-names>H. S.</given-names></name> <name><surname>Mathews</surname> <given-names>S.</given-names></name> <name><surname>Ross</surname> <given-names>T. G.</given-names></name> <name><surname>Graham</surname> <given-names>S. W.</given-names></name> <name><surname>Raubeson</surname> <given-names>L. A.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Does complete plastid genome sequencing improve species discrimination and phylogenetic resolution in <italic>Araucaria</italic>?</article-title> <source><italic>Mol. Ecol. Resour.</italic></source> <volume>15</volume> <fpage>1067</fpage>&#x2013;<lpage>1078</lpage>. <pub-id pub-id-type="doi">10.1111/1755-0998.12375</pub-id> <pub-id pub-id-type="pmid">25611173</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>S&#x00E4;rkinen</surname> <given-names>T.</given-names></name> <name><surname>George</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Predicting plastid marker variation: can complete plastid genomes from closely related species help?</article-title> <source><italic>PLOS ONE</italic></source> <volume>8</volume>:<issue>e82266</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0082266</pub-id> <pub-id pub-id-type="pmid">24312409</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shokralla</surname> <given-names>S.</given-names></name> <name><surname>Hellberg</surname> <given-names>R. S.</given-names></name> <name><surname>Handy</surname> <given-names>S. M.</given-names></name> <name><surname>King</surname> <given-names>I.</given-names></name> <name><surname>Hajibabaei</surname> <given-names>M.</given-names></name></person-group> (<year>2015</year>). <article-title>A DNA mini-barcoding system for authentication of processed fish products.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>5</volume>:<issue>15894</issue>. <pub-id pub-id-type="doi">10.1038/srep15894</pub-id> <pub-id pub-id-type="pmid">26516098</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Staats</surname> <given-names>M.</given-names></name> <name><surname>Arulandhu</surname> <given-names>A. J.</given-names></name> <name><surname>Gravendeel</surname> <given-names>B.</given-names></name> <name><surname>Holst-Jensen</surname> <given-names>A.</given-names></name> <name><surname>Scholtens</surname> <given-names>I.</given-names></name> <name><surname>Peelen</surname> <given-names>T.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Advances in DNA metabarcoding for food and wildlife forensic species identification.</article-title> <source><italic>Anal. Bioanal. Chem.</italic></source> <volume>408</volume> <fpage>4615</fpage>&#x2013;<lpage>4630</lpage>. <pub-id pub-id-type="doi">10.1007/s00216-016-9595-8</pub-id> <pub-id pub-id-type="pmid">27178552</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sugiura</surname> <given-names>M.</given-names></name> <name><surname>Hirose</surname> <given-names>T.</given-names></name> <name><surname>Sugita</surname> <given-names>M.</given-names></name></person-group> (<year>1998</year>). <article-title>Evolution and mechanism of translation in chloroplasts.</article-title> <source><italic>Annu. Rev. Genet.</italic></source> <volume>32</volume> <fpage>437</fpage>&#x2013;<lpage>459</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.genet.32.1.437</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tamura</surname> <given-names>K.</given-names></name> <name><surname>Stecher</surname> <given-names>G.</given-names></name> <name><surname>Peterson</surname> <given-names>D.</given-names></name> <name><surname>Filipski</surname> <given-names>A.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name></person-group> (<year>2013</year>). <article-title>MEGA6: molecular evolutionary genetics analysis version 6.0.</article-title> <source><italic>Mol. Biol. Evol.</italic></source> <volume>30</volume> <fpage>2725</fpage>&#x2013;<lpage>2729</lpage>. <pub-id pub-id-type="doi">10.1093/molbev/mst197</pub-id> <pub-id pub-id-type="pmid">24132122</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tillmar</surname> <given-names>A. O.</given-names></name> <name><surname>Dell&#x2019;Amico</surname> <given-names>B.</given-names></name> <name><surname>Welander</surname> <given-names>J.</given-names></name> <name><surname>Holmlund</surname> <given-names>G.</given-names></name></person-group> (<year>2013</year>). <article-title>A universal method for species identification of mammals utilizing next generation sequencing for the analysis of DNA mixtures.</article-title> <source><italic>PLOS ONE</italic></source> <volume>8</volume>:<issue>e83761</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0083761</pub-id> <pub-id pub-id-type="pmid">24358309</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wallinger</surname> <given-names>C.</given-names></name> <name><surname>Juen</surname> <given-names>A.</given-names></name> <name><surname>Staudacher</surname> <given-names>K.</given-names></name> <name><surname>Schallhart</surname> <given-names>N.</given-names></name> <name><surname>Mitterrutzner</surname> <given-names>E.</given-names></name> <name><surname>Steiner</surname> <given-names>E.-M.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Rapid plant identification using species- and group-specific primers targeting chloroplast DNA.</article-title> <source><italic>PLOS ONE</italic></source> <volume>7</volume>:<issue>e29473</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0029473</pub-id> <pub-id pub-id-type="pmid">22253728</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Liu</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>L.</given-names></name> <name><surname>Han</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>A nucleotide signature for the identification of <italic>Angelicae sinensis</italic> radix (Danggui) and its products.</article-title> <source><italic>Sci. Rep.</italic></source> <volume>6</volume>:<issue>34940</issue>. <pub-id pub-id-type="doi">10.1038/srep34940</pub-id> <pub-id pub-id-type="pmid">27713564</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weigel</surname> <given-names>D.</given-names></name> <name><surname>Mott</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>The 1001 genomes project for <italic>Arabidopsis thaliana</italic>.</article-title> <source><italic>Genome Biol.</italic></source> <volume>10</volume>:<issue>107</issue>. <pub-id pub-id-type="doi">10.1186/gb-2009-10-5-107</pub-id> <pub-id pub-id-type="pmid">19519932</pub-id></citation></ref>
<ref id="B59"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Whittall</surname> <given-names>J. B.</given-names></name> <name><surname>Syring</surname> <given-names>J.</given-names></name> <name><surname>Parks</surname> <given-names>M.</given-names></name> <name><surname>Buenrostro</surname> <given-names>J.</given-names></name> <name><surname>Dick</surname> <given-names>C.</given-names></name> <name><surname>Liston</surname> <given-names>A.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Finding a (pine) needle in a haystack: chloroplast genome sequence divergence in rare and widespread pines.</article-title> <source><italic>Mol. Ecol.</italic></source> <volume>19</volume> <fpage>100</fpage>&#x2013;<lpage>114</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-294X.2009.04474.x</pub-id> <pub-id pub-id-type="pmid">20331774</pub-id></citation></ref>
<ref id="B60"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wolfe</surname> <given-names>K. H.</given-names></name> <name><surname>Li</surname> <given-names>W. H.</given-names></name> <name><surname>Sharp</surname> <given-names>P. M.</given-names></name></person-group> (<year>1987</year>). <article-title>Rates of nucleotide substitution vary greatly among plant mitochondrial, chloroplast, and nuclear DNAs.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>84</volume> <fpage>9054</fpage>&#x2013;<lpage>9058</lpage>. <pub-id pub-id-type="pmid">3480529</pub-id></citation></ref>
<ref id="B61"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wood</surname> <given-names>D. E.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2014</year>). <article-title>Kraken: ultrafast metagenomic sequence classification using exact alignments.</article-title> <source><italic>Genome Biol.</italic></source> <volume>15</volume>:<issue>R46</issue>. <pub-id pub-id-type="doi">10.1186/gb-2014-15-3-r46</pub-id> <pub-id pub-id-type="pmid">24580807</pub-id></citation></ref>
<ref id="B62"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yang</surname> <given-names>J.-B.</given-names></name> <name><surname>Tang</surname> <given-names>M.</given-names></name> <name><surname>Li</surname> <given-names>H.-T.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.-R.</given-names></name> <name><surname>Li</surname> <given-names>D.-Z.</given-names></name></person-group> (<year>2013</year>). <article-title>Complete chloroplast genome of the genus <italic>Cymbidium</italic>: lights into the species identification, phylogenetic implications and population genetic analyses.</article-title> <source><italic>BMC Evol. Biol.</italic></source> <volume>13</volume>:<issue>84</issue>. <pub-id pub-id-type="doi">10.1186/1471-2148-13-84</pub-id> <pub-id pub-id-type="pmid">23597078</pub-id></citation></ref>
<ref id="B63"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>C.-Y.</given-names></name> <name><surname>Wang</surname> <given-names>F.-Y.</given-names></name> <name><surname>Yan</surname> <given-names>H.-F.</given-names></name> <name><surname>Hao</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>C.-M.</given-names></name> <name><surname>Ge</surname> <given-names>X.-J.</given-names></name></person-group> (<year>2012</year>). <article-title>Testing DNA barcoding in closely related groups of <italic>Lysimachia</italic> L. (Myrsinaceae).</article-title> <source><italic>Mol. Ecol. Resour.</italic></source> <volume>12</volume> <fpage>98</fpage>&#x2013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1111/j.1755-0998.2011.03076.x</pub-id> <pub-id pub-id-type="pmid">21967641</pub-id></citation></ref>
<ref id="B64"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Y.-J.</given-names></name> <name><surname>Ma</surname> <given-names>P.-F.</given-names></name> <name><surname>Li</surname> <given-names>D.-Z.</given-names></name></person-group> (<year>2011</year>). <article-title>High-throughput sequencing of six bamboo chloroplast genomes: phylogenetic implications for temperate woody bamboos (Poaceae: Bambusoideae).</article-title> <source><italic>PLOS ONE</italic></source> <volume>6</volume>:<issue>e20596</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0020596</pub-id> <pub-id pub-id-type="pmid">21655229</pub-id></citation></ref>
<ref id="B65"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name> <name><surname>Su</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Ultra-deep sequencing enables high-fidelity recovery of biodiversity for bulk arthropod samples without PCR amplification.</article-title> <source><italic>GigaScience</italic></source> <volume>2</volume>:<issue>4</issue>. <pub-id pub-id-type="doi">10.1186/2047-217X-2-4</pub-id> <pub-id pub-id-type="pmid">23587339</pub-id></citation></ref>
</ref-list>
<glossary>
<title>Abbreviations</title>
<def-list id="DL1">
<def-item>
<term>bp</term>
<def>
<p>base pair</p>
</def>
</def-item>
<def-item>
<term>NCBI</term>
<def>
<p>National Center for Biotechnology Information</p>
</def>
</def-item>
<def-item>
<term>SRA</term>
<def>
<p>Sequence Read Archive</p>
</def>
</def-item>
</def-list>
</glossary>
</back>
</article>