<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2017.01205</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Transcriptome Profiling Using Single-Molecule Direct RNA Sequencing Approach for In-depth Understanding of Genes in Secondary Metabolism Pathways of <italic>Camellia sinensis</italic></article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Xu</surname> <given-names>Qingshan</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/376579/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhu</surname> <given-names>Junyan</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Zhao</surname> <given-names>Shiqi</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/441471/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Hou</surname> <given-names>Yan</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Li</surname> <given-names>Fangdong</given-names></name>
</contrib>
<contrib contrib-type="author">
<name><surname>Tai</surname> <given-names>Yuling</given-names></name>
<uri xlink:href="http://loop.frontiersin.org/people/434143/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wan</surname> <given-names>Xiaochun</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/318342/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Wei</surname> <given-names>ChaoLing</given-names></name>
<xref ref-type="author-notes" rid="fn001"><sup>&#x002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/307516/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><institution>State Key Laboratory of Tea Plant Biology and Utilization, Anhui Agricultural University</institution> <country>Hefei, China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: <italic>Luo Jie, Huazhong Agricultural University, China</italic></p></fn>
<fn fn-type="edited-by"><p>Reviewed by: <italic>Alain Tissier, Leibniz-Institute of Plant Biochemistry, Germany; Vinay Kumar, Central University of Punjab, India</italic></p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x002A;Correspondence: <italic>ChaoLing Wei, <email>weichl@ahau.edu.cn</email> Xiaochun Wan, <email>xcwan@ahau.edu.cn</email></italic></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Plant Metabolism and Chemodiversity, a section of the journal Frontiers in Plant Science</p></fn></author-notes>
<pub-date pub-type="epub">
<day>11</day>
<month>07</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>08</volume>
<elocation-id>1205</elocation-id>
<history>
<date date-type="received">
<day>19</day>
<month>12</month>
<year>2016</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>06</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x00A9; 2017 Xu, Zhu, Zhao, Hou, Li, Tai, Wan and Wei.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Xu, Zhu, Zhao, Hou, Li, Tai, Wan and Wei</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Characteristic secondary metabolites, including flavonoids, theanine and caffeine, are important components of <italic>Camellia sinensis</italic>, and their biosynthesis has attracted widespread interest. Previous studies on the biosynthesis of these major secondary metabolites using next-generation sequencing technologies limited the accurately prediction of full-length (FL) splice isoforms. Herein, we applied single-molecule sequencing to pooled tea plant tissues, to provide a more complete transcriptome of <italic>C. sinensis</italic>. Moreover, we identified 94 FL transcripts and four alternative splicing events for enzyme-coding genes involved in the biosynthesis of flavonoids, theanine and caffeine. According to the comparison between long-read isoforms and assemble transcripts, we improved the quality and accuracy of genes sequenced by short-read next-generation sequencing technology. The resulting FL transcripts, together with the improved assembled transcripts and identified alternative splicing events, enhance our understanding of genes involved in the biosynthesis of characteristic secondary metabolites in <italic>C. sinensis</italic>.</p>
</abstract>
<kwd-group>
<kwd><italic>Camellia sinensis</italic></kwd>
<kwd>single-molecule sequencing</kwd>
<kwd>full-length transcript</kwd>
<kwd>alternative splicing</kwd>
<kwd>characteristic secondary metabolite</kwd>
</kwd-group>
<contract-sponsor id="cn001">National Natural Science Foundation of China<named-content content-type="fundref-id">10.13039/501100001809</named-content></contract-sponsor>
<counts>
<fig-count count="7"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="58"/>
<page-count count="11"/>
<word-count count="0"/>
</counts>
</article-meta>
</front>
<body>
<sec><title>Introduction</title>
<p>The tea plant (<italic>Camellia sinensis</italic>) is an important horticultural crop and source of one of the most popular natural non-alcoholic beverages consumed across the world (<xref ref-type="bibr" rid="B9">Chen et al., 2007</xref>; <xref ref-type="bibr" rid="B57">Zhang et al., 2015</xref>). The rich flavors of tea are mainly attributable to the characteristic secondary metabolites including flavonoids, theanine and caffeine (<xref ref-type="bibr" rid="B26">Liang et al., 2001</xref>; <xref ref-type="bibr" rid="B29">Mamati et al., 2006</xref>). These secondary compounds have been confirmed to be beneficial to human health (<xref ref-type="bibr" rid="B16">Hertog et al., 1993</xref>; <xref ref-type="bibr" rid="B6">Cabrera et al., 2006</xref>; <xref ref-type="bibr" rid="B20">Khan and Mukhtar, 2007</xref>) and contribute to the nutrient content and unique taste of tea (<xref ref-type="bibr" rid="B10">Chu and Juneja, 1997</xref>; <xref ref-type="bibr" rid="B7">Chen et al., 2008</xref>). Flavonoids such as flavanones, flavones, dihydroflavonols, flavonols, and flavin-3-ols (catechins) are derived from multiple branches of the phenylpropanoid pathway (<xref ref-type="bibr" rid="B13">Dixon and Pasinetti, 2010</xref>). Theanine is synthesized from glutamic acid and ethylamine by theanine synthetase (TS) in the roots of the tea plant (<xref ref-type="bibr" rid="B12">Deng et al., 2012</xref>). Caffeine is a purine alkaloid that is abundant in the leaves of tea plant (<xref ref-type="bibr" rid="B45">Takeda, 1994</xref>; <xref ref-type="bibr" rid="B3">Ashihara et al., 1995</xref>). A thorough understanding of the genes underlying the biosynthesis of characteristic metabolites are essential for functional genomic studies.</p>
<p>Due to the large genome size (&#x223C;4.0 Gigabases) (<xref ref-type="bibr" rid="B46">Tanaka et al., 2006</xref>) of <italic>C. sinensis</italic> and genetic barriers in tea plant tissue culture and transformation, little genomic information is available currently. The genes encoding characteristic secondary metabolite biosynthetic enzymes were mostly discovered through Sanger sequencing (<xref ref-type="bibr" rid="B43">Singh et al., 2008</xref>, <xref ref-type="bibr" rid="B42">2009b</xref>) or next-generation sequencing (<xref ref-type="bibr" rid="B40">Shi et al., 2011</xref>; <xref ref-type="bibr" rid="B52">Wu et al., 2013</xref>, <xref ref-type="bibr" rid="B53">2014</xref>; <xref ref-type="bibr" rid="B51">Wang et al., 2014</xref>; <xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). Sanger sequencing of full-length (FL) cDNA clones is the most reliable means of transcript discovery, but this method has fallen out of fashion somewhat following the advent of cheaper next-generation sequencing technologies (<xref ref-type="bibr" rid="B50">Wang et al., 2016</xref>). Using RNA-seq technology, most of the essential genes that regulate theanine, caffeine, and flavonoid biosynthesis were identified from whole tissues of tea (<xref ref-type="bibr" rid="B40">Shi et al., 2011</xref>). By studying transcription profiles of different tissues at different developmental stages, the gene network responsible for the regulation of the secondary metabolic pathways was also elucidated in tea plant (<xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). However, the relatively short length of the reads generated from next-generation sequencing prevented to assemble the FL transcripts accurately (<xref ref-type="bibr" rid="B31">Minoche et al., 2014</xref>; <xref ref-type="bibr" rid="B14">Dong et al., 2015</xref>). Furthermore, in some cases, incorrect annotation can result from the low-quality transcripts generated by short-read RNA-seq sequencing (<xref ref-type="bibr" rid="B5">Au et al., 2012</xref>, <xref ref-type="bibr" rid="B4">2013</xref>).</p>
<p>AS is an important post-transcriptional regulatory mechanism in multicellular eukaryotes that significantly enhances transcriptome diversity (<xref ref-type="bibr" rid="B17">Kalsotra and Cooper, 2011</xref>; <xref ref-type="bibr" rid="B36">Reddy et al., 2013</xref>). Next-generation sequencing revealed that over 60% of multi-exon genes are alternatively spliced in plant, such as <italic>Oryza sativa</italic> (<xref ref-type="bibr" rid="B56">Zhang et al., 2010</xref>), <italic>Arabidopsis thaliana</italic> (<xref ref-type="bibr" rid="B30">Marquez et al., 2012</xref>), and <italic>Glycine max</italic> (<xref ref-type="bibr" rid="B39">Shen et al., 2014</xref>). Up to now, very little was known about the alternative splicing in tea plant for the absence of genome information (<xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). Additionally, short reads generated from next-generation sequencing require computational <italic>de novo</italic> assembly, therefore, identification of gene isoforms are not well supported by direct experimental evidence and may suffer from a high incidence of false positives (<xref ref-type="bibr" rid="B5">Au et al., 2012</xref>). More recently, single-molecule sequencing (SMS) technology eliminates the need for assembly with much longer reads (<xref ref-type="bibr" rid="B38">Sharon et al., 2013</xref>; <xref ref-type="bibr" rid="B47">Tilgner et al., 2014</xref>, <xref ref-type="bibr" rid="B48">2015</xref>), providing direct evidence for transcript isoforms of each gene (<xref ref-type="bibr" rid="B4">Au et al., 2013</xref>; <xref ref-type="bibr" rid="B8">Chen et al., 2014</xref>; <xref ref-type="bibr" rid="B1">Abdelghany et al., 2016</xref>). These long-read transcripts can greatly increase the accuracy of transcriptome characterization compared with transcript tags assembled from short RNA-seq reads (<xref ref-type="bibr" rid="B14">Dong et al., 2015</xref>). Moreover, the higher error rate associated with SMS sequencing has been addressed by self-correction which involves the use of circular-consensus reads (<xref ref-type="bibr" rid="B23">Li Q. et al., 2014</xref>; <xref ref-type="bibr" rid="B55">Xu et al., 2015</xref>).</p>
<p>In this study, we employed an SMS approach to generate a more complete/FL transcriptome of <italic>C. sinensis</italic>. Based on long-read databases and genome sequences from bacterial artificial chromosome (BAC) libraries, we acquired FL transcripts and observed alternative splicing events for flavonoid, theanine and caffeine biosynthetic genes. The longer reads improved the quality and accuracy of transcripts generated from short-read assembly. This is the first study to use SMS technology to get the global overview of FL transcripts and alternative splicing events in tea. These results are necessary to deduce the nature of the encoded protein and in assessing a splice variant&#x2019;s role in gene regulation for <italic>C. sinensis</italic>.</p>
</sec>
<sec id="s1" sec-type="materials|methods">
<title>Materials and Methods</title>
<sec><title>Plant Materials</title>
<p>Tea plants (<italic>C. sinensis cv. Shuchazao</italic>) were grown in the 916 Tea Plantation in Shucheng County, Anhui Province, China. Eight samples of different tissues were collected from exactly the same tea plant. The tissues sampled were as follows: apical bud, first leaf, mature leaf, old leaf, stems, flowers, fruits and roots. Apical bud, first leaf, mature leaf and stems were collected on June 15, 2015; Old leaf were collected on November 13, 2015; Flowers, fruits and roots were collected on October 12, 2015. All samples were immediately frozen in liquid nitrogen and stored at -80&#x00B0;C until further use.</p>
</sec>
<sec><title>RNA Isolation</title>
<p>Total RNA was extracted using the RNeasy Plus Mini kit (Qiagen, Valencia, CA, United States). Total RNA from each sample was quantified and the quality assessed using an Agilent 2100 Bioanalyzer (Agilent Technologies, Palo Alto, CA, United States). Equal amounts of total RNA from each sample were pooled to provide 90 &#x03BC;g of total RNA. PolyA RNAs was isolated from total RNA using Dynal oligo (dT) 25 beads (Invitrogen<sup>TM</sup> Life Technologies, Carlsbad, CA, United States) according to the manufacturer&#x2019;s protocol. The isolated polyA RNAs was eluted with 20 &#x03BC;l of RNase-free water and subjected to RNA-seq library construction.</p>
</sec>
<sec><title>Library Preparation and Single-Molecule Sequencing</title>
<p>The library was prepared according to the PacBio ISO-Seq experimental workflow (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S1</xref>). The first cDNA strand was synthesized from purified polyA RNAs using a Clontech SMARTer PCR cDNA Synthesis Kit (Clontech, Mountain View, CA, United States). After PCR optimization, large-scale PCR was performed to synthesize second strand cDNA for BluePippin size selection (Sage Science, Inc., Beverly, MA, United States) with size ranges of 0-1 kb, 1-2 kb, 2-3 kb, and 3-6 kb. After size selection, another amplification was performed, and amplified, size selected cDNA products were made into SMRTbell template libraries (0-1 kb, 1-2 kb, 2-3 kb, and 3-6 kb) according to the manufacture&#x2019;s instruction.</p>
<p>Libraries were prepared by annealing a sequencing primer (SMRTbell Template Prep Kit 1.0) and binding polymerase to the primer-annealed template. Sequencing was performed on a PacBio RS II platform. A total of seven SMRT cells were conducted in this study (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
</sec>
<sec><title>Data Analyses of Single-Molecule Sequencing Data</title>
<p>Raw data from four libraries produced by Pacific Biosciences RS II were processed following the BGI PacBio transcriptome analysis procedure (SMRT analysis 2.3.0) (<bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>). In this pipeline, the &#x2018;Reads of Insert&#x2019; that could either be a FL transcript (as defined by the presence of 5&#x2032; primer, 3&#x2032; primer, and the polyA tail if applicable) or a non-full-length transcript were generated using a minimum filtering requirement of 0 and a minimum read accuracy of 0.75. In the cluster panel, the options of &#x201C;Predict Consensus Isoforms using the ICE Algorithm&#x201D; and &#x201C;Call Quiver to Polish Consensus Isoforms&#x201D; were applied to get high quality, FL, and polished consensus transcripts. Finally, the high quality consensus transcripts of multiple libraries were merged together and redundancy removed based on CD-HIT-EST (-c 0.98 -T 6 -G 0 -aL 0.90 -AL 100 -aS 0.98 -AS 30) to obtain final FL isoforms.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption><p>SMRT analysis of each library to obtain high-quality consensus isoforms. Qualified sequencing data produced by Pacific Biosciences RS II were processed using SMRT analysis (Reads of Insert, Classify, Cluster) to obtain consensus full-length isoforms.</p></caption>
<graphic xlink:href="fpls-08-01205-g001.tif"/>
</fig>
</sec>
<sec><title>Functional Annotation</title>
<p>Final FL isoforms were searched against NCBI non-redundant (NR), NCBI nucleotide sequence (NT), Swiss-Prot, Cluster of Orthologous Groups (COG) and Kyoto Encyclopedia of Genes and Genomes (KEGG, version 58) databases with a threshold <italic>E</italic>-value &#x2264;10<sup>-5</sup>. Gene Ontology (GO) annotations were determined based on the best BLASTX hit from the NR database using the Blast2GO software version 2.3.5 (<italic>E</italic>-value &#x2264;10<sup>-5</sup>). KEGG pathway analyses were performed using the KEGG Automatic Annotation Server (KAAS<sup><xref ref-type="fn" rid="fn01">1</xref></sup>).</p>
</sec>
<sec><title>Unigene and Isoform Prediction of Major Secondary Metabolites Biosynthetic Genes</title>
<p>To identify candidate genes, isoforms encoding enzymes from characteristic secondary metabolic pathways were clustered using CD-HIT software version 4.6.6 (cd-hit-est <italic>c</italic> = 0.90) (<xref ref-type="bibr" rid="B25">Li et al., 2001</xref>). The longest isoform of each cluster was defined as the candidate gene (<xref ref-type="bibr" rid="B24">Li and Godzik, 2006</xref>).</p>
<p>Alternative splicing isoforms were analyzed using BLAST<sup><xref ref-type="fn" rid="fn02">2</xref></sup> by employing transcripts from each cluster with genome sequences from the BAC library. Alternative splicing isoforms found by BLAST were viewed using the Gene Structure Display Server<sup><xref ref-type="fn" rid="fn03">3</xref></sup>.</p>
</sec>
<sec><title>Validation of Alternative Splicing Isoforms by RT-PCR</title>
<p>For PCR validation of alternative splicing events, 1 &#x03BC;g of total RNA obtained from the eight different tissues was used for reverse transcription (RT) in 20 &#x03BC;l reactions with SuperScript III reverse transcriptase (Invitrogen) and N6 random hexamers (TaKaRa, Dalian, China). Gene-specific primers were designed with Primer Premier 6 to span the predicted splicing events (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>). PCR was performed as follows: 3 min at 94&#x00B0;C, followed by 35 cycles of 94&#x00B0;C for 30 s, 55&#x00B0;C for 30 s, and 72&#x00B0;C for a time period proportional to the predicted product size. PCR amplification was monitored by 2.5% agarose gel electrophoresis.</p>
<p>PCR products were excised from the gel and purified using a gel extraction kit (Qiagen, Hilden, Germany). Purified products were cloned into the pGEM-T easy vector (Promega, United States) and plasmids were isolated using the Qiagen plasmid mini-isolation kit and confirmed by sequencing. Sequences were aligned with related isoforms to confirm the predicted alternative splicing isoforms.</p>
</sec>
<sec><title>Comparison with Short-Read Assemblies</title>
<p>Short-read sequences based on Illumina Hiseq2000 sequencing were selected for comparison with <italic>C. sinensis</italic> FL transcripts. Illumina data were obtained from same eight tea plant (<italic>C. sinensis cv</italic>. <italic>Shuchazao</italic>) tissues (buds, first leaf, mature leaf, old leaf, stems, flowers, fruits and roots) in our previous study (unpublished data). Clean reads for each tissue were assembled and annotated to generate unigenes, which were merged into the final dataset and redundancy removed by CD-HIT-EST (-c 0.98 -T 6 -G 0 -aL 0.90 -AL 100 -aS 0.98 -AS 30).</p>
<p>Candidate secondary metabolic pathway genes were identified using CD-HIT software (cd-hit-est <italic>c</italic> = 0.90) (<xref ref-type="bibr" rid="B24">Li and Godzik, 2006</xref>). Comparison of FL and Illumina-derived candidate secondary metabolic pathway genes was performed using local BLASTN (1e<sup>-10</sup> cut-off).</p>
</sec>
</sec>
<sec><title>Results</title>
<sec><title>High Quality Reads Were Obtained from <italic>Camellia sinensis</italic> by Full-Length Sequencing</title>
<p>To identify as many isoforms as possible, eight different <italic>C. sinensis</italic> tissues were harvested for RNA isolation. Equal amounts of total RNA from each tissue were pooled together and reverse-transcribed. To minimize bias that favors sequencing of shorter transcripts, multiple size-fractionated libraries (&#x003C;1, 1&#x2013;2, 2&#x2013;3 and 3&#x2013;6 kb) were made using BluePippin. Four ISO-Seq libraries were constructed for one sample, and seven cells were sequenced using the Pacific Bioscience RS II platform, generating 361,947 reads. The mean read lengths of inserts from different libraries (&#x003C;1, 1&#x2013;2, 2&#x2013;3, and 3&#x2013;6 kb) produced by SMS sequencing were 768, 2160, 3023, and 3885 bases, respectively (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
<p>SMRT analyses (Reads of Insert, Classify and Cluster) were used to obtain high-quality consensus isoforms (<bold>Figure <xref ref-type="fig" rid="F1">1</xref></bold>). Reads of Insert from different libraries (&#x003C;1, 1&#x2013;2, 2&#x2013;3, and 3&#x2013;6 kb) were classified into 38,131, 83,638, 64,244 and 24,669 FL non-chimeric transcripts, respectively, depending on whether 5&#x2032; and 3&#x2032; primer sequences or polyA tails were detected (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S3</xref>). ICE and Quiver were then used to cluster and polish the non-chimeric transcripts. After clustering and polishing, 21,093, 34,891, 26,633, and 9,021 high quality, FL, and polished consensus transcripts were generated for the four libraries, respectively (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S4</xref>). The quality distribution of consensus isoforms were closed to 1 (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S2</xref>). Finally, 91,638 high-quality consensus isoforms of the four libraries were merged into 80,217 isoforms with an average length of 1,781 bp and N50 of 2,459 bp (<bold>Table <xref ref-type="table" rid="T1">1</xref></bold>). In total, 68,360 isoforms (85.2%) were longer than 500bp, and 59900 isoforms (74.7%) were longer than 1 kb (<bold>Figure <xref ref-type="fig" rid="F2">2</xref></bold>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of final <italic>C. sinensis</italic> consensus isoforms.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<th valign="top" align="left">Sample</th>
<th valign="top" align="center">Total isoforms</th>
<th valign="top" align="center">Total bases (bp)</th>
<th valign="top" align="center">Mean length (bp)</th>
<th valign="top" align="center">N50 (bp)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Total</td>
<td valign="top" align="center">80,217</td>
<td valign="top" align="center">142,878,553</td>
<td valign="top" align="center">1,781</td>
<td valign="top" align="center">2,459</td></tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption><p>Length distribution of <italic>C. sinensis</italic> transcripts.</p></caption>
<graphic xlink:href="fpls-08-01205-g002.tif"/>
</fig>
</sec>
<sec><title>Functional Annotation and Categorization of the Isoforms</title>
<p>To predict and analyze the function of the 80,217 isoforms, we use BLAST (<xref ref-type="bibr" rid="B2">Altschul et al., 1990</xref>), BLAST2GO (<xref ref-type="bibr" rid="B11">Conesa et al., 2005</xref>), and InterProScan 5 (<xref ref-type="bibr" rid="B34">Quevillon et al., 2005</xref>) to perform functional annotation (using NR, NT, SwissProt, KEGG, COG, GO, and InterPro databases). A total of 72,877 isoforms were successfully matched to known proteins in at least one out the five databases, and 21,192 isoforms received high scores with proteins in all five databases (<bold>Figure <xref ref-type="fig" rid="F3">3</xref></bold> and <bold>Table <xref ref-type="table" rid="T2">2</xref></bold>).</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption><p>Venn diagram of NR, COG, KEGG, SwissProt, and InterPro results for the <italic>C. sinensis</italic> transcriptome.</p></caption>
<graphic xlink:href="fpls-08-01205-g003.tif"/>
</fig>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Summary of functional annotation results for <italic>C. sinensis</italic> transcripts.</p></caption>
<table cellspacing="5" cellpadding="5" frame="hsides" rules="groups">
<thead>
<tr>
<td valign="top" align="left"></td>
<td valign="top" align="left"></td>
<th valign="top" align="left">Nr</th>
<th valign="top" align="left">Nt</th>
<th valign="top" align="left">SwissProt</th>
<th valign="top" align="left">KEGG</th>
<th valign="top" align="left">COG</th>
<th valign="top" align="left">InterPro</th>
<th valign="top" align="left">GO</th>
<td valign="top" align="left"></td></tr>
<tr>
<th valign="top" align="left">Values</th>
<th valign="top" align="left">Total</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">annotated</th>
<th valign="top" align="left">Overall</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">Number</td>
<td valign="top" align="left">80,217</td>
<td valign="top" align="left">64,797</td>
<td valign="top" align="left">69,456</td>
<td valign="top" align="left">47,479</td>
<td valign="top" align="left">51,149</td>
<td valign="top" align="left">28,940</td>
<td valign="top" align="left">40,768</td>
<td valign="top" align="left">15,119</td>
<td valign="top" align="left">72,887</td>
</tr>
<tr>
<td valign="top" align="left">Percentage</td>
<td valign="top" align="left">100%</td>
<td valign="top" align="left">80.78%</td>
<td valign="top" align="left">86.59%</td>
<td valign="top" align="left">59.19%</td>
<td valign="top" align="left">63.76%</td>
<td valign="top" align="left">36.08%</td>
<td valign="top" align="left">50.82%</td>
<td valign="top" align="left">18.85%</td>
<td valign="top" align="left">90.86%</td></tr>
</tbody>
</table>
</table-wrap>
<p>To functionally classify the <italic>C. sinensis</italic> transcripts, GO terms were assigned to each isoform using BLAST2GO based on the best BLASTx hit from the NR database. In total, 15,119 isoforms were assigned GO terms, which were classified into three major categories (molecular function, cellular component and biological process; Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S3</xref>). For molecular function classification, major categories were &#x201C;catalytic activity&#x201D; (GO: 0003824) and &#x201C;binding&#x201D; (GO: 0005488). In the cellular component category, isoforms involved in the &#x201C;cell part&#x201D; (5,977, 39.5% of the total), &#x201C;cell&#x201D; (5,977, 39.5%) and &#x201C;organelle&#x201D; (4,278, 28.3%) were highly represented. The major subgroups of biological processes were &#x201C;cellular process&#x201D; (GO: 0009987) and &#x201C;metabolic process&#x201D; (GO: 0008152).</p>
<p>Cluster of Orthologous Group contains protein sequences encoded in 21 prokaryotic and eukaryotic genomes, and this database was used to evaluate the completeness of the isoforms and the validity of the annotations. A total of 28,940 isoforms were assigned to 25 functional clusters (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S4</xref>). &#x201C;General function prediction only&#x201D; (24.8%, 7,177), &#x201C;replication, recombination and repair&#x201D; (15.5%, 4,484), &#x201C;transcription&#x201D; (14.4%, 4,165), &#x201C;post-translational modification, protein turnover, chaperones&#x201D; (12.2%, 3,528), and &#x201C;signal transduction mechanisms&#x201D; (11.0%, 3,169) were the five largest categories. Secondary metabolites are very important for the taste and quality of tea, and approximately 4.1% (1,178) of isoforms were clustered into the secondary metabolism category (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S4</xref>).</p>
<p>In order to explore the biological functions and interactions of genes in <italic>C. sinensis</italic>, isoforms were searched against the KEGG database. A total of 51,149 isoforms were annotated and assigned to 135 functional categories (<bold>Table <xref ref-type="table" rid="T2">2</xref></bold> and Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S5</xref>). Among these pathways, the &#x201C;biosynthesis of secondary metabolites&#x201D; pathway included 2176 isoforms, providing a valuable resource for further gene function research.</p>
</sec>
<sec><title>Unigenes and Isoforms in Flavonoid Pathway</title>
<p>Based on the KEGG database, a total of 301 isoforms were observed in flavonoid pathway (Supplementary Data <xref ref-type="supplementary-material" rid="SM2">S1</xref>) clustering into 90 candidate genes by CD-HIT-EST (<italic>c</italic> = 0.90) software. Flavonoids are synthesized via the phenylpropanoid pathway by the enzymes phenylalanine ammonia lyase (<italic>PAL</italic>), cinnamate 4-hydroxylase (<italic>C4H</italic>) and 4-coumarate CoA ligase (<italic>4CL</italic>). Fifteen, six and seven genes were annotated as <italic>PAL</italic>, <italic>C4H</italic> and <italic>4CL</italic>, respectively (<bold>Figure <xref ref-type="fig" rid="F4">4A</xref></bold>). Among them, fifteen <italic>PAL</italic> were generated from 37 isoforms, and the isoforms &#x201C;tea17336&#x201D; and &#x201C;tea3529&#x201D; in different clusters may be transcribed from the same gene by alternative splicing (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>). Furthermore, <italic>PAL</italic> isoforms, &#x201C;tea20264,&#x201D; &#x201C;tea22666&#x201D; and &#x201C;tea19184&#x201D; shared significant similarity with <italic>CSPAL</italic> (GenBank accession number: <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AY694188">AY694188</ext-link>) which was associated with catechin accumulation (<xref ref-type="bibr" rid="B41">Singh et al., 2009a</xref>), and <italic>PAL</italic> isoform tea22927 showed 81.7% identity with <italic>PtPAL2</italic> (GenBank accession number: <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="AF480620">AF480620</ext-link>) which was expressed in heavily lignified structural cells of Quaking Aspen shoots (<xref ref-type="bibr" rid="B18">Kao et al., 2002</xref>).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption><p><italic>Camellia sinensis</italic> transcripts involved in three secondary metabolite biosynthetic pathways. Numbers in brackets following each gene indicate the number of unigenes identified in C. <italic>sinensis</italic>. <bold>(A)</bold> Flavonoid biosynthesis pathway. <bold>(B)</bold> Theanine biosynthesis pathway. <bold>(C)</bold> Caffeine biosynthesis pathway.</p></caption>
<graphic xlink:href="fpls-08-01205-g004.tif"/>
</fig>
<p>Chalcone synthase (<italic>CHS</italic>) is the first enzyme of the general flavonoid pathway, and this enzyme mediates the influx of substrate from the phenylpropanoid pathway. By mapping isoforms to genome sequences in the BAC library, <italic>CHS</italic> (tea49771 and tea53048) were characterized as alternative 5&#x2032; splice sites (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Subsequently, the stereo-specific cyclization of chalcones into naringenin is catalyzed by chalcone isomerase (<italic>CHI</italic>). Flavonoid 3&#x2032;-hydroxylase (<italic>F3</italic>&#x2032;<italic>H</italic>) and flavonoid 3&#x2032;, 5&#x2032;-hydroxylase (<italic>F3</italic>&#x2032;<italic>5</italic>&#x2032;<italic>H</italic>) catalyze the formation of eriodictyol and dihydrotricetin from naringenin (<xref ref-type="bibr" rid="B51">Wang et al., 2014</xref>). Five <italic>F3</italic>&#x2032;<italic>5</italic>&#x2032;<italic>H</italic> (tea27534, tea24671, tea22363, tea35097, and tea43773) genes were obtained from 18 isoforms in the present study. Of them, one gene (tea43773) with isoform (tea27827) from another gene (tea24671) cluster may undergo alternative splicing events (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>). A BLAST search of &#x201C;tea27827&#x201D; and &#x201C;tea24671&#x201D; revealed 98.5 and 98.4% identity with <italic>CSF3</italic>&#x2032;<italic>5</italic>&#x2032;<italic>H&#x2019;</italic> (GenBank accession number: <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="DQ194358">DQ194358</ext-link>) which played a critical role in the accumulation of tea catechins (<xref ref-type="bibr" rid="B51">Wang et al., 2014</xref>).</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption><p>Alternative splicing (AS) isoforms of characterized genes covered by PacBio long reads.</p></caption>
<graphic xlink:href="fpls-08-01205-g005.tif"/>
</fig>
<p>The formation of flavan-3-ols (e.g., catechin and gallocatechin) can be produced from leucoanthocyanidins by leucoanthocyanidin reductase (<italic>LAR</italic>) (<xref ref-type="bibr" rid="B27">Liu et al., 2016</xref>). Of the 13 <italic>LAR</italic> genes identified in this study, &#x201C;tea53448&#x201D; and &#x201C;tea51087&#x201D; were characterized as intron retention sites. Moreover, some <italic>LAR</italic> isoforms (tea51293, tea55264, tea51087, and tea51953) in our database shared significant similarity with <italic>CSLAR</italic> gene (GenBank accession no. GU992401) whose overexpression in tobacco leading to the accumulation of higher levels of epicatechin and its glucoside than of catechin (<xref ref-type="bibr" rid="B33">Pang et al., 2013</xref>). The generation of epi-flavan-3-ols (epicatechin and epigallocatechin) were achieved through a two-step reaction of leucoanthocyanidin catalyzed by leucoanthocyanidin oxidase (<italic>ANS</italic>) and anthocyanidin reductase (<italic>ANR</italic>) (<xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). There were eight <italic>ANS</italic> genes and eight <italic>ANR</italic> genes from the long-read transcripts. Among them, tea52647 and tea52640 from different clusters were fell into the intron retention class (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>). Furthermore, by aligning long-read sequences with complete CDS data from NCBI, 50 genes were designated as FL transcripts (Supplementary Data <xref ref-type="supplementary-material" rid="SM3">S2</xref>).</p>
</sec>
<sec><title>Unigenes and Isoforms in Theanine Pathway</title>
<p>In total, 123 isoforms involved in theanine biosynthesis were annotated by the KEGG database (Supplementary Data <xref ref-type="supplementary-material" rid="SM2">S1</xref>). Based on clustering analysis with CD-HIT-EST (<italic>c</italic> = 0.90), 44 candidate genes were identified from these isoforms and 28 genes were considered to be FL transcripts followed by the alignment of long-read sequences with complete CDS data from NCBI (Supplementary Data <xref ref-type="supplementary-material" rid="SM3">S2</xref>).</p>
<p>Theanine is synthesized from glutamic acid and ethylamine via <italic>TS</italic>, alanine transaminase (<italic>ALT</italic>), arginine decarboxylase (<italic>ADC</italic>), glutamine synthetase (<italic>GS</italic>), glutamate synthase (<italic>Fe-GOGAT</italic>), and glutamate dehydrogenase (<italic>GDH</italic>) (<xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). There were thirty-three <italic>GSs/TSs</italic>, four <italic>GOGATs (NADPH)</italic>, two <italic>GOGATs (Fe)</italic>, three <italic>ALTs</italic> and two <italic>ADCs</italic> in our database (<bold>Figure <xref ref-type="fig" rid="F4">4B</xref></bold>). Of 33 <italic>GSs/TS</italic> genes, tea48459 and tea 11573 were characterized as intron retention by mapping isoforms to genome sequences of the BAC library (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Additionally, sequence analysis of <italic>GS</italic> isoforms tea47901, tea53333 and tea43896 revealed 99.5, 99.1, and 99.0% identity with <italic>CsGS</italic> (Genbank accession No. EF055882) whose expression was stimulated in response to abscisic acid, salicylic acid, and hydrogen peroxide in tea plant (<xref ref-type="bibr" rid="B35">Rana et al., 2008</xref>).</p>
</sec>
<sec><title>Unigenes and Isoforms in Caffeine Pathway</title>
<p>In our database, 105 isoforms annotated by KEGG database in caffeine pathway were clustered into 37 candidate genes by CD-HIT-EST (<italic>c</italic> = 0.90) (Supplementary Data <xref ref-type="supplementary-material" rid="SM2">S1</xref>). The caffeine biosynthesis pathway is part of purine metabolism and comprises purine biosynthesis and purine modification steps (<xref ref-type="bibr" rid="B58">Zrenner et al., 2005</xref>). Purine biosynthesis starts from adenosine, and involves adenosine nucleosidase (<italic>Anase</italic>), adenine phosphoribosyltransferase (<italic>APRT</italic>), AMP deaminase (<italic>AMPDA</italic>), IMP dehydrogenase (<italic>IMPDH</italic>), and 5&#x2032;-nucleotidase (<italic>5&#x2032;-Nase</italic>). Eight <italic>APRTs</italic>, eight <italic>AMPDs</italic>, two <italic>IMPDHs</italic> and nine <italic>5&#x2032;-Nases</italic> were identified in the present study. Among them, nine <italic>5&#x2032;-Nase</italic> were yielded from 15 isoforms. Of the 15 isoforms tea14721, tea35654 and tea56228 from the same cluster appeared to undergo exon skipping and intron retention (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Moreover, tea28112 and tea31199 from different <italic>5&#x2032;-Nase</italic> were characterized as exon skipping (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>).</p>
<p>Purine modification steps include one nucleosidase reaction and three methylations (<xref ref-type="bibr" rid="B40">Shi et al., 2011</xref>). Caffeine is derived from xanthosine (<italic>XR</italic>) via 7-methylxanthosine synthase (<italic>7-NMT</italic>), N-methylnucleotidase (<italic>N-MeNase</italic>), theobromine synthase (<italic>MXMT</italic>), and tea caffeine synthase (<italic>TCS</italic>). There were one <italic>MXMT</italic> and nine <italic>TCSs</italic> in our database (<bold>Figure <xref ref-type="fig" rid="F4">4C</xref></bold>). Among nine <italic>TCSs</italic>, tea47446 and tea43456 were identified as alternative 5&#x2032; splicing event (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Another <italic>TCS</italic> isoforms (tea39349 and tea26065) from different clusters appeared to undergo intron retention (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S6</xref>). In addition, 16 genes were detected as FL transcripts by aligning long-read sequences with complete CDS data from NCBI (Supplementary Data <xref ref-type="supplementary-material" rid="SM3">S2</xref>).</p>
</sec>
<sec><title>PacBio Isoforms Improved the Quality of Transcripts from Short-Read Assembly</title>
<p>A total of 208 short-read transcripts annotated by KEGG database were obtained from our previous study (Supplementary Data <xref ref-type="supplementary-material" rid="SM2">S1</xref>), of which 143, 35, and 30 transcripts were involved in flavonoid, theanine and caffeine biosynthesis, respectively. These 208 transcripts were then clustered into 147 candidate genes by CD-HIT-EST (<italic>c</italic> = 0.90), including 105 genes in the flavonoid pathway, 18 genes in the theanine pathway, and 24 genes in the caffeine pathway (Supplementary Data <xref ref-type="supplementary-material" rid="SM2">S1</xref>).</p>
<p>We compared 147 candidate genes from Illumina sequencing with our 171 long-read genes using local BLASTN. The comparison revealed a good agreement between the short-read unigenes and the long-read database at the nucleotide level (<bold>Figure <xref ref-type="fig" rid="F6">6</xref></bold> and Supplementary Data <xref ref-type="supplementary-material" rid="SM4">S3</xref>). For the flavonoid pathway, we identified 74 (70.5%) unigenes from Short-Seq with a high degree of consistency with our database. For the theanine pathway, 15 (83.3%) short-seq unigenes were mapped to 10 long-read genes with high homology. For the caffeine pathway, 23 (95.8%) short-seq unigenes shared significant homology with 17 long-read genes.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption><p>BLAST comparison of Short-Seq sequences and long-read data.</p></caption>
<graphic xlink:href="fpls-08-01205-g006.tif"/>
</fig>
<p>However, SMS-Seq genes were longer than Short-Seq unigenes. Approximately 25.9% of the assembled unigenes from short-seq reads were &#x003C;500 bases, whereas only 2.3% of the isoforms from the PACBIO reads were &#x003C;500 bases (<bold>Figure <xref ref-type="fig" rid="F7">7</xref></bold>). Notably, many of these short-seq unigenes were completely mapped to the same genes of long-read sequences. For example, CL9445.Contig2 (ANS), CL9445.Contig3 (ANS), Unigene6902 (ANS), and Unigene22214 (ANS) were all mapped to tea48033 (ANS) with 100% homology (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S7</xref>). These results suggest that most candidate genes assembled from the Short-Seq reads did not represent FL cDNAs. Our long-read data therefore improved the quality of transcripts assembled from Illumina short reads.</p>
<fig id="F7" position="float">
<label>FIGURE 7</label>
<caption><p>Comparison of transcript length distribution from different sequencing platforms.</p></caption>
<graphic xlink:href="fpls-08-01205-g007.tif"/>
</fig>
<p>In addition, a few SMS genes shared significant homology with much longer Short-Seq unigenes. Interestingly, these Short-Seq unigenes included two complete CDS regions encoding two identical or different genes, whereas the related SMS genes were only homologous to part of the Short-Seq unigenes. For instance, tea5819 sequences were completely mapped to CL10428.Contig2, and the overlapping regions share the same complete CDS as GOGAT-Fe. However, the region of CL10428.Contig2 without SMS sequence coverage includes another complete CDS region encoding the ATP-dependent RNA helicase-like protein DB10. Tea9026 (GOGAT-Fe) sequences were mapped to two regions of CL10260.Contig7 that have the same CDS region encoding GOGAT-Fe (data not shown). This result suggests that transcripts generated from Short-Seq data may be susceptible to misassembly.</p>
</sec>
<sec><title>Validation of Alternative Splicing Events Identified by Multiple Alignment</title>
<p>To experimentally confirm the accuracy of the identified alternative splicing isoforms, three genes involved in flavonoid, theanine and caffeine biosynthesis annotated as a single transcript but present as two or more isoforms were selected for RT-PCR analysis. Primers were designed and synthesized (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S2</xref>) and used for RT-PCR using RNA from seven different tissues. The results showed that size of the fragments and the bands on the agarose gel were consistent with the alternative splicing isoforms (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). We cloned the DNA fragments corresponding to the predicted sizes and verified the isoforms by sequencing. Sequences and alignment of the verified isoforms are shown in Supplementary Data <xref ref-type="supplementary-material" rid="SM5">S4</xref>.</p>
</sec>
</sec>
<sec><title>Discussion</title>
<p>High-throughput mRNA sequencing studies using next-generation sequencing technologies have opened up a new era of transcriptome-wide research (<xref ref-type="bibr" rid="B5">Au et al., 2012</xref>; <xref ref-type="bibr" rid="B32">Mutz et al., 2013</xref>). Such approaches are particularly suitable for transcription profiling in non-model organisms that lack genomic sequences (<xref ref-type="bibr" rid="B19">Kawaharamiki et al., 2011</xref>; <xref ref-type="bibr" rid="B40">Shi et al., 2011</xref>). To date, most <italic>C. sinensis</italic> transcript studies have been based on next-generation sequencing (<xref ref-type="bibr" rid="B53">Wu et al., 2014</xref>; <xref ref-type="bibr" rid="B57">Zhang et al., 2015</xref>), and the short reads resulting from this approach have prevented the accurate assembly of FL transcripts in the absence of genomic sequence information (<xref ref-type="bibr" rid="B4">Au et al., 2013</xref>). In the present study, several Short-Seq reads from next-generation sequencing can be completely aligned to the same gene in our dataset (Supplementary Figure <xref ref-type="supplementary-material" rid="SM1">S7</xref>). Moreover, some misassembly transcripts were found by comparing with the long-read isoforms in our dataset. This result confirmed previous studies that transcripts generated from next-generation sequencing may suffer from misassembly (<xref ref-type="bibr" rid="B37">Schliesky et al., 2012</xref>; <xref ref-type="bibr" rid="B21">Li B. et al., 2014</xref>) and long reads produced by SMS sequencing technology can facilitate gene identification and annotation (<xref ref-type="bibr" rid="B14">Dong et al., 2015</xref>).</p>
<p>Previous studies have demonstrated the ability of SMS sequencing technology to generate continuous long reads (<xref ref-type="bibr" rid="B1">Abdelghany et al., 2016</xref>; <xref ref-type="bibr" rid="B50">Wang et al., 2016</xref>; <xref ref-type="bibr" rid="B54">Xu et al., 2016</xref>). Similar results were also observed in our study, resulting 80,217 isoforms with an average length of 1,781 bp were directly obtained by using SMS (Supplementary Table <xref ref-type="supplementary-material" rid="SM1">S5</xref>). By contrast, 55,088 transcripts were assembled and annotated from mixed tissue samples of <italic>C. sinensis</italic> based on next-generation sequencing, with an average unigene length of 355 bp (<xref ref-type="bibr" rid="B40">Shi et al., 2011</xref>). A total of 347,827 assembly transcripts were yielded from 13 different tea samples of various organs and developmental stages, with an average size of 791.2 bp (<xref ref-type="bibr" rid="B22">Li et al., 2015</xref>). On the other hand, we also identified 94 FL transcripts involved in the biosynthesis of flavonoids, theanine and caffeine by employing NCBI complete CDS. The above evidence indicated our results included a large number of longer transcripts specific to <italic>C</italic>. <italic>sinensis</italic> with known functions, which will be useful for improving the accuracy and quality of <italic>C.</italic> <italic>sinensis</italic> transcripts.</p>
<p>Due to its advantages, such as the highly accurate reads and the low costs, Illumina based RNA-seq is widely used for transcriptome analysis (<xref ref-type="bibr" rid="B28">Liu et al., 2012</xref>). However, for the alternative splicing events analysis, short reads require additional computational <italic>de novo</italic> assembly, therefore, it is difficult to infer the accuracy of gene model prediction (<xref ref-type="bibr" rid="B44">Steijger et al., 2013</xref>; <xref ref-type="bibr" rid="B49">Tilgner et al., 2013</xref>; <xref ref-type="bibr" rid="B50">Wang et al., 2016</xref>). These limitations were overcome with the emergence of SMS sequencing technology which can generate kilobase-sized sequencing reads in the absence of PCR amplification, where one read usually represents one FL transcript (<xref ref-type="bibr" rid="B38">Sharon et al., 2013</xref>; <xref ref-type="bibr" rid="B15">Gordon et al., 2015</xref>). In this research, alternative splicing events of some characteristic metabolic genes was predicted followed by the alignment with the BAC library, and confirmed by RT-PCR and sanger sequencing (<bold>Figure <xref ref-type="fig" rid="F5">5</xref></bold>). Our results demonstrated that SMS sequencing is highly powerful in alternative splicing event discovery and provides a rich data resource for later functional studies of different isoforms in <italic>C. sinensis</italic>.</p>
</sec>
<sec><title>Conclusion</title>
<p>In summary, we identified numerous long-read isoforms specific to <italic>C. sinensis</italic> and characterized FL transcripts and alternative splicing events related to flavonoid, theanine and caffeine biosynthesis. The availability of FL isoforms can improve <italic>C. sinensis</italic> transcriptome characterization. The identification of alternative splicing events can deduce the nature of the encoded protein and in assessing a splice variant&#x2019;s role in gene regulation for <italic>C. sinensis</italic>. Furthermore, the FL transcripts generated in our study provide a more accurate depiction of gene transcription and will greatly improve <italic>C. sinensis</italic> genome annotation in the future.</p>
</sec>
<sec><title>Accession Codes</title>
<p>Raw data and 529 isoforms generated from SMRT sequencing and 208 transcripts produced by ILLUMINA HiSeq have been submitted to the Sequence Read Archive (SRA) of the National Center for Biotechnology Information (NCBI) under accession numbers <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="SRR5460108">SRR5460108</ext-link> and <ext-link ext-link-type="DDBJ/EMBL/GenBank" xlink:href="SRX2433645">SRX2433645</ext-link>.</p>
</sec>
<sec><title>Author Contributions</title>
<p>CW and XW conceived and designed the study. QX analyzed the data and wrote the manuscript. JZ performed PCR validation experiments. YH, FL, and SZ given the advice for data analyzing. YT provided bacterial artificial chromosome libraries. All authors have read and approved the final version of the manuscript.</p>
</sec>
<sec><title>Conflict of Interest Statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="financial-disclosure">
<p><bold>Funding.</bold> This work was supported by the National Natural Science Foundation of China [grant number 31171608], the Special Innovative Province Construction in Anhui Province in 2015 [grant number 15czs08032], the Vitalizing Plan of Tea Industry in Anhui Province [2012&#x2013;2015], and the Program of Changjiang Scholars and Innovative Research Team in University [grant number IRT1101].</p>
</fn>
</fn-group>
<ack>
<p>We would like to thank the 916 Tea Plantation in Shucheng, Anhui Province, China for providing samples of tea plants.</p>
</ack>
<sec sec-type="supplementary material">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fpls.2017.01205/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fpls.2017.01205/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Presentation_1.PDF" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_1.XLSX" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_2.XLSX" id="SM3" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_3.XLSX" id="SM4" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/>
<supplementary-material xlink:href="Data_Sheet_4.DOC" id="SM5" mimetype="application/msword" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abdelghany</surname> <given-names>S. E.</given-names></name> <name><surname>Hamilton</surname> <given-names>M.</given-names></name> <name><surname>Jacobi</surname> <given-names>J. L.</given-names></name> <name><surname>Ngam</surname> <given-names>P.</given-names></name> <name><surname>Devitt</surname> <given-names>N.</given-names></name> <name><surname>Schilkey</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>A survey of the sorghum transcriptome using single-molecule long reads.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>7</volume>:<issue>11706</issue>. <pub-id pub-id-type="doi">10.1038/ncomms11706</pub-id></citation></ref>
<ref id="B2"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Altschul</surname> <given-names>S. F.</given-names></name> <name><surname>Gish</surname> <given-names>W.</given-names></name> <name><surname>Miller</surname> <given-names>W.</given-names></name> <name><surname>Myers</surname> <given-names>E. W.</given-names></name> <name><surname>Lipman</surname> <given-names>D. J.</given-names></name></person-group> (<year>1990</year>). <article-title>Basic local alignment search tool.</article-title> <source><italic>J. Mol. Biol.</italic></source> <volume>215</volume> <fpage>403</fpage>&#x2013;<lpage>410</lpage>. <pub-id pub-id-type="doi">10.1016/S0022-2836(05)80360-2</pub-id></citation></ref>
<ref id="B3"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ashihara</surname> <given-names>H.</given-names></name> <name><surname>Shimizu</surname> <given-names>H.</given-names></name> <name><surname>Takeda</surname> <given-names>Y.</given-names></name> <name><surname>Suzuki</surname> <given-names>T.</given-names></name> <name><surname>Gillies</surname> <given-names>F. M.</given-names></name> <name><surname>Crozier</surname> <given-names>A.</given-names></name></person-group> (<year>1995</year>). <article-title>Caffeine metabolism in high and low caffeine containing cultivars of <italic>Camellia sinensis</italic>.</article-title> <source><italic>Z. Naturforsch. C</italic></source> <volume>50</volume> <fpage>602</fpage>&#x2013;<lpage>607</lpage>.</citation></ref>
<ref id="B4"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Au</surname> <given-names>K. F.</given-names></name> <name><surname>Sebastiano</surname> <given-names>V.</given-names></name> <name><surname>Afshar</surname> <given-names>P. T.</given-names></name> <name><surname>Durruthy</surname> <given-names>J. D.</given-names></name> <name><surname>Lee</surname> <given-names>L.</given-names></name> <name><surname>Williams</surname> <given-names>B. A.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Characterization of the human ESC transcriptome by hybrid sequencing.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>110</volume> <fpage>4821</fpage>&#x2013;<lpage>4830</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1320101110</pub-id></citation></ref>
<ref id="B5"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Au</surname> <given-names>K. F.</given-names></name> <name><surname>Underwood</surname> <given-names>J. G.</given-names></name> <name><surname>Lee</surname> <given-names>L.</given-names></name> <name><surname>Wong</surname> <given-names>W. H.</given-names></name></person-group> (<year>2012</year>). <article-title>Improving PacBio long read accuracy by short read alignment.</article-title> <source><italic>PLoS ONE</italic></source> <volume>7</volume>:<issue>e46679</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0046679</pub-id></citation></ref>
<ref id="B6"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cabrera</surname> <given-names>C.</given-names></name> <name><surname>Artacho</surname> <given-names>R.</given-names></name> <name><surname>Gim&#x00E9;nez</surname> <given-names>R.</given-names></name></person-group> (<year>2006</year>). <article-title>Beneficial effects of green tea&#x2014;a review.</article-title> <source><italic>J. Am. Coll. Nutr.</italic></source> <volume>25</volume> <fpage>79</fpage>&#x2013;<lpage>99</lpage>. <pub-id pub-id-type="doi">10.1080/07315724.2006.10719518</pub-id></citation></ref>
<ref id="B7"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Milacic</surname> <given-names>V.</given-names></name> <name><surname>Chen</surname> <given-names>M. S.</given-names></name> <name><surname>Wan</surname> <given-names>S. B.</given-names></name> <name><surname>Lam</surname> <given-names>W. H.</given-names></name> <name><surname>Huo</surname> <given-names>C.</given-names></name><etal/></person-group> (<year>2008</year>). <article-title>Tea polyphenols, their biological effects and potential molecular targets.</article-title> <source><italic>Histol. Histopathol.</italic></source> <volume>23</volume> <fpage>487</fpage>&#x2013;<lpage>496</lpage>. <pub-id pub-id-type="doi">10.14670/HH-23.487</pub-id></citation></ref>
<ref id="B8"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Kostadima</surname> <given-names>M.</given-names></name> <name><surname>Martens</surname> <given-names>J. H. A.</given-names></name> <name><surname>Canu</surname> <given-names>G.</given-names></name> <name><surname>Garcia</surname> <given-names>S. P.</given-names></name> <name><surname>Turro</surname> <given-names>E.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Transcriptional diversity during lineage commitment of human blood progenitors.</article-title> <source><italic>Science</italic></source> <volume>345</volume> <fpage>1543</fpage>&#x2013;<lpage>1549</lpage>. <pub-id pub-id-type="doi">10.1126/science.1251033</pub-id></citation></ref>
<ref id="B9"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>Z. X.</given-names></name> <name><surname>Yang</surname> <given-names>Y. J.</given-names></name></person-group> (<year>2007</year>). <article-title>Genetic improvement and breeding of tea plant (<italic>Camellia sinensis</italic>) in China: from individual selection to hybridization and molecular breeding.</article-title> <source><italic>Euphytica</italic></source> <volume>154</volume> <fpage>239</fpage>&#x2013;<lpage>248</lpage>. <pub-id pub-id-type="doi">10.1007/s10681-006-9292-3</pub-id></citation></ref>
<ref id="B10"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chu</surname> <given-names>D. C.</given-names></name> <name><surname>Juneja</surname> <given-names>L. R.</given-names></name></person-group> (<year>1997</year>). <article-title>&#x201C;General chemical composition of green tea and its infusion,&#x201D; in</article-title> <source><italic>Chemistry &#x0026; Applications of Green Tea</italic></source>, <role>eds</role> <person-group person-group-type="editor"><name><surname>Yamamoto</surname> <given-names>T.</given-names></name> <name><surname>Juneja</surname> <given-names>L. R.</given-names></name> <name><surname>Chu</surname> <given-names>D. C.</given-names></name> <name><surname>Kim</surname> <given-names>M.</given-names></name></person-group> (<publisher-loc>Boca Raton, FL</publisher-loc>: <publisher-name>CRC Press</publisher-name>), <fpage>13</fpage>&#x2013;<lpage>22</lpage>.</citation></ref>
<ref id="B11"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Conesa</surname> <given-names>A.</given-names></name> <name><surname>G&#x00F6;tz</surname> <given-names>S.</given-names></name> <name><surname>Garc&#x00ED;ag&#x00F3;mez</surname> <given-names>J. M.</given-names></name> <name><surname>Terol</surname> <given-names>J.</given-names></name> <name><surname>Tal&#x00F3;n</surname> <given-names>M.</given-names></name> <name><surname>Robles</surname> <given-names>M.</given-names></name></person-group> (<year>2005</year>). <article-title>Blast2GO: a universal tool for annotation and visualization in functional genomics research.</article-title> <source><italic>Bioinformatics</italic></source> <volume>21</volume> <fpage>3674</fpage>&#x2013;<lpage>3676</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti610</pub-id></citation></ref>
<ref id="B12"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Deng</surname> <given-names>W. W.</given-names></name> <name><surname>Wang</surname> <given-names>S.</given-names></name> <name><surname>Qi</surname> <given-names>C.</given-names></name> <name><surname>Zhang</surname> <given-names>Z. Z.</given-names></name> <name><surname>Hu</surname> <given-names>X. Y.</given-names></name></person-group> (<year>2012</year>). <article-title>Effect of salt treatment on theanine biosynthesis in <italic>Camellia sinensis</italic> seedlings.</article-title> <source><italic>Plant Physiol. Biochem.</italic></source> <volume>56</volume> <fpage>35</fpage>&#x2013;<lpage>40</lpage>. <pub-id pub-id-type="doi">10.1016/j.plaphy.2012.04.003</pub-id></citation></ref>
<ref id="B13"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dixon</surname> <given-names>R. A.</given-names></name> <name><surname>Pasinetti</surname> <given-names>G. M.</given-names></name></person-group> (<year>2010</year>). <article-title>Flavonoids and isoflavonoids: from plant biology to agriculture and neuroscience.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>154</volume> <fpage>453</fpage>&#x2013;<lpage>457</lpage>. <pub-id pub-id-type="doi">10.1104/pp.110.161430</pub-id></citation></ref>
<ref id="B14"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dong</surname> <given-names>L.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Yang</surname> <given-names>S.</given-names></name> <name><surname>Kong</surname> <given-names>G.</given-names></name> <name><surname>Chu</surname> <given-names>J. S. C.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Single-molecule real-time transcript sequencing facilitates common wheat genome annotation and grain transcriptome research.</article-title> <source><italic>BMC Genomics</italic></source> <volume>16</volume>:<issue>1039</issue>. <pub-id pub-id-type="doi">10.1186/s12864-015-2257-y</pub-id></citation></ref>
<ref id="B15"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gordon</surname> <given-names>S. P.</given-names></name> <name><surname>Tseng</surname> <given-names>E.</given-names></name> <name><surname>Salamov</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Meng</surname> <given-names>X.</given-names></name> <name><surname>Zhao</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Widespread polycistronic transcripts in fungi revealed by single-molecule mRNA sequencing.</article-title> <source><italic>PLoS ONE</italic></source> <volume>10</volume>:<issue>e0132628</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0132628</pub-id></citation></ref>
<ref id="B16"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hertog</surname> <given-names>M. G.</given-names></name> <name><surname>Hollman</surname> <given-names>P. C.</given-names></name> <name><surname>Katan</surname> <given-names>M. B.</given-names></name> <name><surname>Kromhout</surname> <given-names>D.</given-names></name></person-group> (<year>1993</year>). <article-title>Intake of potentially anticarcinogenic flavonoids and their determinants in adults in The Netherlands.</article-title> <source><italic>Nutr. Cancer</italic></source> <volume>20</volume> <fpage>21</fpage>&#x2013;<lpage>29</lpage>. <pub-id pub-id-type="doi">10.1080/01635589309514267</pub-id></citation></ref>
<ref id="B17"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kalsotra</surname> <given-names>A.</given-names></name> <name><surname>Cooper</surname> <given-names>T. A.</given-names></name></person-group> (<year>2011</year>). <article-title>Functional consequences of developmentally regulated alternative splicing.</article-title> <source><italic>Nat. Rev. Genet.</italic></source> <volume>12</volume> <fpage>715</fpage>&#x2013;<lpage>729</lpage>. <pub-id pub-id-type="doi">10.1038/nrg3052</pub-id></citation></ref>
<ref id="B18"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kao</surname> <given-names>Y. Y.</given-names></name> <name><surname>Harding</surname> <given-names>S. A.</given-names></name> <name><surname>Tsai</surname> <given-names>C. J.</given-names></name></person-group> (<year>2002</year>). <article-title>Differential expression of two distinct phenylalanine ammonia-lyase genes in condensed tannin-accumulating and lignifying cells of quaking aspen.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>130</volume> <fpage>796</fpage>&#x2013;<lpage>807</lpage>. <pub-id pub-id-type="doi">10.1104/pp.006262</pub-id></citation></ref>
<ref id="B19"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kawaharamiki</surname> <given-names>R.</given-names></name> <name><surname>Wada</surname> <given-names>K.</given-names></name> <name><surname>Azuma</surname> <given-names>N.</given-names></name> <name><surname>Chiba</surname> <given-names>S.</given-names></name></person-group> (<year>2011</year>). <article-title>Expression profiling without genome sequence information in a non-model species, pandalid shrimp (<italic>Pandalus latirostris</italic>), by next-generation sequencing.</article-title> <source><italic>PLoS ONE</italic></source> <volume>6</volume>:<issue>e26043</issue>. <pub-id pub-id-type="doi">10.1371/journal.pone.0026043</pub-id></citation></ref>
<ref id="B20"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Khan</surname> <given-names>N.</given-names></name> <name><surname>Mukhtar</surname> <given-names>H.</given-names></name></person-group> (<year>2007</year>). <article-title>Tea polyphenols for health promotion.</article-title> <source><italic>Life Sci.</italic></source> <volume>81</volume> <fpage>519</fpage>&#x2013;<lpage>533</lpage>. <pub-id pub-id-type="doi">10.1016/j.lfs.2007.06.011</pub-id></citation></ref>
<ref id="B21"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>B.</given-names></name> <name><surname>Fillmore</surname> <given-names>N.</given-names></name> <name><surname>Bai</surname> <given-names>Y.</given-names></name> <name><surname>Collins</surname> <given-names>M.</given-names></name> <name><surname>Thomson</surname> <given-names>J. A.</given-names></name> <name><surname>Stewart</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Evaluation of de novo transcriptome assemblies from RNA-Seq data.</article-title> <source><italic>Genome Biol.</italic></source> <volume>15</volume> <issue>553</issue>. <pub-id pub-id-type="doi">10.1186/s13059-014-0553-5</pub-id></citation></ref>
<ref id="B22"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>C. F.</given-names></name> <name><surname>Yan</surname> <given-names>Z.</given-names></name> <name><surname>Yao</surname> <given-names>Y.</given-names></name> <name><surname>Zhao</surname> <given-names>Q. Y.</given-names></name> <name><surname>Wang</surname> <given-names>S. J.</given-names></name> <name><surname>Wang</surname> <given-names>X. C.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Global transcriptome and gene regulation network for secondary metabolite biosynthesis of tea plant (<italic>Camellia sinensis</italic>).</article-title> <source><italic>BMC Genomics</italic></source> <volume>16</volume>:<issue>560</issue>. <pub-id pub-id-type="doi">10.1186/s12864-015-1773-0</pub-id></citation></ref>
<ref id="B23"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Song</surname> <given-names>J.</given-names></name> <name><surname>Xu</surname> <given-names>H.</given-names></name> <name><surname>Xu</surname> <given-names>J.</given-names></name> <name><surname>Zhu</surname> <given-names>Y.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>High-accuracy de novo assembly and SNP detection of chloroplast genomes using a SMRT circular consensus sequencing strategy.</article-title> <source><italic>New Phytol.</italic></source> <volume>204</volume> <fpage>1041</fpage>&#x2013;<lpage>1049</lpage>. <pub-id pub-id-type="doi">10.1111/nph.12966</pub-id></citation></ref>
<ref id="B24"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Godzik</surname> <given-names>A.</given-names></name></person-group> (<year>2006</year>). <article-title>Cd-Hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.</article-title> <source><italic>Bioinformatics</italic></source> <volume>22</volume> <fpage>1658</fpage>&#x2013;<lpage>1659</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl158</pub-id></citation></ref>
<ref id="B25"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Jaroszewski</surname> <given-names>L.</given-names></name> <name><surname>Godzik</surname> <given-names>A.</given-names></name></person-group> (<year>2001</year>). <article-title>Clustering of highly homologous sequences to reduce the size of large protein databases.</article-title> <source><italic>Bioinformatics</italic></source> <volume>17</volume> <fpage>282</fpage>&#x2013;<lpage>283</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/17.3.282</pub-id></citation></ref>
<ref id="B26"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Ma</surname> <given-names>W.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name> <name><surname>Ying</surname> <given-names>W.</given-names></name></person-group> (<year>2001</year>). <article-title>Comparison of chemical compositions of <italic>Ilex latifolia</italic> Thumb and <italic>Camellia sinensis</italic> L.</article-title> <source><italic>Food Chem.</italic></source> <volume>75</volume> <fpage>339</fpage>&#x2013;<lpage>343</lpage>. <pub-id pub-id-type="doi">10.1016/S0308-8146(01)00209-6</pub-id></citation></ref>
<ref id="B27"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>C.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Shulaev</surname> <given-names>V.</given-names></name> <name><surname>Dixon</surname> <given-names>R. A.</given-names></name></person-group> (<year>2016</year>). <article-title>A role for leucoanthocyanidin reductase in the extension of proanthocyanidins.</article-title> <source><italic>Nat. Plants</italic></source> <volume>2</volume>:<issue>16182</issue>. <pub-id pub-id-type="doi">10.1038/nplants.2016.182</pub-id></citation></ref>
<ref id="B28"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>Y.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Hu</surname> <given-names>N.</given-names></name> <name><surname>He</surname> <given-names>Y.</given-names></name> <name><surname>Ray</surname> <given-names>P.</given-names></name><etal/></person-group> (<year>2012</year>). <article-title>Comparison of next-generation sequencing systems.</article-title> <source><italic>Biomed Res. Int.</italic></source> <volume>2012</volume>:<issue>251364</issue>. <pub-id pub-id-type="doi">10.1155/2012/251364</pub-id></citation></ref>
<ref id="B29"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mamati</surname> <given-names>G. E.</given-names></name> <name><surname>Liang</surname> <given-names>Y.</given-names></name> <name><surname>Lu</surname> <given-names>J.</given-names></name></person-group> (<year>2006</year>). <article-title>Expression of basic genes involved in tea polyphenol synthesis in relation to accumulation of catechins and total tea polyphenols.</article-title> <source><italic>J. Sci. Food Agric.</italic></source> <volume>86</volume> <fpage>459</fpage>&#x2013;<lpage>464</lpage>. <pub-id pub-id-type="doi">10.1002/jsfa.2368</pub-id></citation></ref>
<ref id="B30"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Marquez</surname> <given-names>Y.</given-names></name> <name><surname>Brown</surname> <given-names>J. W.</given-names></name> <name><surname>Simpson</surname> <given-names>C.</given-names></name> <name><surname>Barta</surname> <given-names>A.</given-names></name> <name><surname>Kalyna</surname> <given-names>M.</given-names></name></person-group> (<year>2012</year>). <article-title>Transcriptome survey reveals increased complexity of the alternative splicing landscape in <italic>Arabidopsis</italic>.</article-title> <source><italic>Genome Res.</italic></source> <volume>22</volume> <fpage>1184</fpage>&#x2013;<lpage>1195</lpage>. <pub-id pub-id-type="doi">10.1101/gr.134106.111</pub-id></citation></ref>
<ref id="B31"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Minoche</surname> <given-names>A. E.</given-names></name> <name><surname>Dohm</surname> <given-names>J. C.</given-names></name> <name><surname>Schneider</surname> <given-names>J.</given-names></name> <name><surname>Holtgr&#x00E4;we</surname> <given-names>D.</given-names></name> <name><surname>Vieh&#x00F6;ver</surname> <given-names>P.</given-names></name> <name><surname>Montfort</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Exploiting single-molecule transcript sequencing for eukaryotic gene prediction.</article-title> <source><italic>Genome Biol.</italic></source> <volume>16</volume> <issue>184</issue>. <pub-id pub-id-type="doi">10.1186/s13059-015-0729-7</pub-id></citation></ref>
<ref id="B32"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mutz</surname> <given-names>K. O.</given-names></name> <name><surname>Heilkenbrinker</surname> <given-names>A.</given-names></name> <name><surname>L&#x00F6;nne</surname> <given-names>M.</given-names></name> <name><surname>Walter</surname> <given-names>J. G.</given-names></name> <name><surname>Stahl</surname> <given-names>F.</given-names></name></person-group> (<year>2013</year>). <article-title>Transcriptome analysis using next-generation sequencing.</article-title> <source><italic>Curr. Opin. Biotechnol.</italic></source> <volume>24</volume> <fpage>22</fpage>&#x2013;<lpage>30</lpage>. <pub-id pub-id-type="doi">10.1016/j.copbio.2012.09.004</pub-id></citation></ref>
<ref id="B33"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pang</surname> <given-names>Y.</given-names></name> <name><surname>Abeysinghe</surname> <given-names>I. S. B.</given-names></name> <name><surname>He</surname> <given-names>J.</given-names></name> <name><surname>He</surname> <given-names>X. Z.</given-names></name> <name><surname>Huhman</surname> <given-names>D.</given-names></name> <name><surname>Mewan</surname> <given-names>K. M.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Functional characterization of proanthocyanidin pathway enzymes from tea and their application for metabolic engineering.</article-title> <source><italic>Plant Physiol.</italic></source> <volume>161</volume> <fpage>1103</fpage>&#x2013;<lpage>1116</lpage>. <pub-id pub-id-type="doi">10.1104/pp.112.212050</pub-id></citation></ref>
<ref id="B34"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Quevillon</surname> <given-names>E.</given-names></name> <name><surname>Silventoinen</surname> <given-names>V.</given-names></name> <name><surname>Pillai</surname> <given-names>S.</given-names></name> <name><surname>Harte</surname> <given-names>N.</given-names></name> <name><surname>Mulder</surname> <given-names>N.</given-names></name> <name><surname>Apweiler</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2005</year>). <article-title>InterProScan: protein domains identifier.</article-title> <source><italic>Nucleic Acids Res.</italic></source> <volume>33</volume> <fpage>116</fpage>&#x2013;<lpage>120</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gki442</pub-id></citation></ref>
<ref id="B35"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rana</surname> <given-names>N. K.</given-names></name> <name><surname>Mohanpuria</surname> <given-names>P.</given-names></name> <name><surname>Yadav</surname> <given-names>S. K.</given-names></name></person-group> (<year>2008</year>). <article-title>Cloning and characterization of a cytosolic glutamine synthetase from <italic>Camellia sinensis</italic> (L.) O. Kuntze that is Upregulated by ABA, SA, and H<sub>2</sub>O<sub>2</sub>.</article-title> <source><italic>Mol. Biotechnol.</italic></source> <volume>39</volume> <fpage>49</fpage>&#x2013;<lpage>56</lpage>. <pub-id pub-id-type="doi">10.1007/s12033-007-9027-2</pub-id></citation></ref>
<ref id="B36"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Reddy</surname> <given-names>A. S. N.</given-names></name> <name><surname>Marquez</surname> <given-names>Y.</given-names></name> <name><surname>Kalyna</surname> <given-names>M.</given-names></name> <name><surname>Barta</surname> <given-names>A.</given-names></name></person-group> (<year>2013</year>). <article-title>Complexity of the alternative splicing landscape in plants.</article-title> <source><italic>Plant Cell</italic></source> <volume>25</volume> <fpage>3657</fpage>&#x2013;<lpage>3683</lpage>. <pub-id pub-id-type="doi">10.1105/tpc.113.117523</pub-id></citation></ref>
<ref id="B37"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schliesky</surname> <given-names>S.</given-names></name> <name><surname>Gowik</surname> <given-names>U.</given-names></name> <name><surname>Weber</surname> <given-names>A. P. M.</given-names></name> <name><surname>Br&#x00E4;utigam</surname> <given-names>A.</given-names></name></person-group> (<year>2012</year>). <article-title>RNA-seq assembly &#x2013; are we there yet?</article-title> <source><italic>Front. Plant Sci.</italic></source> <volume>3</volume>:<issue>220</issue>. <pub-id pub-id-type="doi">10.3389/fpls.2012.00220</pub-id></citation></ref>
<ref id="B38"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sharon</surname> <given-names>D.</given-names></name> <name><surname>Tilgner</surname> <given-names>H.</given-names></name> <name><surname>Grubert</surname> <given-names>F.</given-names></name> <name><surname>Snyder</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>A single-molecule long-read survey of the human transcriptome.</article-title> <source><italic>Nat. Biotechnol.</italic></source> <volume>31</volume> <fpage>1009</fpage>&#x2013;<lpage>1014</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.2705</pub-id></citation></ref>
<ref id="B39"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shen</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Wang</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>W.</given-names></name> <name><surname>Fang</surname> <given-names>C.</given-names></name> <name><surname>Wu</surname> <given-names>M.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Global dissection of alternative splicing in paleopolyploid soybean.</article-title> <source><italic>Plant Cell</italic></source> <volume>26</volume> <fpage>996</fpage>&#x2013;<lpage>1008</lpage>. <pub-id pub-id-type="doi">10.1105/tpc.114.122739</pub-id></citation></ref>
<ref id="B40"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>C.-Y.</given-names></name> <name><surname>Yang</surname> <given-names>H.</given-names></name> <name><surname>Wei</surname> <given-names>C.-L.</given-names></name> <name><surname>Yu</surname> <given-names>O.</given-names></name> <name><surname>Zhang</surname> <given-names>Z.-Z.</given-names></name> <name><surname>Jiang</surname> <given-names>C.-J.</given-names></name><etal/></person-group> (<year>2011</year>). <article-title>Deep sequencing of the <italic>Camellia sinensis</italic> transcriptome revealed candidate genes for major metabolic pathways of tea-specific compounds.</article-title> <source><italic>BMC Genomics</italic></source> <volume>12</volume>:<issue>131</issue>. <pub-id pub-id-type="doi">10.1186/1471-2164-12-131</pub-id></citation></ref>
<ref id="B41"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>K.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name> <name><surname>Rani</surname> <given-names>A.</given-names></name> <name><surname>Gulati</surname> <given-names>A.</given-names></name> <name><surname>Ahuja</surname> <given-names>P. S.</given-names></name></person-group> (<year>2009a</year>). <article-title>Phenylalanine ammonia-lyase (PAL) and cinnamate 4-hydroxylase (C4H) and catechins (flavan-3-ols) accumulation in tea.</article-title> <source><italic>Funct. Integr. Genomics</italic></source> <volume>9</volume> <fpage>125</fpage>&#x2013;<lpage>134</lpage>. <pub-id pub-id-type="doi">10.1007/s10142-008-0092-9</pub-id></citation></ref>
<ref id="B42"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>K.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name> <name><surname>Yadav</surname> <given-names>S. K.</given-names></name> <name><surname>Ahuja</surname> <given-names>P. S.</given-names></name></person-group> (<year>2009b</year>). <article-title>Characterization of dihydroflavonol 4-reductase cDNA in tea [<italic>Camellia sinensis</italic> (L.) O. Kuntze].</article-title> <source><italic>Plant Biotechnol. Rep.</italic></source> <volume>3</volume> <fpage>95</fpage>&#x2013;<lpage>101</lpage>. <pub-id pub-id-type="doi">10.1007/s11816-008-0079-y</pub-id></citation></ref>
<ref id="B43"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Singh</surname> <given-names>K.</given-names></name> <name><surname>Rani</surname> <given-names>A.</given-names></name> <name><surname>Kumar</surname> <given-names>S.</given-names></name> <name><surname>Sood</surname> <given-names>P.</given-names></name> <name><surname>Mahajan</surname> <given-names>M.</given-names></name> <name><surname>Yadav</surname> <given-names>S. K.</given-names></name><etal/></person-group> (<year>2008</year>). <article-title>early gene of the flavonoid pathway, flavanone 3-hydroxylase, exhibits a positive relationship with the concentration of catechins in tea (<italic>Camellia sinensis</italic>).</article-title> <source><italic>Tree Physiol.</italic></source> <volume>28</volume> <fpage>1349</fpage>&#x2013;<lpage>1356</lpage>. <pub-id pub-id-type="doi">10.1093/treephys/28.9.1349</pub-id></citation></ref>
<ref id="B44"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steijger</surname> <given-names>T.</given-names></name> <name><surname>Abril</surname> <given-names>J. F.</given-names></name> <name><surname>Engstrom</surname> <given-names>P. G.</given-names></name> <name><surname>Kokocinski</surname> <given-names>F.</given-names></name> <name><surname>Hubbard</surname> <given-names>T. J.</given-names></name> <name><surname>Guigo</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>Assessment of transcript reconstruction methods for RNA-seq.</article-title> <source><italic>Nat. Methods</italic></source> <volume>10</volume> <fpage>1177</fpage>&#x2013;<lpage>1184</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.2714</pub-id></citation></ref>
<ref id="B45"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Takeda</surname> <given-names>Y.</given-names></name></person-group> (<year>1994</year>). <article-title>Differences in caffeine and tannin contents between tea [<italic>Camellia sinensis</italic>] cultivars, and application to tea breeding.</article-title> <source><italic>Japan Agric. Res. Q.</italic></source> <volume>28</volume> <fpage>117</fpage>&#x2013;<lpage>123</lpage>.</citation></ref>
<ref id="B46"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tanaka</surname> <given-names>J.</given-names></name> <name><surname>Taniguchi</surname> <given-names>F.</given-names></name> <name><surname>Hirai</surname> <given-names>N.</given-names></name> <name><surname>Yamaguchi</surname> <given-names>S.</given-names></name></person-group> (<year>2006</year>). <article-title>Estimation of the genome size of tea (<italic>Camellia sinensis</italic>), camellia (<italic>C. japonica</italic>), and their interspecific hybrids by flow cytometry.</article-title> <source><italic>Tea Res. J.</italic></source> <volume>101</volume> <fpage>1</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.5979/cha.2006.1</pub-id></citation></ref>
<ref id="B47"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tilgner</surname> <given-names>H.</given-names></name> <name><surname>Gubert</surname> <given-names>F.</given-names></name> <name><surname>Sharon</surname> <given-names>D.</given-names></name> <name><surname>Snyder</surname> <given-names>M. P.</given-names></name></person-group> (<year>2014</year>). <article-title>Defining a personal, allele-specific, and single-molecule long-read transcriptome.</article-title> <source><italic>Proc. Natl. Acad. Sci. U.S.A.</italic></source> <volume>111</volume> <fpage>9869</fpage>&#x2013;<lpage>9874</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1400447111</pub-id></citation></ref>
<ref id="B48"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tilgner</surname> <given-names>H.</given-names></name> <name><surname>Jahanbani</surname> <given-names>F.</given-names></name> <name><surname>Blauwkamp</surname> <given-names>T.</given-names></name> <name><surname>Moshrefi</surname> <given-names>A.</given-names></name> <name><surname>Jaeger</surname> <given-names>E.</given-names></name> <name><surname>Chen</surname> <given-names>F.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Comprehensive transcriptome analysis using synthetic long-read sequencing reveals molecular co-association of distant splicing events.</article-title> <source><italic>Nat. Biotechnol.</italic></source> <volume>33</volume> <fpage>736</fpage>&#x2013;<lpage>742</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.3242</pub-id></citation></ref>
<ref id="B49"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tilgner</surname> <given-names>H.</given-names></name> <name><surname>Raha</surname> <given-names>D.</given-names></name> <name><surname>Habegger</surname> <given-names>L.</given-names></name> <name><surname>Mohiuddin</surname> <given-names>M.</given-names></name> <name><surname>Gerstein</surname> <given-names>M.</given-names></name> <name><surname>Snyder</surname> <given-names>M.</given-names></name></person-group> (<year>2013</year>). <article-title>Accurate identification and analysis of human mRNA isoforms using deep long read sequencing.</article-title> <source><italic>G3</italic></source>, <fpage>387</fpage>&#x2013;<lpage>397</lpage>. <pub-id pub-id-type="doi">10.1534/g3.112.004812</pub-id></citation></ref>
<ref id="B50"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>B.</given-names></name> <name><surname>Tseng</surname> <given-names>E.</given-names></name> <name><surname>Regulski</surname> <given-names>M.</given-names></name> <name><surname>Clark</surname> <given-names>T. A.</given-names></name> <name><surname>Hon</surname> <given-names>T.</given-names></name> <name><surname>Jiao</surname> <given-names>Y.</given-names></name><etal/></person-group> (<year>2016</year>). <article-title>Unveiling the complexity of the maize transcriptome by single-molecule long-read sequencing.</article-title> <source><italic>Nat. Commun.</italic></source> <volume>7</volume>:<issue>11708</issue>. <pub-id pub-id-type="doi">10.1038/ncomms11708</pub-id></citation></ref>
<ref id="B51"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>Y. S.</given-names></name> <name><surname>Xu</surname> <given-names>Y. J.</given-names></name> <name><surname>Gao</surname> <given-names>L. P.</given-names></name> <name><surname>Yu</surname> <given-names>O.</given-names></name> <name><surname>Wang</surname> <given-names>X. Z.</given-names></name> <name><surname>He</surname> <given-names>X. J.</given-names></name><etal/></person-group> (<year>2014</year>). <article-title>Functional analysis of flavonoid 3&#x2019;,5&#x2019;-hydroxylase from tea plant (<italic>Camellia sinensis</italic>): critical role in the accumulation of catechins.</article-title> <source><italic>BMC Plant Biol.</italic></source> <volume>14</volume>:<issue>347</issue>. <pub-id pub-id-type="doi">10.1186/s12870-014-0347-7</pub-id></citation></ref>
<ref id="B52"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>H.</given-names></name> <name><surname>Chen</surname> <given-names>D.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Yu</surname> <given-names>B.</given-names></name> <name><surname>Qiao</surname> <given-names>X.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name><etal/></person-group> (<year>2013</year>). <article-title>De novo characterization of leaf transcriptome using 454 sequencing and development of EST-SSR markers in tea (<italic>Camellia sinensis</italic>).</article-title> <source><italic>Plant Mol. Biol. Rep.</italic></source> <volume>31</volume> <fpage>524</fpage>&#x2013;<lpage>538</lpage>. <pub-id pub-id-type="doi">10.1186/s12870-014-0347-7</pub-id></citation></ref>
<ref id="B53"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>Z. J.</given-names></name> <name><surname>Li</surname> <given-names>X. H.</given-names></name> <name><surname>Liu</surname> <given-names>Z. W.</given-names></name> <name><surname>Xu</surname> <given-names>Z. S.</given-names></name> <name><surname>Zhuang</surname> <given-names>J.</given-names></name></person-group> (<year>2014</year>). <article-title>De novo assembly and transcriptome characterization: novel insights into catechins biosynthesis in <italic>Camellia sinensis</italic>.</article-title> <source><italic>BMC Plant Biol.</italic></source> <volume>14</volume>:<issue>277</issue>. <pub-id pub-id-type="doi">10.1186/s12870-014-0277-4</pub-id></citation></ref>
<ref id="B54"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Z.</given-names></name> <name><surname>Luo</surname> <given-names>H.</given-names></name> <name><surname>Ji</surname> <given-names>A.</given-names></name> <name><surname>Zhang</surname> <given-names>X.</given-names></name> <name><surname>Song</surname> <given-names>J.</given-names></name> <name><surname>Chen</surname> <given-names>S.</given-names></name></person-group> (<year>2016</year>). <article-title>Global identification of the full-length transcripts and alternative splicing related to phenolic acid biosynthetic genes in <italic>Salvia miltiorrhiza</italic>.</article-title> <source><italic>Front. Plant Sci.</italic></source> <volume>7</volume>:<issue>100</issue>. <pub-id pub-id-type="doi">10.3389/fpls.2016.00100</pub-id></citation></ref>
<ref id="B55"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Xu</surname> <given-names>Z.</given-names></name> <name><surname>Peters</surname> <given-names>R. J.</given-names></name> <name><surname>Weirather</surname> <given-names>J.</given-names></name> <name><surname>Luo</surname> <given-names>H.</given-names></name> <name><surname>Liao</surname> <given-names>B.</given-names></name> <name><surname>Xin</surname> <given-names>Z.</given-names></name><etal/></person-group> (<year>2015</year>). <article-title>Full-length transcriptome sequences and splice variants obtained by a combination of sequencing platforms applied to different root tissues of <italic>Salvia miltiorrhiza</italic> and tanshinone biosynthesis.</article-title> <source><italic>Plant J.</italic></source> <volume>82</volume> <fpage>951</fpage>&#x2013;<lpage>961</lpage>. <pub-id pub-id-type="doi">10.1111/tpj.12865</pub-id></citation></ref>
<ref id="B56"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>G.</given-names></name> <name><surname>Guo</surname> <given-names>G.</given-names></name> <name><surname>Hu</surname> <given-names>X.</given-names></name> <name><surname>Yong</surname> <given-names>Z.</given-names></name> <name><surname>Li</surname> <given-names>Q.</given-names></name> <name><surname>Li</surname> <given-names>R.</given-names></name><etal/></person-group> (<year>2010</year>). <article-title>Deep RNA sequencing at single base-pair resolution reveals high complexity of the rice transcriptome.</article-title> <source><italic>Genome Res.</italic></source> <volume>20</volume> <fpage>646</fpage>&#x2013;<lpage>654</lpage>. <pub-id pub-id-type="doi">10.1101/gr.100677.109</pub-id></citation></ref>
<ref id="B57"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>H. B.</given-names></name> <name><surname>Xia</surname> <given-names>E. H.</given-names></name> <name><surname>Huang</surname> <given-names>H.</given-names></name> <name><surname>Jiang</surname> <given-names>J. J.</given-names></name> <name><surname>Liu</surname> <given-names>B. Y.</given-names></name> <name><surname>Gao</surname> <given-names>L. Z.</given-names></name></person-group> (<year>2015</year>). <article-title>De novo transcriptome assembly of the wild relative of tea tree (<italic>Camellia taliensis</italic>) and comparative analysis with tea transcriptome identified putative genes associated with tea quality and stress response.</article-title> <source><italic>BMC Genomics</italic></source> <volume>16</volume>:<issue>298</issue>. <pub-id pub-id-type="doi">10.1186/s12864-015-1494-4</pub-id></citation></ref>
<ref id="B58"><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zrenner</surname> <given-names>R.</given-names></name> <name><surname>Stitt</surname> <given-names>M.</given-names></name> <name><surname>Sonnewald</surname> <given-names>U.</given-names></name> <name><surname>Boldt</surname> <given-names>R.</given-names></name></person-group> (<year>2005</year>). <article-title>Pyrimidine and purine biosynthesis and degradation in plants.</article-title> <source><italic>Annu. Rev. Plant Biol.</italic></source> <volume>57</volume> <fpage>805</fpage>&#x2013;<lpage>836</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.arplant.57.032905.105421</pub-id></citation></ref>
</ref-list>
<fn-group>
<fn id="fn01"><label>1</label><p><ext-link ext-link-type="uri" xlink:href="http://www.genome.jp/tools/kaas/">www.genome.jp/tools/kaas/</ext-link></p></fn>
<fn id="fn02"><label>2</label><p><ext-link ext-link-type="uri" xlink:href="https://blast.ncbi.nlm.nih.gov/Blast.cgi">https://blast.ncbi.nlm.nih.gov/Blast.cgi</ext-link></p></fn>
<fn id="fn03"><label>3</label><p><ext-link ext-link-type="uri" xlink:href="http://gsds.cbi.pku.edu.cn/">http://gsds.cbi.pku.edu.cn/</ext-link></p></fn>
</fn-group>
</back>
</article>