<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="editorial" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1114542</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.1114542</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Editorial</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Editorial: Long-read sequencing&#x2014;Pitfalls, benefits and success stories</article-title>
<alt-title alt-title-type="left-running-head">Wohlers et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fgene.2022.1114542">10.3389/fgene.2022.1114542</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Wohlers</surname>
<given-names>Inken</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1164417/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Garg</surname>
<given-names>Shilpa</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/755172/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Hehir-Kwa</surname>
<given-names>Jayne Y.</given-names>
</name>
<xref ref-type="aff" rid="aff4">
<sup>4</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1440337/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Medical Systems Biology</institution>, <institution>L&#xfc;beck Institute for Experimental Dermatology (LIED) and Institute for Cardiogenetics</institution>, <institution>University of L&#xfc;beck</institution>, <addr-line>L&#xfc;beck</addr-line>, <country>Germany</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Biology</institution>, <institution>University of Copenhagen</institution>, <addr-line>Copenhagen</addr-line>, <country>Denmark</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>NNF Center for Biosustainability</institution>, <institution>Technical University Denmark</institution>, <addr-line>Copenhagen</addr-line>, <country>Denmark</country>
</aff>
<aff id="aff4">
<sup>4</sup>
<institution>Princess Maxima Center for Pediatric Oncology</institution>, <addr-line>Utrecht</addr-line>, <country>Netherlands</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited and Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/23877/overview">Richard D. Emes</ext-link>, University of Nottingham, United Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Inken Wohlers, <email>Inken.Wohlers@uni-luebeck.de</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>04</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>1114542</elocation-id>
<history>
<date date-type="received">
<day>02</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Wohlers, Garg and Hehir-Kwa.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Wohlers, Garg and Hehir-Kwa</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<related-article id="RA1" related-article-type="commentary-article" journal-id="Front. Genet." xlink:href="https://www.frontiersin.org/researchtopic/25078" ext-link-type="uri">Editorial on the Research Topic <article-title>Long-read sequencing &#x2014; Pitfalls, benefits and success stories</article-title>
</related-article>
<kwd-group>
<kwd>sequencing</kwd>
<kwd>third-generation sequencing</kwd>
<kwd>long-read sequencing</kwd>
<kwd>nanopore</kwd>
<kwd>PacBio</kwd>
<kwd>genetic variants</kwd>
<kwd>genomics</kwd>
<kwd>transcriptomics</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<p>Long-read sequencing is an approach that holds promise to obtain genomic information for the first time in its entirety, accurately and resolved by haplotypes. In recent years the accuracy of long-read sequencing techniques, such as Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio), has significantly improved. Already, long-read sequencing can be used to reach an unprecedented genetic resolution (<xref ref-type="bibr" rid="B5">Garg, 2021</xref>), e.g. by covering previously inaccessible genomic regions or entire RNA transcripts in individual reads. Hence, long-read sequencing is a game changer in genetic research for many fields. Advancements in sequencing technology require major, simultaneous algorithmic developments, for example, in the context of base calling, variant calling and assembly.</p>
<p>The latest ONT Q20&#x2b; and PacBio high fidelity (HiFi) protocols have revolutionized sequencing by producing reads with lengths in the order of tens of kilobases up to megabases and with base accuracies exceeding 99% (<xref ref-type="bibr" rid="B6">Logsdon et al., 2020</xref>). While the HiFi technology produces relatively more accurate sequences than ONT, ONT can generate ultra-long reads and is scalable from portable devices to benchtop sequencers. The key requirement of both technologies is high molecular weight DNA.</p>
<p>Rapidly decreasing sequencing costs have fueled the generation of increasing amounts of genomic data for biodiversity applications and human health (<xref ref-type="bibr" rid="B4">De Coster et al., 2021</xref>; <xref ref-type="bibr" rid="B7">Miller et al., 2021</xref>). An example is the implementation of next-generation sequencing (NGS) into clinical practice for diagnosis, prognosis, and therapy selection in various medical fields, of which oncology has been a pioneer (<xref ref-type="bibr" rid="B2">Berger and Mardis, 2018</xref>). Undoubtedly, challenges still exist, such as input material requirements and establishing robust data analysis methods. Further, extensive method development and benchmarking are needed and ongoing to fully unlock the potential of long reads for these and other applications (<xref ref-type="bibr" rid="B1">Amarasinghe et al., 2020</xref>).</p>
<p>In this Research Topic, we cover a range of investigations that all benefit in different ways from long reads: Transcriptome profiling (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.931996/full">Vogeley et al.</ext-link>), variant calling (specifically, structural variants (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.761791/full">Bolognini and Magi</ext-link>) and somatic mitochondrial variants (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.887644/full">L&#xfc;th et al.</ext-link>), detection of circular DNA (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.867018/full">T&#xfc;ns et al.</ext-link>) and metagenome assembly (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.868280/full">Luo et al.</ext-link>). In the following, we outline the long-read sequencing context and contribution of these works.</p>
<p>The portability of ONT&#x2019;s MinION long-read sequencing device has increased IT infrastructure challenges associated with integrating these sequencers into a diagnostic laboratory setting whilst increasing the flexibility where sequencing can occur, ultimately resulting in better sample access and rapid results for patients. Addressing the corresponding need for immediate data analysis, (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.931996/full">Vogeley et al.</ext-link>) provide a self-contained transcriptome workflow for RNA sequencing data, which runs <italic>via</italic> a local webserver and explicitly supports long-read RNA-Seq. Since long reads span entire transcripts, isoform-level expression is more accurate than short read-based expression estimates (<xref ref-type="bibr" rid="B3">Chen et al., 2021</xref>). Accordingly, the workflow of (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.931996/full">Vogeley et al.</ext-link>) quantifies expression, computes differential expression, performs pathway enrichment analyses, and generates visualizations such as expression heatmaps.</p>
<p>With respect to resolving genomic variation, particularly the confident detection of structural variation from NGS data remains a challenge with no single best practice approach. There is much to be gained from using long-read sequencing data for structural variant (SV) calling, as the data is free of technical limitations such as genome coverage bias and alignment uncertainty, which plague the detection of structural variation in short-read sequencing. However, SV detection from long-read sequencing data needs dedicated tools different from those used for detecting SVs from short-read data. Meanwhile, these recent tools still require comprehensive benchmarking. Accordingly, for both real and simulated ONT data, (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.761791/full">Bolognini and Magi</ext-link>) compare five long read SV callers across four long read aligners and assess the effect of sequencing parameters on performance.</p>
<p>Another type of genetic alteration is somatic variation, i.e., mutations acquired during life course. Somatic variant calling in the context of human mitochondrial genetics has been investigated by (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.887644/full">L&#xfc;th et al.</ext-link>). Mitochondria are maternally inherited cell organelles that carry a circular genome. <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.887644/full">L&#xfc;th et al.</ext-link> generated mixtures of two mitochondrial haplotypes at different ratios for benchmarking somatic variant calling from Nanopore sequencing data with respect to variant calls obtained from deep short-read sequencing. They compared two mappers and three callers and found that basecaller, mapper, and variant caller choices affect performance. Further, somatic variants at allele frequencies of 5% are largely accurately detected, but performance decreases significantly for lower frequencies.</p>
<p>Another Research Topic article investigates the detection of circular DNA from Nanopore sequencing data. Extra-chromosomal DNA (ecDNA) is common in cancer and plays a crucial role in tumor progression. (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.867018/full">T&#xfc;ns et al.</ext-link>) present a novel open-source workflow that processes Nanopore data for detecting circular ecDNA. Its key step is a dedicated graph-theoretic approach devised by the authors. They demonstrated the workflow for the MYCN oncogene and found ecDNA&#x2019;s breakpoints reliably at the base level. The work of (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.867018/full">T&#xfc;ns et al.</ext-link>) provides readily and comprehensive detection of circular DNAs from long-read Nanopore sequencing, which facilitates biomarker discovery for cancer progression.</p>
<p>Finally, (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.868280/full">Luo et al.</ext-link>) improve long-read sequencing-based metagenome assembly by performing read correction using two complementary correction tools and then merging corrected reads before assembly. While conceptually simple, this strategy considerably improves a range of assembly quality criteria compared to stand-alone state-of-the-art metagenome assemblers, both on simulated and real data. Consequently, their corresponding workflow MetaBooster generates assemblies down to the level of individual strains.</p>
<p>In summary, our Research Topic captures the current state of long-read sequencing: All investigations covered (RNA-Seq, SVs, somatic variants, metagenomics) clearly benefit from using long reads, some are even only possible using long-read sequencing (e.g., detection of circular DNAs). However, most applications still need benchmarking to validate them with respect to short-read sequencing (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.887644/full">L&#xfc;th et al.</ext-link>) or to find a good selection or combination of analysis tools (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.761791/full">Bolognini and Magi</ext-link>; <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.868280/full">Luo et al.</ext-link>; <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.887644/full">L&#xfc;th et al.</ext-link>). Finally, developments allowing rapid, automated analysis of long-read sequencing data are brought forward (<ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.931996/full">Vogeley et al.</ext-link>), highlighting that broad implementation of tools and workflows into biomedical and clinical settings is close.</p>
</body>
<back>
<sec id="s1">
<title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p>
</sec>
<sec sec-type="COI-statement" id="s2">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s3">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Amarasinghe</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Su</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Dong</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zappia</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Ritchie</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Gouil</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Opportunities and challenges in long-read sequencing data analysis</article-title>. <source>Genome Biol.</source> <volume>21</volume> (<issue>1</issue>), <fpage>30</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-020-1935-5</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Berger</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Mardis</surname>
<given-names>E. R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>The emerging clinical relevance of genomics in cancer medicine</article-title>. <source>Nat. Rev. Clin. Oncol.</source> <volume>15</volume> (<issue>6</issue>), <fpage>353</fpage>&#x2013;<lpage>365</lpage>. <pub-id pub-id-type="doi">10.1038/s41571-018-0002-6</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Davidson</surname>
<given-names>N. M.</given-names>
</name>
<name>
<surname>Wan</surname>
<given-names>Y. K.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Yao</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Low</surname>
<given-names>H. M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>A systematic benchmark of Nanopore long read RNA sequencing for transcript level analysis in human cell lines</article-title>. <source>bioRxiv</source> <volume>2021</volume>, <fpage>440736</fpage>. <pub-id pub-id-type="doi">10.1101/2021.04.21.440736</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>De Coster</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Weissensteiner</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Sedlazeck</surname>
<given-names>F. J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Towards population-scale long-read sequencing</article-title>. <source>Nat. Rev. Genet.</source> <volume>22</volume> (<issue>9</issue>), <fpage>572</fpage>&#x2013;<lpage>587</lpage>. <pub-id pub-id-type="doi">10.1038/s41576-021-00367-3</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Garg</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Computational methods for chromosome-scale haplotype reconstruction</article-title>. <source>Genome Biol.</source> <volume>22</volume> (<issue>1</issue>), <fpage>101</fpage>. <pub-id pub-id-type="doi">10.1186/s13059-021-02328-9</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Logsdon</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Vollger</surname>
<given-names>M. R.</given-names>
</name>
<name>
<surname>Eichler</surname>
<given-names>E. E.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Long-read human genome sequencing and its applications</article-title>. <source>Nat. Rev. Genet.</source> <volume>21</volume> (<issue>10</issue>), <fpage>597</fpage>&#x2013;<lpage>614</lpage>. <pub-id pub-id-type="doi">10.1038/s41576-020-0236-x</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Miller</surname>
<given-names>D. E.</given-names>
</name>
<name>
<surname>Sulovari</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Loucks</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hoekzema</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Munson</surname>
<given-names>K. M.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Targeted long-read sequencing identifies missing disease-causing variation</article-title>. <source>Am. J. Hum. Genet.</source> <volume>108</volume> (<issue>8</issue>), <fpage>1436</fpage>&#x2013;<lpage>1449</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2021.06.006</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>