<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="methods-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fgene.2017.00137</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>confFuse: High-Confidence Fusion Gene Detection across Tumor Entities</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name><surname>Huang</surname> <given-names>Zhiqin</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/460117/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Jones</surname> <given-names>David T. W.</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Wu</surname> <given-names>Yonghe</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/383170/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Lichter</surname> <given-names>Peter</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff4"><sup>4</sup></xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name><surname>Zapatka</surname> <given-names>Marc</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="author-notes" rid="fn001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/183296/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Division of Molecular Genetics, German Cancer Research Center</institution>, <addr-line>Heidelberg</addr-line>, <country>Germany</country></aff>
<aff id="aff2"><sup>2</sup><institution>Division of Pediatric Neurooncology, German Cancer Research Center</institution>, <addr-line>Heidelberg</addr-line>, <country>Germany</country></aff>
<aff id="aff3"><sup>3</sup><institution>Hopp-Children&#x00027;s Cancer Center at the NCT Heidelberg</institution>, <addr-line>Heidelberg</addr-line>, <country>Germany</country></aff>
<aff id="aff4"><sup>4</sup><institution>DKFZ-Heidelberg Center for Personalized Oncology (DKFZ-HIPO)</institution>, <addr-line>Heidelberg</addr-line>, <country>Germany</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mehdi Pirooznia, National Heart Lung and Blood Institute (NIH), United States</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Juntaro Matsuzaki, National Cancer Centre, Japan; Hayfa Hadi Hassani, University of Baghdad, Iraq</p></fn>
<fn fn-type="corresp" id="fn001"><p>&#x0002A;Correspondence: Marc Zapatka <email>m.zapatka&#x00040;dkfz-heidelberg.de</email></p></fn>
<fn fn-type="other" id="fn002"><p>This article was submitted to Bioinformatics and Computational Biology, a section of the journal Frontiers in Genetics</p></fn></author-notes>
<pub-date pub-type="epub">
<day>29</day>
<month>09</month>
<year>2017</year>
</pub-date>
<pub-date pub-type="collection">
<year>2017</year>
</pub-date>
<volume>8</volume>
<elocation-id>137</elocation-id>
<history>
<date date-type="received">
<day>28</day>
<month>07</month>
<year>2017</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>09</month>
<year>2017</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2017 Huang, Jones, Wu, Lichter and Zapatka.</copyright-statement>
<copyright-year>2017</copyright-year>
<copyright-holder>Huang, Jones, Wu, Lichter and Zapatka</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) or licensor are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract><p><bold>Background:</bold> Fusion genes play an important role in the tumorigenesis of many cancers. Next-generation sequencing (NGS) technologies have been successfully applied in fusion gene detection for the last several years, and a number of NGS-based tools have been developed for identifying fusion genes during this period. Most fusion gene detection tools based on RNA-seq data report a large number of candidates (mostly false positives), making it hard to prioritize candidates for experimental validation and further analysis. Selection of reliable fusion genes for downstream analysis becomes very important in cancer research. We therefore developed confFuse, a scoring algorithm to reliably select high-confidence fusion genes which are likely to be biologically relevant.</p>
<p><bold>Results:</bold> confFuse takes multiple parameters into account in order to assign each fusion candidate a confidence score, of which score &#x02265;8 indicates high-confidence fusion gene predictions. These parameters were manually curated based on our experience and on certain structural motifs of fusion genes. Compared with alternative tools, based on 96 published RNA-seq samples from different tumor entities, our method can significantly reduce the number of fusion candidates (301 high-confidence from 8,083 total predicted fusion genes) and keep high detection accuracy (recovery rate 85.7%). Validation of 18 novel, high-confidence fusions detected in three breast tumor samples resulted in a 100% validation rate.</p>
<p><bold>Conclusions:</bold> confFuse is a novel downstream filtering method that allows selection of highly reliable fusion gene candidates for further downstream analysis and experimental validations. confFuse is available at <ext-link ext-link-type="uri" xlink:href="https://github.com/Zhiqin-HUANG/confFuse">https://github.com/Zhiqin-HUANG/confFuse</ext-link>.</p></abstract>
<kwd-group>
<kwd>RNA-seq</kwd>
<kwd>next-generation sequencing</kwd>
<kwd>fusion gene</kwd>
<kwd>biomarkers</kwd>
<kwd>bioinformatics</kwd>
</kwd-group>
<counts>
<fig-count count="8"/>
<table-count count="0"/>
<equation-count count="0"/>
<ref-count count="31"/>
<page-count count="10"/>
<word-count count="5351"/>
</counts>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>A fusion gene is typically generated from two different genes due to genomic aberrations, or rarely at the transcript level (e.g., read-through co-transcript events). It can lead to enhanced expression or altered activity of an oncogene, or deregulation of a tumor suppressor gene (Abate et al., <xref ref-type="bibr" rid="B1">2014</xref>). Several technologies such as chromosome banding analysis and fluorescence <italic>in situ</italic> hybridization (FISH) have been successfully applied in detection of chromosomal alterations in the past (reviewed in e.g. Mertens et al., <xref ref-type="bibr" rid="B20">2015</xref>). More recently, next-generation sequencing (NGS) technologies such as paired-end RNA-seq have enabled the generation of accurate, high-resolution data in a single experiment, allowing for unbiased genome-wide fusion detection (Steidl et al., <xref ref-type="bibr" rid="B25">2011</xref>; Seshagiri et al., <xref ref-type="bibr" rid="B24">2012</xref>; Chmielecki et al., <xref ref-type="bibr" rid="B5">2013</xref>; Weischenfeldt et al., <xref ref-type="bibr" rid="B30">2013</xref>; Lilljebj&#x000F6;rn et al., <xref ref-type="bibr" rid="B18">2014</xref>). A great number of fusion gene detection tools/pipelines have been developed to interrogate data from NGS, particularly paired-end RNA-seq (Carrara et al., <xref ref-type="bibr" rid="B4">2013</xref>; Kumar et al., <xref ref-type="bibr" rid="B14">2016</xref>). The performance of the tools differs in terms of sensitivity and specificity, depending on the individual algorithms and filtering methods applied (Kumar et al., <xref ref-type="bibr" rid="B14">2016</xref>). Each of these tools/pipelines has its own advantages and weaknesses. A tool/pipeline should be properly chosen for each user&#x00027;s requirements, since one single tool/pipeline may not work best for all different data sets.</p>
<p>Fusion gene detection tools/pipelines generally consist of three major parts: firstly, mapping genomic data on reference genome/transcriptome based on existing alignment tools such as Bowtie (Langmead et al., <xref ref-type="bibr" rid="B16">2009</xref>; Langmead and Salzberg, <xref ref-type="bibr" rid="B15">2012</xref>) and BWA (Li and Durbin, <xref ref-type="bibr" rid="B17">2009</xref>); second, individual methods for generating fusion candidates such as deFuse (McPherson et al., <xref ref-type="bibr" rid="B19">2011</xref>), FusionMap (Ge et al., <xref ref-type="bibr" rid="B7">2011</xref>), and SOAPfuse (Jia et al., <xref ref-type="bibr" rid="B10">2013</xref>); and third, additional filtering algorithms to remove false positive candidates. The sensitivity of fusion gene detection mainly depends on the mapping ability in the alignment step and the specificity mostly depends on the methods of generating fusion candidates and the individual filtering methods.</p>
<p>Most of those tools/pipelines generate a large number of putative fusion transcripts even after filtering, of which most are likely to be false positives or of low biological interest (e.g., precursor read-through transcripts), making it hard to prioritize candidates for experimental validation. Additional filtering methods were developed based on individual datasets in order to select reliable candidates (Cancer Genome Atlas Research Network, <xref ref-type="bibr" rid="B3">2013</xref>; Torres-Garc&#x000ED;a et al., <xref ref-type="bibr" rid="B26">2014</xref>). Those individual filters of fusion gene candidates, however, may have a bias toward cancer or cell type-specific artifacts. A method which can work across different data sets would be very helpful for users. Some false positive fusion predictions may be due to sequencing/alignment artifacts or sequencing library preparation (Mertens et al., <xref ref-type="bibr" rid="B20">2015</xref>). Furthermore, strict filtering can decrease sensitivity of true fusion detection (Torres-Garc&#x000ED;a et al., <xref ref-type="bibr" rid="B26">2014</xref>). Therefore, we developed confFuse, a new scoring algorithm, which can be applied on paired-end RNA-seq across tumor entities with both high sensitivity and high detection accuracy.</p>
</sec>
<sec sec-type="materials and methods" id="s2">
<title>Materials and methods</title>
<p>confFuse was designed to rank fusion candidates based on deFuse output by assigning each fusion candidate a confidence score, with the aim of markedly reducing the total number of fusion candidates while retaining a high recall rate for true positives. It takes multiple features into account, including some from the standard deFuse output and also newly generated features, with each given a specific score weight. These features are closely relevant to mapping performance and fusion-related structure. The final confidence score is the sum of the score weights of different single/combined features (the initial baseline score is 10). These parameter weightings were manually optimized in comparison to a known validated fusion list, in order to achieve a balance between eliminating false positives whilst retaining true fusions. Fusion candidates scoring between 8 and 10 are considered as being high-confidence candidates. The main features used to calculate these score weights are described below and summarized in Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>.</p>
<sec>
<title>Training data</title>
<p>Sixteen recently published pediatric glioblastoma RNA-seq samples were chosen as the first training data (International Cancer Genome Consortium PedBrain Tumor Project, <xref ref-type="bibr" rid="B9">2016</xref>). Fusion gene candidates in these 16 samples were first identified by tools SOAPfuse and TopHat2-Fusion (Kim et al., <xref ref-type="bibr" rid="B13">2013</xref>). High-confidence candidates were then filtered for common artifacts and by visual inspection of fusion break points of exons between two fusion partners (International Cancer Genome Consortium PedBrain Tumor Project, <xref ref-type="bibr" rid="B9">2016</xref>). In total, 40 fusion genes were successfully verified by RT-PCR among the 16 samples. The first training data was mainly used to identify and select features from deFuse reports.</p>
<p>The second training data contains 96 RNA-seq samples from seven studies, including pilocytic astrocytoma (<italic>n</italic> &#x0003D; 7; Jones et al., <xref ref-type="bibr" rid="B11">2013</xref>), thyroid cancer (<italic>n</italic> &#x0003D; 5; Ricarte-Filho et al., <xref ref-type="bibr" rid="B22">2013</xref>), glioblastoma (<italic>n</italic> &#x0003D; 47; Bao et al., <xref ref-type="bibr" rid="B2">2014</xref>), lung adenocarcinoma (<italic>n</italic> &#x0003D; 28; Seo et al., <xref ref-type="bibr" rid="B23">2012</xref>), ependymoma (<italic>n</italic> &#x0003D; 7; Pajtler et al., <xref ref-type="bibr" rid="B21">2015</xref>), lung cancer liver metastasis (<italic>n</italic> &#x0003D; 1; Ju et al., <xref ref-type="bibr" rid="B12">2012</xref>), and biphenotypic sinonasal sarcoma (<italic>n</italic> &#x0003D; 1; Wang et al., <xref ref-type="bibr" rid="B29">2014</xref>). Those samples, including 126 experimentally verified fusion genes, were used to optimize the score weights of varied features.</p>
</sec>
<sec>
<title>Validation data</title>
<p>A published study of early-onset prostate cancer including 11 RNA-seq samples were chosen for validation <italic>in silico</italic> (Weischenfeldt et al., <xref ref-type="bibr" rid="B30">2013</xref>) and three primary breast cancer samples were used for experimental validation.</p>
</sec>
<sec>
<title>Artifact list</title>
<p>Despite a prominent role for oncogenic gene fusions in multiple cancer types, it is relatively rare for the exact same fusion to be detected across multiple, unrelated tumor entities. Fusions identified in multiple samples from different tumor entities based on currently available fusion detection tools are therefore mostly considered to be of high false positive rate. This high false positive prediction may be due to genomic complexity such as repeat regions or mapping artifacts in the alignment step. In total, 171 paired-end RNA-seq samples from 15 different entities were used to generate an artifact list of fusions identified in multiple samples from several different entities (Figure <xref ref-type="fig" rid="F1">1</xref>). To increase the sensitivity, a small number of verified fusions were manually excluded from the artifact list. We aim to assign high-confidence fusion genes a score between 8 and 10, and consider fusions contained in the artifact list to be of high false positive rate. confFuse therefore assigns a negative score (&#x02212;6) to those fusions in the artifact list in order to rank them outside of the range of confident predictions.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>One hundred and seventy-one paired&#x02013;end RNA&#x02013;seq samples from 15 different entities were used to generate an artifact list of fusion genes. ETMR, embryonal tumor with multilayered rosettes; CLL, chronic lymphocytic leukemia; ATRT, atypical teratoid rhabdoid tumor.</p></caption>
<graphic xlink:href="fgene-08-00137-g0001.tif"/>
</fig>
<p>When taking fusion candidates identified by deFuse (version v0.6.1) in no less than three entities (recurrence &#x02265; 3), 2190 fusions were included in the artifact list (Figure <xref ref-type="fig" rid="F2">2</xref>). For candidates identified in &#x02265;4 and &#x02265;5 entities, there are 1,409 and 995 fusions in the artifact list, respectively. In this study, we chose a threshold of three entities for the final artifact list. Of note, a small number of additional putative artifacts are still identified with each increase in the number of different tumor types, suggesting that accuracy could be further improved by increasing the complexity of the data set used to generate the artifact list.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>The number of recurrent fusion genes in 15 different entities. Fusions identified in more than two entities were selected for the final artifact list, resulting in 2,190 artifact fusions labeled with red color.</p></caption>
<graphic xlink:href="fgene-08-00137-g0002.tif"/>
</fig>
<p>In total, 64% (5881/9169) of putative fusion transcripts (<italic>n</italic> &#x0003D; 9169) identified in the second training data were found in the artifact list (<italic>n</italic> &#x0003D; 2190). Among them, 62.3% (3666/5881) were fusions from adjacent genes and 91.5% (5378/5881) were identified by deFuse as likely being a product of alternative splicing.</p>
</sec>
<sec>
<title>Split reads and spanning reads</title>
<p>One of the most important features supporting a true fusion event is the number of split reads and spanning reads. Since this is related not just to mapping performance, but also to fusion gene expression levels and sequencing depth, we found that setting a simple threshold on the number of split and spanning reads could not best distinguish true and false positive predictions. For example, a true fusion gene with low expression and low coverage sequencing depth may have only a few detectable split and spanning reads. A false positive fusion gene may have a large number of reads due to mapping artifacts and/or unreliable reads aligned to multiple genomic locations. Comparing verified fusions with all initial calls, the distribution of number of split reads and spanning reads between them is similar, of which the majority are of low read numbers (Figure <xref ref-type="fig" rid="F3">3</xref>). Most of the verified fusions in the first training data have &#x0003C;200 split reads and 50 spanning reads. A threshold purely on the number of split and spanning reads therefore cannot distinguish true and false positive fusion predictions.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>The number of split reads, spanning reads and breakpoint homology between verified and identified fusions. deFuse reported most of the verified fusion genes in 16 glioblastoma samples as containing &#x0003C;200 split reads, &#x0003C;50 spanning reads and &#x0003C;10 bp breakpoint homology. The maximum proportion of spanning reads in fusion partners aligned on a repeat region (repeat_proportion) is &#x0003C;80% among most of the verified fusions.</p></caption>
<graphic xlink:href="fgene-08-00137-g0003.tif"/>
</fig>
<p>In addition, the mapping quality of reads should also be considered. Some spanning reads can be aligned to more than one genomic position, indicating reads of low mapping quality which do not reliably support a fusion event. Breakpoint homology is the number of nucleotides near the fusion break point which can map equally well to both fusion partners, with very high breakpoint homology therefore suggesting more ambiguous support for a fusion event. Most of the verified fusions contain &#x0003C;10 homologous bases at the fusion breakpoint (Figure <xref ref-type="fig" rid="F3">3</xref>). confFuse therefore assigns a negative score (&#x02212;1) when breakpoint homology is &#x02265;10. If spanning reads are mapped on a repeat region, it is difficult to identify where they are originally from, and thus those spanning reads are considered as low mapping quality. Therefore, confFuse assigns a negative score to those fusions with the majority of spanning reads aligned on repeat regions, e.g., &#x02212;0.5 score for fusions of 80% up to 90% of spanning reads aligned on a repeat region and &#x02212;1 score for those fusions of &#x0003E;90% (Figure <xref ref-type="fig" rid="F3">3</xref>; Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
<p>Fusion genes with different fusion transcripts (i.e., splice variants) in the same sample may be of high true positive rate, especially those fusion transcripts with a high count of split reads and spanning reads. We observed that deFuse sometimes reports multiple fusion transcripts for the same pair of fusion partner genes (occurrence of the same fusion pairs) indicating high probability of true fusion events. confFuse thus considers the multiple fusion transcripts from a fusion gene as a positive factor supporting a true fusion event.</p>
<p>By combining with the occurrence of the same fusion pairs, the number of supporting reads, mapping quality of reads, possible mapping artifacts and other fusion structure related features mentioned below, confFuse also assigns a positive score from 0.5 to 2.5 to fusions with a high number of split and uniquely mapped spanning reads or a negative score from &#x02212;0.5 to &#x02212;2.0 otherwise, such as &#x02212;1.5 score when all the spanning reads are mapped on more than one genomic location (Table <xref ref-type="supplementary-material" rid="SM1">S1</xref>).</p>
</sec>
<sec>
<title>Fusion structure related features</title>
<p>Two adjacent genes in the same orientation may give rise to an apparent fusion due to read-through transcription or aberrant splicing rather than genomic rearrangement. Although some may acquire novel function, the vast majority are expected to be false positives in terms of their biological relevance. deFuse reports an altsplice feature, indicating that a fusion may arise from alternative splicing between adjacent genes. In the first training data, verified fusion genes do not contain any read-through or alternative splicing events (Figure <xref ref-type="fig" rid="F4">4</xref>). More than 75% of initially identified fusions are, however, from an alternative splicing event. Therefore, confFuse takes those fusions with read-through or alternative splicing as high false positive fusion candidates by assigning a negative score (&#x02212;4) in order to separate them from high confidence candidates. An adjacent gene fusion ESR1:CCDC170 was reported from 22 of 990 tumor samples (Veeraraghavan et al., <xref ref-type="bibr" rid="B28">2014</xref>), showing the possibility of true fusions from adjacent genes. Fusion candidates from adjacent genes but without a read-through or altsplice tag are therefore given a reduced penalty of only &#x02212;0.5. Furthermore, a higher ratio of inter- rather than intra-chromosomal fusions were detected in the verified fusions comparing with all identified fusions (Figure <xref ref-type="fig" rid="F4">4</xref>), and confFuse therefore also assigns a modest negative score (&#x02212;0.5) for intrachromosomal predictions (users can reset the score weight to zero if distance between fusion gene partners is large).</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Important feature distributions between verified and identified fusions. In the first training dataset, the verified fusions (<italic>n</italic> &#x0003D; 38) do not have any feature of read-through, alternative splicing and adjacent genes (which may also be partly due to a selection bias in those fusions selected for verification). Comparing with identified fusions, higher ratios of interchromosomal, open reading frame (orf) and exon boundaries in verified fusions were detected.</p></caption>
<graphic xlink:href="fgene-08-00137-g0004.tif"/>
</fig>
<p>True oncogenic fusions typically preserve the open reading frame in order to form a functional fusion protein, and the precise location of a fusion breakpoint point plays a critical role in demonstrating evidence supporting true positive fusions. When the location of a fusion splicing point is at a known exon boundary, such a fusion is more likely to be a true positive. Comparing with all identified fusions, we observed higher ratios of verified fusions preserving an open reading frame and showing fusion splicing points at exon boundaries (Figure <xref ref-type="fig" rid="F4">4</xref>). confFuse assigns a negative score to fusions with non-detected open reading frame (&#x02212;1 score) and with splicing point not at an exon boundary (&#x02212;1.5 score). It is also more likely to be of low biological interest when a break point is located downstream of the 3&#x02032; fusion partner because such predictive fusions may not have biological function or may arise from mapping artifacts. confFuse takes those fusions as low-confidence ones by assigning a negative score (&#x02212;4).</p>
</sec>
</sec>
<sec id="s3">
<title>Results and discussion</title>
<sec>
<title>Recovery rate of verified fusions</title>
<p>In the first training data (<italic>n</italic> &#x0003D; 16), 77.5% (31/40) of verified fusions were scored &#x02265;8 by confFuse, 15% (6/40) is of 6 &#x02264; score &#x0003C;8, and 2.5% (1/40) is &#x0003C;6 (Table <xref ref-type="supplementary-material" rid="SM2">S2</xref>). In total, 8,083 fusion gene candidates (9,169 putative transcripts) from the second training data (<italic>n</italic> &#x0003D; 96) were identified by deFuse using default settings, of which 126 fusions were previously validated by RT-PCR (Table <xref ref-type="supplementary-material" rid="SM3">S3</xref>). confFuse called 301 high-confidence fusion genes (score &#x02265; 8, 301/8,083, 3.7%). Among the 301 fusions were 108 of the 126 validated fusions, resulting in a recovery rate of 85.7% (108/126). The remaining previously validated fusions were either scored &#x0003C;8 (<italic>n</italic> &#x0003D; 5) or were not detected or were already filtered by default deFuse parameters prior to application of confFuse (<italic>n</italic> &#x0003D; 13; Table <xref ref-type="supplementary-material" rid="SM4">S4</xref>). The correlation between recovery rate and confFuse score in the second training data is given in Figure <xref ref-type="fig" rid="F5">5</xref>. As annotation features such as read-through are not provided by fusionMap and soapFuse, the number of supporting reads was chosen to compare recovery rate. As the threshold for the required minimum number of supporting reads increases, the recovery rate by soapFuse and fusionMap decreases (Figure <xref ref-type="supplementary-material" rid="SM11">S1</xref>, Tables <xref ref-type="supplementary-material" rid="SM9">S9</xref>, <xref ref-type="supplementary-material" rid="SM10">S10</xref>). For fusions with more than five supporting reads, fusionMap predicted 964 fusions, of which 101 were previously verified, resulting in a recovery rate &#x0007E;80% (101/126); soapFuse predicted 562 fusions, of which 65 were verified before, resulting in a recovery rate &#x0007E;51.6% (65/126). Thus, setting a threshold only on the number of supporting reads cannot retain a high recovery rate and results in decreasing sensitivity.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>confFuse score and recovery rate in 96 published samples. In total, 126 fusion genes were validated, 113 of which were identified by deFuse. confFuse detected 108 of 126 known validated fusions with score threshold 8. The right-hand figure shows the correlation between confFuse score and the number of fusions identified by deFuse.</p></caption>
<graphic xlink:href="fgene-08-00137-g0005.tif"/>
</fig>
</sec>
<sec>
<title>Comparison of defuse probability and confFuse confidence score</title>
<p>The distribution of deFuse&#x00027;s own probability score was compared with our confFuse confidence score for all putative fusion transcripts in the second training data (Figure <xref ref-type="fig" rid="F6">6</xref>). Notably, there are many putative transcripts with a high deFuse probability which were assigned a low score (&#x0007E; &#x02212;8) by confFuse, of which most are in the artifact list or of alternative splicing feature. None of these are in the list of 126 known validated fusions in the second training data. Most of the verified fusions (108/126) are located in the range of confFuse confidence score no less than eight. It demonstrates that confFuse is able to identify the verified fusions among hundreds of putative fusions with high deFuse probability.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Comparison of probability predicted by deFuse and confidence score by confFuse in 96 published RNA-seq samples. High density is of red color and low density is of blue color.</p></caption>
<graphic xlink:href="fgene-08-00137-g0006.tif"/>
</fig>
</sec>
<sec>
<title>Comparison of alternative fusion detection tools</title>
<p>Comparing across tools, fusionMap, deFuse (default probability score threshold &#x0003D; 0.5), deFuse-0.81 (deFuse with probability score threshold set to &#x02265;0.81, as used in McPherson et al., <xref ref-type="bibr" rid="B19">2011</xref>), soapFuse, confFuse-6.5 (score &#x02265; 6.5) and confFuse-8 (high-confidence candidates scored &#x02265; 8) showed recovery rates of 91.3, 89.7, 84.9, 73, 88.1, and 85.7% respectively for the 126 validated fusions (Figure <xref ref-type="fig" rid="F7">7</xref>; Tables <xref ref-type="supplementary-material" rid="SM4">S4</xref>&#x02013;<xref ref-type="supplementary-material" rid="SM6">S6</xref>, <xref ref-type="supplementary-material" rid="SM9">S9</xref>, <xref ref-type="supplementary-material" rid="SM10">S10</xref>), indicating confFuse can dramatically reduce the number of candidates (from 8,083 to only 301) without compromising detection accuracy when compared with other available tools.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Identified fusion genes and recovery rate of validated fusions among different tools. One hundred and twenty-six fusions were previously validated by RT-PCR. Five methods (fusionMap, deFuse, deFuse-0.81, confFuse-6.5, and confFuse-8) performed similarly in terms of recovery rate. confFuse generated much less fusion candidates than the others (higher specificity) while identifying comparable number of validated fusions (similar sensitivity).</p></caption>
<graphic xlink:href="fgene-08-00137-g0007.tif"/>
</fig>
</sec>
<sec>
<title>Validation of confFuse predicted candidates</title>
<p>To evaluate the accuracy of high-confidence candidate predictions (score &#x02265; 8), three primary breast tumor samples were sequenced to generate paired-end RNA-seq data. In total, deFuse predicted 1,026 fusion genes in the three samples, of which 18 scored &#x02265;8 by confFuse. All 18 high-confidence candidates were validated with RT-PCR followed by Sanger sequencing, resulting in 100% validation rate (Figure <xref ref-type="fig" rid="F8">8</xref>; Table <xref ref-type="supplementary-material" rid="SM7">S7</xref>). To the best of our knowledge, the 18 novel fusion genes haven&#x00027;t been validated before. Interestingly, one of them (QKI:PACRG) was predicted in three of 1,019 breast cancer samples from TCGA (<ext-link ext-link-type="uri" xlink:href="http://www.tumorfusions.org">www.tumorfusions.org</ext-link>), indicating that QKI:PACRG may be a novel recurrent fusion in breast cancer. QKI can suppress cell proliferation and prevent inappropriate activation of the Notch signaling pathway in lung cancer (Zong et al., <xref ref-type="bibr" rid="B31">2014</xref>) and PACRG is an evolutionarily conserved protein with currently unclear function (Dawe et al., <xref ref-type="bibr" rid="B6">2005</xref>). Furthermore, we randomly chose some candidates scored &#x0003C;8 for validation. Nine of 13 candidates (6.5 &#x02264; score &#x02264; 7.5, medium confidence) and three of 10 candidates (0.5 &#x02264; score &#x02264; 6, low confidence) were experimentally verified (Table <xref ref-type="supplementary-material" rid="SM7">S7</xref>). Most of the genes involved in verified fusions were in the expression level of FPKM &#x0003C;50 (median&#x02248;6.44; Figure <xref ref-type="supplementary-material" rid="SM11">S2</xref>), indicating that it is not only very highly-expressed candidates being detected but also lowly-expressed ones.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Eighteen high-confidence fusions validated by RT-PCR followed by Sanger sequencing in three primary breast tumor samples. Circular layout is based on tool (Gu et al., <xref ref-type="bibr" rid="B8">2014</xref>).</p></caption>
<graphic xlink:href="fgene-08-00137-g0008.tif"/>
</fig>
<p>To find out whether or not those verified fusions are individual tumor specific rather than simply artifacts of pan-breast tumor expression, the primers of verified fusions in each sample were used for validation in the other two breast primary tumor samples as control. In total, primers (Untergasser et al., <xref ref-type="bibr" rid="B27">2012</xref>) of 14 verified fusions were used, of which 13 showed true negative in two control samples and one showed a false negative in one control sample DR8V (Figure <xref ref-type="supplementary-material" rid="SM11">S3</xref>). This &#x0201C;false negative&#x0201D; fusion (CD3D:TOM1L2) in sample DR8V shows a much weaker band in gel compared with the one in sample AC72 where the fusion was true positive, and appears as a double band of different size to AC72, possibly indicating an unspecific PCR product.</p>
<p>In addition, 11 published early-onset prostate cancer samples were used as an <italic>in silico</italic> validation dataset. Eight hundred and forty-nine fusion genes were called by deFuse, 24 of which were identified as high-confidence fusion (score &#x02265; 8) by confFuse. The well-known E26 transformation-specific (ETS) fusions were detected by confFuse in 10 of 11 samples, which are the same as previously published results (Weischenfeldt et al., <xref ref-type="bibr" rid="B30">2013</xref>). Among the 24 fusions, 17 were confirmed by DNA-seq, FISH validation or known ETS fusions, resulting in a &#x0007E;70% (17/24) recovery rate (Table <xref ref-type="supplementary-material" rid="SM8">S8</xref>).</p>
<p>Overall, confFuse can classify fusion candidates into three groups, namely high-confidence (8 &#x02264; score &#x02264; 10), medium-confidence (6.5 &#x02264; score &#x02264; 7.5) and low-confidence (score &#x02264; 6) fusions, which makes the users more convenient to prioritize candidates for validations. Not only novel and biologically relevant fusions can be identified by confFuse, but also well-known fusions across tumor entities can be detected by confFuse.</p>
</sec>
</sec>
<sec sec-type="conclusions" id="s4">
<title>Conclusions</title>
<p>Based on deFuse reports, the scoring algorithm confFuse assigns each putative fusion transcript a confidence score. In three breast tumor samples, we achieved 100% true positive rate for high-confidence fusion candidates. Once more verified fusion genes are available as reference data, score weight optimization could be further improved. Users can also customize the score weights based on their experience to better analyze their specific data. In summary, confFuse can reliably select high-confidence fusion genes that are more likely to be biologically relevant, achieving both high validation rate and high detection accuracy, while reducing the number of candidates to a realistic number for validation.</p>
</sec>
<sec id="s5">
<title>Availability of data and material</title>
<p>Training dataset:</p>
<list list-type="simple">
<list-item><p>Pediatric glioblastoma: EGAS00001001139 (International Cancer Genome Consortium PedBrain Tumor Project, <xref ref-type="bibr" rid="B9">2016</xref>)</p></list-item>
<list-item><p>Pilocytic astrocytoma: EGAS00001000381 (Jones et al., <xref ref-type="bibr" rid="B11">2013</xref>)</p></list-item>
<list-item><p>Thyroid cancer: SRP027364 (Ricarte-Filho et al., <xref ref-type="bibr" rid="B22">2013</xref>)</p></list-item>
<list-item><p>Glioblastomas: GSE48865 (Bao et al., <xref ref-type="bibr" rid="B2">2014</xref>)</p></list-item>
<list-item><p>Lund adenocarcinoma: ERP001058 (Seo et al., <xref ref-type="bibr" rid="B23">2012</xref>)</p></list-item>
<list-item><p>Ependymoma: Application from the authors (Pajtler et al., <xref ref-type="bibr" rid="B21">2015</xref>)</p></list-item>
<list-item><p>Lung cancer liver metastasis: ERP001058 (Ju et al., <xref ref-type="bibr" rid="B12">2012</xref>)</p></list-item>
<list-item><p>Biphenotypic sinonasal sarcoma: GSE52257 (Wang et al., <xref ref-type="bibr" rid="B29">2014</xref>)</p></list-item>
</list>
<p>Validation dataset:</p>
<list list-type="simple">
<list-item><p>Prostate cancer: EGAS00001000258 (Weischenfeldt et al., <xref ref-type="bibr" rid="B30">2013</xref>)</p></list-item>
</list>
</sec>
<sec id="s6">
<title>Ethics statement</title>
<p>This study was carried out in accordance with the recommendations of the ethics committee of the University of Heidelberg, with written informed consent obtained from all subjects in accordance with the Declaration of Helsinki.</p>
</sec>
<sec id="s7">
<title>Author contributions</title>
<p>ZH wrote the manuscript and implemented the method. PL, YW, DJ, ZH, and MZ designed the experimental validation and interpreted the results. ZH designed the primers and YW performed the validation. DJ, YW, and MZ revised the manuscript. All the authors read and approved the final manuscript.</p>
<sec>
<title>Conflict of interest statement</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
</sec>
</body>
<back>
<ack><p>We would like to thank Dr. Zuguang Gu and Dr. Barbara Worst for valuable suggestions and discussion, and thank Achim Stephan for technical assistance. We also thank Prof. Dr. Roland Eils&#x00027;s group for IT infrastructure support.</p>
</ack>
<sec sec-type="supplementary-material" id="s8">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="http://journal.frontiersin.org/article/10.3389/fgene.2017.00137/full#supplementary-material">http://journal.frontiersin.org/article/10.3389/fgene.2017.00137/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Table1.XLS" id="SM1" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S1</label>
<caption><p>The score weights of single and combined features in confFuse scoring algorithm.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM2" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S2</label>
<caption><p>The verified fusion genes in the first training data.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM3" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S3</label>
<caption><p>The verified fusion genes in the second training data.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM4" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S4</label>
<caption><p>The recovery of verified fusions in deFuse, confFuse-6.5, and confFuse-8.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM5" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S5</label>
<caption><p>The recovery of verified fusions in fusionMap.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM6" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S6</label>
<caption><p>The recovery of verified fusions in soapFuse.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM7" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S7</label>
<caption><p>The experimental validations in three breast tumor samples.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM8" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S8</label>
<caption><p>The high-confidence fusions in prostate cancer samples.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM9" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S9</label>
<caption><p>The correlation of verified fusions and supporting reads in fusionMap.</p></caption></supplementary-material>
<supplementary-material xlink:href="Table1.XLS" id="SM10" mimetype="application/vnd.ms-excel" xmlns:xlink="http://www.w3.org/1999/xlink">
<label>Table S10</label>
<caption><p>The correlation of verified fusions and supporting reads in soapFuse.</p></caption></supplementary-material>
<supplementary-material xlink:href="Image1.PDF" id="SM11" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abate</surname> <given-names>F.</given-names></name> <name><surname>Zairis</surname> <given-names>S.</given-names></name> <name><surname>Ficarra</surname> <given-names>E.</given-names></name> <name><surname>Acquaviva</surname> <given-names>A.</given-names></name> <name><surname>Wiggins</surname> <given-names>C. H.</given-names></name> <name><surname>Frattini</surname> <given-names>V.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Pegasus: a comprehensive annotation and prediction tool for detection of driver gene fusions in cancer</article-title>. <source>BMC Syst. Biol.</source> <volume>8</volume>:<fpage>97</fpage>. <pub-id pub-id-type="doi">10.1186/s12918-014-0097-z</pub-id><pub-id pub-id-type="pmid">25183062</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bao</surname> <given-names>Z.-S.</given-names></name> <name><surname>Chen</surname> <given-names>H.-M.</given-names></name> <name><surname>Yang</surname> <given-names>M.-Y.</given-names></name> <name><surname>Zhang</surname> <given-names>C.-B.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name> <name><surname>Ye</surname> <given-names>W.-L.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>RNA-seq of 272 gliomas revealed a novel, recurrent PTPRZ1-MET fusion transcript in secondary glioblastomas</article-title>. <source>Genome Res.</source> <volume>24</volume>, <fpage>1765</fpage>&#x02013;<lpage>1773</lpage>. <pub-id pub-id-type="doi">10.1101/gr.165126.113</pub-id><pub-id pub-id-type="pmid">25135958</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><collab>Cancer Genome Atlas Research Network</collab></person-group> (<year>2013</year>). <article-title>Comprehensive molecular characterization of clear cell renal cell carcinoma</article-title>. <source>Nature</source> <volume>499</volume>, <fpage>43</fpage>&#x02013;<lpage>49</lpage>. <pub-id pub-id-type="doi">10.1038/nature12222</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Carrara</surname> <given-names>M.</given-names></name> <name><surname>Beccuti</surname> <given-names>M.</given-names></name> <name><surname>Cavallo</surname> <given-names>F.</given-names></name> <name><surname>Donatelli</surname> <given-names>S.</given-names></name> <name><surname>Lazzarato</surname> <given-names>F.</given-names></name> <name><surname>Cordero</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>State of art fusion-finder algorithms are suitable to detect transcription-induced chimeras in normal tissues?</article-title> <source>BMC Bioinformatics</source> <volume>14</volume>(<supplement>Suppl. 7</supplement>):<fpage>S2</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-14-S7-S2</pub-id><pub-id pub-id-type="pmid">23815381</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chmielecki</surname> <given-names>J.</given-names></name> <name><surname>Crago</surname> <given-names>A. M.</given-names></name> <name><surname>Rosenberg</surname> <given-names>M.</given-names></name> <name><surname>O&#x00027;Connor</surname> <given-names>R.</given-names></name> <name><surname>Walker</surname> <given-names>S. R.</given-names></name> <name><surname>Ambrogio</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Whole-exome sequencing identifies a recurrent NAB2-STAT6 fusion in solitary fibrous tumors</article-title>. <source>Nat. Genet.</source> <volume>45</volume>, <fpage>131</fpage>&#x02013;<lpage>132</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2522</pub-id><pub-id pub-id-type="pmid">23313954</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dawe</surname> <given-names>H. R.</given-names></name> <name><surname>Farr</surname> <given-names>H.</given-names></name> <name><surname>Portman</surname> <given-names>N.</given-names></name> <name><surname>Shaw</surname> <given-names>M. K.</given-names></name> <name><surname>Gull</surname> <given-names>K.</given-names></name></person-group> (<year>2005</year>). <article-title>The Parkin co-regulated gene product, PACRG, is an evolutionarily conserved axonemal protein that functions in outer-doublet microtubule morphogenesis</article-title>. <source>J. Cell Sci.</source> <volume>118</volume>(<issue>Pt 23</issue>), <fpage>5421</fpage>&#x02013;<lpage>5430</lpage>. <pub-id pub-id-type="doi">10.1242/jcs.02659</pub-id><pub-id pub-id-type="pmid">16278296</pub-id></citation></ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ge</surname> <given-names>H.</given-names></name> <name><surname>Liu</surname> <given-names>K.</given-names></name> <name><surname>Juan</surname> <given-names>T.</given-names></name> <name><surname>Fang</surname> <given-names>F.</given-names></name> <name><surname>Newman</surname> <given-names>M.</given-names></name> <name><surname>Hoeck</surname> <given-names>W.</given-names></name></person-group> (<year>2011</year>). <article-title>FusionMap: detecting fusion genes from next-generation sequencing data at base-pair resolution</article-title>. <source>Bioinforma Oxf. Engl</source>. <volume>27</volume>, <fpage>1922</fpage>&#x02013;<lpage>1928</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btr310</pub-id><pub-id pub-id-type="pmid">21593131</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Gu</surname> <given-names>Z.</given-names></name> <name><surname>Gu</surname> <given-names>L.</given-names></name> <name><surname>Eils</surname> <given-names>R.</given-names></name> <name><surname>Schlesner</surname> <given-names>M.</given-names></name> <name><surname>Brors</surname> <given-names>B.</given-names></name></person-group> (<year>2014</year>). <article-title>circlize implements and enhances circular visualization in R</article-title>. <source>Bioinforma Oxf. Engl</source>. <volume>30</volume>, <fpage>2811</fpage>&#x02013;<lpage>2812</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu393</pub-id><pub-id pub-id-type="pmid">24930139</pub-id></citation></ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><collab>International Cancer Genome Consortium PedBrain Tumor Project</collab></person-group> (<year>2016</year>). <article-title>Recurrent MET fusion genes represent a drug target in pediatric glioblastoma</article-title>. <source>Nat. Med.</source> <volume>22</volume>, <fpage>1314</fpage>&#x02013;<lpage>1320</lpage>. <pub-id pub-id-type="doi">10.1038/nm.4204</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jia</surname> <given-names>W.</given-names></name> <name><surname>Qiu</surname> <given-names>K.</given-names></name> <name><surname>He</surname> <given-names>M.</given-names></name> <name><surname>Song</surname> <given-names>P.</given-names></name> <name><surname>Zhou</surname> <given-names>Q.</given-names></name> <name><surname>Zhou</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>SOAPfuse: an algorithm for identifying fusion transcripts from paired-end RNA-Seq data</article-title>. <source>Genome Biol.</source> <volume>14</volume>:<fpage>R12</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2013-14-2-r12</pub-id><pub-id pub-id-type="pmid">23409703</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jones</surname> <given-names>D. T. W.</given-names></name> <name><surname>Hutter</surname> <given-names>B.</given-names></name> <name><surname>J&#x000E4;ger</surname> <given-names>N.</given-names></name> <name><surname>Korshunov</surname> <given-names>A.</given-names></name> <name><surname>Kool</surname> <given-names>M.</given-names></name> <name><surname>Warnatz</surname> <given-names>H.-J.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Recurrent somatic alterations of FGFR1 and NTRK2 in pilocytic astrocytoma</article-title>. <source>Nat. Genet.</source> <volume>45</volume>, <fpage>927</fpage>&#x02013;<lpage>932</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2682</pub-id><pub-id pub-id-type="pmid">23817572</pub-id></citation></ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ju</surname> <given-names>Y. S.</given-names></name> <name><surname>Lee</surname> <given-names>W.-C.</given-names></name> <name><surname>Shin</surname> <given-names>J.-Y.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Bleazard</surname> <given-names>T.</given-names></name> <name><surname>Won</surname> <given-names>J.-K.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>A transforming KIF5B and RET gene fusion in lung adenocarcinoma revealed from whole-genome and transcriptome sequencing</article-title>. <source>Genome Res.</source> <volume>22</volume>, <fpage>436</fpage>&#x02013;<lpage>445</lpage>. <pub-id pub-id-type="doi">10.1101/gr.133645.111</pub-id><pub-id pub-id-type="pmid">22194472</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kim</surname> <given-names>D.</given-names></name> <name><surname>Pertea</surname> <given-names>G.</given-names></name> <name><surname>Trapnell</surname> <given-names>C.</given-names></name> <name><surname>Pimentel</surname> <given-names>H.</given-names></name> <name><surname>Kelley</surname> <given-names>R.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2013</year>). <article-title>TopHat2: accurate alignment of transcriptomes in the presence of insertions, deletions and gene fusions</article-title>. <source>Genome Biol.</source> <volume>14</volume>:<fpage>R36</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2013-14-4-r36</pub-id><pub-id pub-id-type="pmid">23618408</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kumar</surname> <given-names>S.</given-names></name> <name><surname>Vo</surname> <given-names>A. D.</given-names></name> <name><surname>Qin</surname> <given-names>F.</given-names></name> <name><surname>Li</surname> <given-names>H.</given-names></name></person-group> (<year>2016</year>). <article-title>Comparative assessment of methods for the fusion transcripts detection from RNA-Seq data</article-title>. <source>Sci. Rep.</source> <volume>6</volume>:<fpage>21597</fpage>. <pub-id pub-id-type="doi">10.1038/srep21597</pub-id><pub-id pub-id-type="pmid">26862001</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Langmead</surname> <given-names>B.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2012</year>). <article-title>Fast gapped-read alignment with Bowtie 2</article-title>. <source>Nat. Methods</source> <volume>9</volume>, <fpage>357</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.1923</pub-id><pub-id pub-id-type="pmid">22388286</pub-id></citation></ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Langmead</surname> <given-names>B.</given-names></name> <name><surname>Trapnell</surname> <given-names>C.</given-names></name> <name><surname>Pop</surname> <given-names>M.</given-names></name> <name><surname>Salzberg</surname> <given-names>S. L.</given-names></name></person-group> (<year>2009</year>). <article-title>Ultrafast and memory-efficient alignment of short DNA sequences to the human genome</article-title>. <source>Genome Biol.</source> <volume>10</volume>:<fpage>R25</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2009-10-3-r25</pub-id><pub-id pub-id-type="pmid">19261174</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Li</surname> <given-names>H.</given-names></name> <name><surname>Durbin</surname> <given-names>R.</given-names></name></person-group> (<year>2009</year>). <article-title>Fast and accurate short read alignment with Burrows-Wheeler transform</article-title>. <source>Bioinforma Oxf. Engl</source>. <volume>25</volume>, <fpage>1754</fpage>&#x02013;<lpage>1760</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp324</pub-id><pub-id pub-id-type="pmid">19451168</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lilljebj&#x000F6;rn</surname> <given-names>H.</given-names></name> <name><surname>&#x000C5;gerstam</surname> <given-names>H.</given-names></name> <name><surname>Orsmark-Pietras</surname> <given-names>C.</given-names></name> <name><surname>Rissler</surname> <given-names>M.</given-names></name> <name><surname>Ehrencrona</surname> <given-names>H.</given-names></name> <name><surname>Nilsson</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>RNA-seq identifies clinically relevant fusion genes in leukemia including a novel MEF2D/CSF1R fusion responsive to imatinib</article-title>. <source>Leukemia</source> <volume>28</volume>, <fpage>977</fpage>&#x02013;<lpage>979</lpage>. <pub-id pub-id-type="doi">10.1038/leu.2013.324</pub-id><pub-id pub-id-type="pmid">24186003</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>McPherson</surname> <given-names>A.</given-names></name> <name><surname>Hormozdiari</surname> <given-names>F.</given-names></name> <name><surname>Zayed</surname> <given-names>A.</given-names></name> <name><surname>Giuliany</surname> <given-names>R.</given-names></name> <name><surname>Ha</surname> <given-names>G.</given-names></name> <name><surname>Sun</surname> <given-names>M. G. F.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>deFuse: an algorithm for gene fusion discovery in tumor RNA-Seq data</article-title>. <source>PLoS Comput. Biol.</source> <volume>7</volume>:<fpage>e1001138</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pcbi.1001138</pub-id><pub-id pub-id-type="pmid">21625565</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Mertens</surname> <given-names>F.</given-names></name> <name><surname>Johansson</surname> <given-names>B.</given-names></name> <name><surname>Fioretos</surname> <given-names>T.</given-names></name> <name><surname>Mitelman</surname> <given-names>F.</given-names></name></person-group> (<year>2015</year>). <article-title>The emerging complexity of gene fusions in cancer</article-title>. <source>Nat. Rev. Cancer</source> <volume>15</volume>, <fpage>371</fpage>&#x02013;<lpage>381</lpage>. <pub-id pub-id-type="doi">10.1038/nrc3947</pub-id><pub-id pub-id-type="pmid">25998716</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pajtler</surname> <given-names>K. W.</given-names></name> <name><surname>Witt</surname> <given-names>H.</given-names></name> <name><surname>Sill</surname> <given-names>M.</given-names></name> <name><surname>Jones</surname> <given-names>D. T. W.</given-names></name> <name><surname>Hovestadt</surname> <given-names>V.</given-names></name> <name><surname>Kratochwil</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Molecular classification of ependymal tumors across all CNS compartments, histopathological grades, and age groups</article-title>. <source>Cancer Cell</source>. <volume>27</volume>, <fpage>728</fpage>&#x02013;<lpage>743</lpage>. <pub-id pub-id-type="doi">10.1016/j.ccell.2015.04.002</pub-id><pub-id pub-id-type="pmid">25965575</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ricarte-Filho</surname> <given-names>J. C.</given-names></name> <name><surname>Li</surname> <given-names>S.</given-names></name> <name><surname>Garcia-Rendueles</surname> <given-names>M. E. R.</given-names></name> <name><surname>Montero-Conde</surname> <given-names>C.</given-names></name> <name><surname>Voza</surname> <given-names>F.</given-names></name> <name><surname>Knauf</surname> <given-names>J. A.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Identification of kinase fusion oncogenes in post-Chernobyl radiation-induced thyroid cancers</article-title>. <source>J. Clin. Invest.</source> <volume>123</volume>, <fpage>4935</fpage>&#x02013;<lpage>4944</lpage>. <pub-id pub-id-type="doi">10.1172/JCI69766</pub-id><pub-id pub-id-type="pmid">24135138</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seo</surname> <given-names>J.-S.</given-names></name> <name><surname>Ju</surname> <given-names>Y. S.</given-names></name> <name><surname>Lee</surname> <given-names>W.-C.</given-names></name> <name><surname>Shin</surname> <given-names>J.-Y.</given-names></name> <name><surname>Lee</surname> <given-names>J. K.</given-names></name> <name><surname>Bleazard</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>The transcriptional landscape and mutational profile of lung adenocarcinoma</article-title>. <source>Genome Res.</source> <volume>22</volume>, <fpage>2109</fpage>&#x02013;<lpage>2119</lpage>. <pub-id pub-id-type="doi">10.1101/gr.145144.112</pub-id><pub-id pub-id-type="pmid">22975805</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seshagiri</surname> <given-names>S.</given-names></name> <name><surname>Stawiski</surname> <given-names>E. W.</given-names></name> <name><surname>Durinck</surname> <given-names>S.</given-names></name> <name><surname>Modrusan</surname> <given-names>Z.</given-names></name> <name><surname>Storm</surname> <given-names>E. E.</given-names></name> <name><surname>Conboy</surname> <given-names>C. B.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Recurrent R-spondin fusions in colon cancer</article-title>. <source>Nature</source> <volume>488</volume>, <fpage>660</fpage>&#x02013;<lpage>664</lpage>. <pub-id pub-id-type="doi">10.1038/nature11282</pub-id><pub-id pub-id-type="pmid">22895193</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Steidl</surname> <given-names>C.</given-names></name> <name><surname>Shah</surname> <given-names>S. P.</given-names></name> <name><surname>Woolcock</surname> <given-names>B. W.</given-names></name> <name><surname>Rui</surname> <given-names>L.</given-names></name> <name><surname>Kawahara</surname> <given-names>M.</given-names></name> <name><surname>Farinha</surname> <given-names>P.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>MHC class II transactivator CIITA is a recurrent gene fusion partner in lymphoid cancers</article-title>. <source>Nature</source> <volume>471</volume>, <fpage>377</fpage>&#x02013;<lpage>381</lpage>. <pub-id pub-id-type="doi">10.1038/nature09754</pub-id><pub-id pub-id-type="pmid">21368758</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Torres-Garc&#x000ED;a</surname> <given-names>W.</given-names></name> <name><surname>Zheng</surname> <given-names>S.</given-names></name> <name><surname>Sivachenko</surname> <given-names>A.</given-names></name> <name><surname>Vegesna</surname> <given-names>R.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <name><surname>Yao</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>PRADA: pipeline for RNA sequencing data analysis</article-title>. <source>Bioinforma Oxf. Engl</source>. <volume>30</volume>, <fpage>2224</fpage>&#x02013;<lpage>2226</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu169</pub-id><pub-id pub-id-type="pmid">24695405</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Untergasser</surname> <given-names>A.</given-names></name> <name><surname>Cutcutache</surname> <given-names>I.</given-names></name> <name><surname>Koressaar</surname> <given-names>T.</given-names></name> <name><surname>Ye</surname> <given-names>J.</given-names></name> <name><surname>Faircloth</surname> <given-names>B. C.</given-names></name> <name><surname>Remm</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2012</year>). <article-title>Primer3&#x02013;new capabilities and interfaces</article-title>. <source>Nucleic Acids Res.</source> <volume>40</volume>:<fpage>e115</fpage>. <pub-id pub-id-type="doi">10.1093/nar/gks596</pub-id><pub-id pub-id-type="pmid">22730293</pub-id></citation></ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Veeraraghavan</surname> <given-names>J.</given-names></name> <name><surname>Tan</surname> <given-names>Y.</given-names></name> <name><surname>Cao</surname> <given-names>X.-X.</given-names></name> <name><surname>Kim</surname> <given-names>J. A.</given-names></name> <name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Chamness</surname> <given-names>G. C.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Recurrent ESR1-CCDC170 rearrangements in an aggressive subset of oestrogen receptor-positive breast cancers</article-title>. <source>Nat. Commun.</source> <volume>5</volume>:<fpage>4577</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms5577</pub-id><pub-id pub-id-type="pmid">25099679</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname> <given-names>X.</given-names></name> <name><surname>Bledsoe</surname> <given-names>K. L.</given-names></name> <name><surname>Graham</surname> <given-names>R. P.</given-names></name> <name><surname>Asmann</surname> <given-names>Y. W.</given-names></name> <name><surname>Viswanatha</surname> <given-names>D. S.</given-names></name> <name><surname>Lewis</surname> <given-names>J. E.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Recurrent PAX3-MAML3 fusion in biphenotypic sinonasal sarcoma</article-title>. <source>Nat. Genet.</source> <volume>46</volume>, <fpage>666</fpage>&#x02013;<lpage>668</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2989</pub-id><pub-id pub-id-type="pmid">24859338</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Weischenfeldt</surname> <given-names>J.</given-names></name> <name><surname>Simon</surname> <given-names>R.</given-names></name> <name><surname>Feuerbach</surname> <given-names>L.</given-names></name> <name><surname>Schlangen</surname> <given-names>K.</given-names></name> <name><surname>Weichenhan</surname> <given-names>D.</given-names></name> <name><surname>Minner</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Integrative genomic analyses reveal an androgen-driven somatic alteration landscape in early-onset prostate cancer</article-title>. <source>Cancer Cell</source>. <volume>23</volume>, <fpage>159</fpage>&#x02013;<lpage>170</lpage>. <pub-id pub-id-type="doi">10.1016/j.ccr.2013.01.002</pub-id><pub-id pub-id-type="pmid">23410972</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zong</surname> <given-names>F.-Y.</given-names></name> <name><surname>Fu</surname> <given-names>X.</given-names></name> <name><surname>Wei</surname> <given-names>W.-J.</given-names></name> <name><surname>Luo</surname> <given-names>Y.-G.</given-names></name> <name><surname>Heiner</surname> <given-names>M.</given-names></name> <name><surname>Cao</surname> <given-names>L.-J.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>The RNA-binding protein QKI suppresses cancer-associated aberrant splicing</article-title>. <source>PLoS Genet.</source> <volume>10</volume>:<fpage>e1004289</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1004289</pub-id><pub-id pub-id-type="pmid">24722255</pub-id></citation></ref>
</ref-list>
<glossary>
<def-list>
<title>Abbreviations</title>
<def-item><term>FPKM</term><def><p>Fragments Per Kilobase of exon per Million mapped reads.</p></def></def-item>
</def-list>
</glossary>
<fn-group>
<fn fn-type="financial-disclosure"><p><bold>Funding.</bold> This study was supported by the DKFZ-Heidelberg Center for Personalized Oncology (DKFZ-HIPO) through HIPO-017.</p>
</fn>
</fn-group>
</back>
</article>