<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Archiving and Interchange DTD v2.3 20070202//EN" "archivearticle.dtd">
<article article-type="methods-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">1084974</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.1084974</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A shortest path-based approach for copy number variation detection from next-generation sequencing data</article-title>
<alt-title alt-title-type="left-running-head">Liu et al.</alt-title>
<alt-title alt-title-type="right-running-head">
<ext-link ext-link-type="uri" xlink:href="https://doi.org/10.3389/fgene.2022.1084974">10.3389/fgene.2022.1084974</ext-link>
</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Liu</surname>
<given-names>Guojun</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/995307/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yang</surname>
<given-names>Hongzhi</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Yuan</surname>
<given-names>Xiguo</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">
<sup>&#x2a;</sup>
</xref>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>School of Statistics</institution>, <institution>Xi&#x2019;an University of Finance and Economics</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Medical Imaging Center</institution>, <institution>Xidian Group Hospital</institution>, <addr-line>Xi&#x2019;an</addr-line>, <country>China</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Hangzhou Institute of Technology</institution>, <institution>Xidian University</institution>, <addr-line>Hangzhou</addr-line>, <country>China</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/356107/overview">Piyush Pandey</ext-link>, Assam University, India</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/832518/overview">Junwei Luo</ext-link>, Henan Polytechnic University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1068283/overview">Aimin Li</ext-link>, Xi&#x2019;an University of Technology, China</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Guojun Liu, <email>yaguojun@163.com</email>; Xiguo Yuan, <email>xiguoyuan@mail.xidian.edu.cn</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Computational Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>17</day>
<month>01</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>1084974</elocation-id>
<history>
<date date-type="received">
<day>31</day>
<month>10</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>27</day>
<month>12</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Liu, Yang and Yuan.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Liu, Yang and Yuan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Copy number variation (CNV) is one of the main structural variations in the human genome and accounts for a considerable proportion of variations. As CNVs can directly or indirectly cause cancer, mental illness, and genetic disease in humans, their effective detection in humans is of great interest in the fields of oncogene discovery, clinical decision-making, bioinformatics, and drug discovery. The advent of next-generation sequencing data makes CNV detection possible, and a large number of CNV detection tools are based on next-generation sequencing data. Due to the complexity (e.g., bias, noise, alignment errors) of next-generation sequencing data and CNV structures, the accuracy of existing methods in detecting CNVs remains low. In this work, we design a new CNV detection approach, called shortest path-based Copy number variation (SPCNV), to improve the detection accuracy of CNVs. SPCNV calculates the k nearest neighbors of each read depth and defines the shortest path, shortest path relation, and shortest path cost sets based on which further calculates the mean shortest path cost of each read depth and its k nearest neighbors. We utilize the ratio between the mean shortest path cost for each read depth and the mean of the mean shortest path cost of its k nearest neighbors to construct a relative shortest path score formula that is able to determine a score for each read depth. Based on the score profile, a boxplot is then applied to predict CNVs. The performance of the proposed method is verified by simulation data experiments and compared against several popular methods of the same type. Experimental results show that the proposed method achieves the best balance between recall and precision in each set of simulated samples. To further verify the performance of the proposed method in real application scenarios, we then select real sample data from the 1,000 Genomes Project to conduct experiments. The proposed method achieves the best F1-scores in almost all samples. Therefore, the proposed method can be used as a more reliable tool for the routine detection of CNVs.</p>
</abstract>
<kwd-group>
<kwd>copy number variation</kwd>
<kwd>next-generation sequencing data</kwd>
<kwd>k nearest neighbors</kwd>
<kwd>shortest path</kwd>
<kwd>read depth</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>As one type of structural variation, copy number variation (CNV) plays an important role in the formation and development of human cancers and diseases (<xref ref-type="bibr" rid="B19">Mccarroll and Altshuler, 2007</xref>; <xref ref-type="bibr" rid="B25">Stefansson et al., 2008</xref>; <xref ref-type="bibr" rid="B3">Beroukhim et al., 2010</xref>; <xref ref-type="bibr" rid="B36">Yuan et al., 2021a</xref>). Generally, CNV is defined as a deletion or amplification of a genomic sequence that is no less than 1,000 to several megabase pairs in length compared to a reference genome (<xref ref-type="bibr" rid="B10">Freeman et al., 2006</xref>; <xref ref-type="bibr" rid="B38">Zhao et al., 2013</xref>; <xref ref-type="bibr" rid="B37">Yuan et al., 2021b</xref>). The deletion and amplification of a copy number can lead to the reorganization of the genome structure and the change of base content, which further affects the level of human gene expression (<xref ref-type="bibr" rid="B22">Sebat et al., 2004</xref>; <xref ref-type="bibr" rid="B24">Sharp et al., 2005</xref>). Studies have shown that the occurrence of some common human diseases is closely related to CNV, such as ovarian cancer (<xref ref-type="bibr" rid="B2">Adam and David, 2009</xref>; <xref ref-type="bibr" rid="B11">Fridley et al., 2012</xref>), breast cancer (<xref ref-type="bibr" rid="B28">Tchatchou and Burwinkel, 2008</xref>; <xref ref-type="bibr" rid="B14">Kumaran et al., 2017</xref>), autism (<xref ref-type="bibr" rid="B23">Sebat et al., 2007</xref>; <xref ref-type="bibr" rid="B21">Pinto et al., 2010</xref>), schizophrenia (<xref ref-type="bibr" rid="B25">Stefansson et al., 2008</xref>; <xref ref-type="bibr" rid="B26">Stone et al., 2008</xref>), etc. In this context, next-generation sequencing (NGS) technology has developed rapidly and is able to provide rich data resources for the accurate detection of CNVs in the human genome, higher resolution, and more flexible detection methods (<xref ref-type="bibr" rid="B20">Meyerson et al., 2010</xref>). However, due to various factors, such as bias, noise, and the uneven distribution of NGS data, the existing detection methods are still not accurate for CNV detection.</p>
<p>A large number of CNV detection methods have been developed around NGS data, and the vast majority of them are based on the read depth (RD) method. The basic principle of RD-based CNV detection methods is that the number of reads aligned at each position in the reference genome is proportional to the copy number at that position (<xref ref-type="bibr" rid="B33">Yoon et al., 2009</xref>). Compared with the normal region, the number of reads in the copy number amplification region is higher, and the number of reads in the copy number deletion region is lower (<xref ref-type="bibr" rid="B29">Teo et al., 2012</xref>). The RD method can use single-end sequencing reads or paired-end sequencing reads to detect CNVs. In principle, it can detect the amplified and deleted regions of CNVs of any length. In practical applications, this method is more sensitive for detecting long CNVs. Therefore, it is more suitable for detecting copy number amplified regions but cannot accurately detect variant boundaries, resulting in detection results that contain a large number of false-positive positions.</p>
<p>The basic process of RD-based methods for detecting CNVs includes: 1) Reading the alignment and extracting RDs; 2) Preprocessing the RDs; 3) Building the detection model (statistical model, machine learning algorithm, etc.); 4) Selecting a reasonable threshold strategy and predicting CNVs. Based on the above workflow, some well-known RD-based CNV methods have been proposed, mainly including FREEC (<xref ref-type="bibr" rid="B4">Boeva et al., 2012</xref>), CNV-LOF (<xref ref-type="bibr" rid="B36">Yuan et al., 2021a</xref>), CNVnator (<xref ref-type="bibr" rid="B1">Abyzov et al., 2011</xref>), BIC-seq2 (<xref ref-type="bibr" rid="B31">Xi et al., 2016</xref>), SeqCNV (<xref ref-type="bibr" rid="B6">Chen et al., 2017</xref>), CNV_IFTV (<xref ref-type="bibr" rid="B37">Yuan et al., 2021b</xref>), and iCopyDAV (<xref ref-type="bibr" rid="B8">Dharanipragada et al., 2018</xref>). FREEC obtains the normalized read count profile by using GC-content or mappability profiles and employs a lasso-based algorithm to produce a smooth copy number profile that predicts genotype status for each genomic segment. FREEC is more sensitive to copy number gain regions than loss regions, and the detection results have a large number of false-positive positions, resulting in lower precision. CNV-LOF performs a segmentation procedure on the RD profiles to obtain consecutive and non-overlapping RD segment profiles. On this basis, a cyclic binary segmentation (CBS) algorithm (<xref ref-type="bibr" rid="B30">Venkatraman and Olshen, 2007</xref>) is performed on each segment to divide each one into a set of segments. CNV-LOF utilizes the idea of a local outlier factor to assign an outlier score for each RD segment. Based on the anomaly score profile, it predicts CNVs using a boxplot procedure. It is not sensitive to the detection of loss regions, and its performance is not well balanced between recall and precision. CNVnator calibrates the GC content to normalize the RD profile and uses a mean-shift approach to segment the RD profile to predict CNVs. CNVnator is able to detect a large number of long CNVs, the vast majority of which are false-positive events. Therefore, it achieves low precision, especially in the detection of low-purity samples. BIC-seq2 normalizes the RD profile at the nucleotide level and uses the bayesian information criterion to predict CNVs. While its performance is balanced between recall and precision, it has low precision in detecting high-purity samples. SeqCNV extracts the RD signal from paired samples, establishes a maximum penalized likelihood estimation model, and selects a threshold interval to predict CNVs. It is sensitive to short CNV detection and is not suitable for the detection of low-purity samples. CNV_IFTV utilizes the isolation forest algorithm to calculate an anomaly score for each RD, smooths the anomaly score profile using a total variation model, and uses the anomaly score to fit a gamma distribution to predict CNVs. The difference between the established statistical model and the actual distribution of RDs affects the accuracy of the CNV_IFTV detection. iCopyDAV automatically estimates bin size, calibrates GC-content and mappability bias using the median method and mappability score file, and performs segmentation using the CBS algorithm to predict CNVs. It is suitable for testing high purity and medium coverage samples. In general application scenarios, the above methods can effectively detect a large number of CNVs. However, their performance is uneven in the detection of samples of different purity.</p>
<p>With consideration of the above issues, we propose a new approach in this work to accurately detect CNVs using NGS data from the whole genome. The method is called shortest path-based CNV (SPCNV). The SPCNV calculates the k nearest neighbors of each RD and defines the shortest path, the shortest path relation, and the shortest path cost sets. Based on these three types of shortest path sets, we calculate the mean shortest path cost of each RD and its k nearest neighbors. A relative shortest path score equation is then built using the ratio between the mean shortest path cost of each RD and the mean of mean shortest path cost of its k nearest neighbors, which can calculate a score for each RD (<xref ref-type="bibr" rid="B27">Tang et al., 2002</xref>). Based on the score profile, a boxplot program is used to predict CNVs(<xref ref-type="bibr" rid="B39">Zijlstra et al., 2007</xref>). The main contributions of the proposed method are as follows: 1) According to the basic principle of the RD method, the copy number gain and loss correspond to larger and smaller RDs compared to normal RDs, respectively. The two types of RDs have fewer ratios among all RDs. Therefore, we treat the two types of RDs as outliers and successfully transform a traditional outlier detection method into a CNV detection method. 2) By extracting two features, the RD ratio and the difference between adjacent RD ratios, we can observe the difference between RDs from a global and local perspective, which is conducive to detecting isolated variants and local small cluster variants. 3) The proposed method uses the difference between the shortest path of each RD and the average shortest path of its k nearest neighbors to identify CNVs, which is beneficial for identifying a local cluster of insignificant variations. As the traditional machine learning method only relies on the distance between each RD to distinguish the difference between them, it cannot detect a small local cluster variation because their differences are very small.</p>
<p>The remainder of this work is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> includes the workflow of SPCNV, data preprocessing, construction of relative shortest path score formula, and the forecasting of CNVs. <xref ref-type="sec" rid="s3">Section 3</xref> presents the simulation data and real data experiments and analyzes and discusses the experimental results. <xref ref-type="sec" rid="s4">Section 4</xref> addresses the shortcomings of the work and presents future work ideas.</p>
</sec>
<sec sec-type="materials|methods" id="s2">
<title>2 Method and materials</title>
<sec id="s2-1">
<title>2.1 Overview of SPCNV</title>
<p>SPCNV is an RD-based CNV detection method that is suitable for the detection of a single sample. <xref ref-type="fig" rid="F1">Figure 1</xref> shows the workflow of SPCNV in detail, which consists of the following main five steps: 1) The sequenced donor samples (Fastq) and reference genome (Fasta) are prepared for input; 2) The reads are aligned to the reference genome using BWA (<xref ref-type="bibr" rid="B15">Li and Durbin, 2010</xref>) to generate sequence alignment files (SAM), which are converted to BAM format using SAMtools (<xref ref-type="bibr" rid="B16">Li et al., 2009</xref>); 3) The data is preprocessed. This step mainly includes read count (RC) profile extraction with SAMtools, bin definition (<xref ref-type="bibr" rid="B35">Yuan et al., 2018</xref>), anomalous bin removal, obtaining the read depth (RD) profiles, GC bias calibration, noise removal, and the dimension transformation of the RD profiles; 4) The relative shortest path score is built and assigned for each RD; 5) Based on the score profile, a boxplot is utilized to predict CNVs. The SPCNV software is developed in R and Python languages. It can be downloaded from <ext-link ext-link-type="uri" xlink:href="https://github.com/gj-123/SPCNV/releases">https://github.com/gj-123/SPCNV/releases</ext-link> and is easy to install and use after reading the user manual. In the following section, each step in the workflow of SPCNV is analyzed and discussed in detail.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>The workflow of SPCNV.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g001.tif"/>
</fig>
</sec>
<sec id="s2-2">
<title>2.2 Data preprocessing</title>
<p>The sequenced donor samples are aligned to the reference genome using the BWA tool, which generates sequence alignment files in the SAM format. The SAM files are further converted into binary sequence alignment files in BAM format using SAMtools. The read count (RC) profiles are then extracted with SAMtools from the BAM files. We define a sliding window procedure (bin) (<xref ref-type="bibr" rid="B35">Yuan et al., 2018</xref>), with which the RC profiles are divided continuously and are non-overlapping to generate the RD profiles. This process is described using Eq. <xref ref-type="disp-formula" rid="e1">1</xref>.<disp-formula id="e1">
<mml:math id="m1">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(1)</label>
</disp-formula>where <inline-formula id="inf1">
<mml:math id="m2">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:msub>
<mml:mi>D</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the RD value of the n-th bin, which is equal to the mean RC in a bin. Since the reference genome contains a large number of &#x201c;N" positions, reads aligning to these positions will result in RCs equal to 0 that will be mistaken for a loss at that position. Therefore, if a bin contains &#x201c;N" positions, we remove the bin from the genome sequence (<xref ref-type="bibr" rid="B36">Yuan et al., 2021a</xref>). Due to the complexity of the human genome, the distribution of GC-content is uneven, which can lead to the misidentification of copy number deletions. The GC content of each bin is calibrated using the median method (<xref ref-type="bibr" rid="B33">Yoon et al., 2009</xref>). Factors such as alignment errors and biases can cause the resulting sequencing data to be noisy. RD signal noise will seriously affect the detection accuracy, which is a key step in the detection of CNV. Here, the total variation model (<xref ref-type="bibr" rid="B7">Condat, 2013</xref>; <xref ref-type="bibr" rid="B9">Duan et al., 2013</xref>) is used to smooth and segment the RD profile to generate an RD segment (RDS) profile, which is represented by Eq. <xref ref-type="disp-formula" rid="e2">2</xref>.<disp-formula id="e2">
<mml:math id="m3">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>S</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>2</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>where <inline-formula id="inf2">
<mml:math id="m4">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>n</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the value of the n-th read depth segment, which is equal to the mean of all RDs contained in this segment. The RDS profile is converted to two-dimensional space to generate the <inline-formula id="inf3">
<mml:math id="m5">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> profile (<xref ref-type="bibr" rid="B17">Liu et al., 2020</xref>), which is composed of the RDS ratio and differences between adjacent RDS ratios and is expressed by Eq. <xref ref-type="disp-formula" rid="e3">3</xref>.<disp-formula id="e3">
<mml:math id="m6">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo>,</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(3)</label>
</disp-formula>where <inline-formula id="inf4">
<mml:math id="m7">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>X</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the value of the i-th RD ratio, <inline-formula id="inf5">
<mml:math id="m8">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Y</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents the difference between the i-th RD ratio and its adjacent RD ratios.</p>
<p>This transformation process can detect differences in RD from two perspectives. The first dimension can approximately reflect the copy number status corresponding to each RD from a global perspective. The second dimension can approximately reflect the difference between an RDS and its adjacent RDSs from a local perspective. By extracting two features of RD, the proposed method can more easily discover globally isolated and local small cluster variants. At the same time, this step also provides an effective data platform for constructing the relative shortest path score in the next section.</p>
</sec>
<sec id="s2-3">
<title>2.3 Establishment of relative shortest path score</title>
<p>Based on <inline-formula id="inf6">
<mml:math id="m9">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> profile, we construct a relative shortest path score (RSPS) to evaluate the degree of anomaly of each RDS. Here, we regard each element in <inline-formula id="inf7">
<mml:math id="m10">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> as an object represented by <italic>o</italic>. The RSPS fully reflects the closeness between an object and its surrounding objects and is highly suitable for application in CNV detection scenarios. The RSPS of an object depends on the ratio between the object&#x2019;s shortest path and the mean of its k nearest neighbors&#x2019; shortest paths. Some related basic concepts and definitions must be introduced before giving the definition of RSPS, which mainly include the k-distance of an object, the k-distance neighborhood of an object (<xref ref-type="bibr" rid="B5">Breunig et al., 2000</xref>), the shortest path set, the shortest path relation set, and the shortest path cost set.</p>
<p>
<statement content-type="definition" id="Definition_1">
<label>Definition 1</label>
<p>The <italic>k</italic>-distance of an object <italic>o,</italic> which is defined using Eq. <xref ref-type="disp-formula" rid="e4">4</xref>.<disp-formula id="e4">
<mml:math id="m11">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>o</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <inline-formula id="inf8">
<mml:math id="m12">
<mml:mrow>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mi>o</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents euclidean distance between object <inline-formula id="inf9">
<mml:math id="m13">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> and object <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:msup>
<mml:mi>o</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>\</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf11">
<mml:math id="m15">
<mml:mrow>
<mml:msup>
<mml:mi>o</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
</mml:mrow>
</mml:math>
</inline-formula> indicates that the k-th object closest to <italic>o</italic> is sorted in ascending order, <inline-formula id="inf12">
<mml:math id="m16">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the <italic>k</italic>-distance of an object <italic>o</italic>. Here, k is a positive integer.</p>
</statement>
</p>
<p>
<statement content-type="definition" id="Definition_2">
<label>Definition 2</label>
<p>The <italic>k</italic>-distance neighborhood of an object <italic>o</italic> is a collection of objects whose distance from <italic>o</italic> is less than or equal to <inline-formula id="inf13">
<mml:math id="m17">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, which is defined using Eq. <xref ref-type="disp-formula" rid="e5">5</xref>.<disp-formula id="e5">
<mml:math id="m18">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>&#x7c;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>D</mml:mi>
<mml:msup>
<mml:mi>S</mml:mi>
<mml:mo>&#x2032;</mml:mo>
</mml:msup>
<mml:mo>\</mml:mo>
<mml:mrow>
<mml:mfenced open="{" close="}" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
<mml:mo>,</mml:mo>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x2264;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(5)</label>
</disp-formula>where <inline-formula id="inf14">
<mml:math id="m19">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the <italic>k</italic>-distance neighborhood of an object <italic>o</italic> and a collection of objects, and the distance between each object in the collection and o is not greater than <inline-formula id="inf15">
<mml:math id="m20">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>The shortest path set (SPS) of an object <italic>o</italic> is composed of object <italic>o</italic> and the <italic>k</italic>-distance neighborhood of object <italic>o</italic>, which are connected to form a path with the shortest distance. The shortest path relation set (SPRS) of an object <italic>o</italic> is defined as the edge between two objects on the shortest path. The shortest path cost set (SPCS) of an object <italic>o</italic> is defined as the distance between two objects on the shortest path. <xref ref-type="statement" rid="Algorithm_1">Algorithm 1</xref> describes the calculation process of SPS(<italic>o</italic>), SPRS(<italic>o</italic>), and SPCS(<italic>o</italic>) in detail.</p>
<p>The following example is used to clearly explain the SPS, SPRS, and SPCS calculation process of an object. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, there are a total of 10 objects {1, 2, 3, 4, 5, 6, 7, 8, 9, 10}, and their corresponding values are {(0.5, 0.5), (0.13, 0.78), (0.18, 0.13), (0.2, 0.4), (0.32, 0.38), (0.47, 0.44), (0.56, 0.6), (0.6, 0.56), (0.87, 0.5), and (0.91, 0.93)}. We calculate the SPS, SPRS, and SPCS of object one here, where the value of k is set to 5. Eq. <xref ref-type="disp-formula" rid="e4">4</xref> and <xref ref-type="disp-formula" rid="e5">5</xref> are employed to obtain the five nearest neighbors of object 1 (6, 7, 8, 5, 4). The distance from object one to object six is the smallest. According to <xref ref-type="statement" rid="Algorithm_1">Algorithm 1</xref>, SPS (1) is equal to {1, 6}, SPRS (1) is equal to {(1, 6)}, and SPC S(1) is equal to {0.07}. The distance between objects 7, 8, 5, 4, and objects 1, six is calculated to obtain a minimum distance, and the distance between objects one and seven is the smallest. Similarly, SPS (1) is equal to {1, 6, 7}, SPRS (1) is equal to {(1, 6), (1, 7)}, and SPCS (1) is equal to {0.07, 0.12}. The procedure ends when each neighbor of object one finds its closest object from SPS(1). Finally, SPS (1) is equal to {1, 6, 7, 8, 5, 4}, SPRS (1) is equal to {(1, 6), (1, 7), (7, 8), (6, 5), (5, 4)}, and SPCS (1) is equal to {0.07, 0.12, 0.06, 0.16, 0.12}. As shown in <xref ref-type="fig" rid="F2">Figure 2</xref>, the red line segments form the final shortest path.</p>
</statement>
</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>An example of calculating the SPS, SPRS, and SPCS of object 1.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g002.tif"/>
</fig>
<p>
<statement content-type="algorithm" id="Algorithm_1">
<label>Algorithm 1</label>
<p>Calculate SPS, SPRS and SPCS of object <italic>o</italic>. <inline-graphic xlink:href="fgene-13-1084974-fx1.tif"/>
</p>
</statement>
</p>
<p>Based on the above definitions, the mean shortest path cost of an object <italic>o</italic> (<inline-formula id="inf26">
<mml:math id="m31">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>) is defined by Eq. <xref ref-type="disp-formula" rid="e6">6</xref>.<disp-formula id="e6">
<mml:math id="m32">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mrow>
<mml:munderover>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>k</mml:mi>
</mml:munderover>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf27">
<mml:math id="m33">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the i-th element in <inline-formula id="inf28">
<mml:math id="m34">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>C</mml:mi>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>.If the value of <inline-formula id="inf29">
<mml:math id="m35">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is larger, the distance between <italic>o</italic> and its k nearest neighbors is sparser; if it is smaller, they are closer together.</p>
<p>After estimating the mean shortest path cost of all objects, we further construct the relative shortest path score (RSPS) to measure the degree of deviation between object <italic>o</italic> and its k nearest neighbors, which is defined by Eq. <xref ref-type="disp-formula" rid="e7">7</xref>.<disp-formula id="e7">
<mml:math id="m36">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>&#x22c5;</mml:mo>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:munder>
<mml:mstyle displaystyle="true">
<mml:mo>&#x2211;</mml:mo>
</mml:mstyle>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:munder>
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(7)</label>
</disp-formula>where <inline-formula id="inf30">
<mml:math id="m37">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the mean shortest path cost of <italic>o</italic>, <inline-formula id="inf31">
<mml:math id="m38">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the mean shortest path cost of one of its nearest neighbors, <inline-formula id="inf32">
<mml:math id="m39">
<mml:mrow>
<mml:mfenced open="|" close="|" separators="|">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:math>
</inline-formula> represents the number of elements in <inline-formula id="inf33">
<mml:math id="m40">
<mml:mrow>
<mml:msub>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>d</mml:mi>
<mml:mi>i</mml:mi>
<mml:mi>s</mml:mi>
<mml:mi>t</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula>, <inline-formula id="inf34">
<mml:math id="m41">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> represents the ratio between <inline-formula id="inf35">
<mml:math id="m42">
<mml:mrow>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>C</mml:mi>
<mml:mi>m</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> and the mean of mean shortest path cost of its k nearest neighbors. If the value of <inline-formula id="inf36">
<mml:math id="m43">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>o</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mrow>
</mml:math>
</inline-formula> is larger, the distance between <italic>o</italic> and its k nearest neighbors is sparser; if it is smaller, they are closer together. This means that the higher the RSPS of an object, the more likely the object is a CNV.</p>
</sec>
<sec id="s2-4">
<title>2.4 Prediction of CNVs</title>
<p>Although we have evaluated the relative shortest path score for each object, it is not yet possible to distinguish abnormal objects from normal objects. This step is critical in CNV detection, and a reasonable threshold selection strategy will significantly improve the accuracy of the detection results. Traditional threshold selection strategies mainly include: 1) Fitting a statistical model using the score profile to evaluate the significance of each object and using hypothesis testing to predict CNVs; 2) Selecting an empirical value as a threshold to identify CNVs. The limitation of the first strategy is that due to the bias and noise of the sequencing data, the actual distribution of the data and the fitted model are quite different, resulting in inaccurate detection results. The limitation of the second strategy is that a fixed threshold can effectively identify abnormal objects in some scenarios. However, the performance of the method may drop significantly in certain application scenarios. Considering the above two scenarios, we use a boxplot to determine thresholds based on the score files. This method does not require an assumption that the score profile obeys a certain distribution in advance and is able to dynamically determine thresholds according to different score files. Here, we use Eq. <xref ref-type="disp-formula" rid="e8">8</xref> to estimate a threshold for judging anomalies.<disp-formula id="e8">
<mml:math id="m44">
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3bb;</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:mrow>
<mml:mfenced open="(" close=")" separators="|">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:mrow>
</mml:math>
<label>(8)</label>
</disp-formula>where <inline-formula id="inf37">
<mml:math id="m45">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mn>3</mml:mn>
</mml:msub>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents upper quartile of RSPS, <inline-formula id="inf38">
<mml:math id="m46">
<mml:mrow>
<mml:mi>R</mml:mi>
<mml:mi>S</mml:mi>
<mml:mi>P</mml:mi>
<mml:msub>
<mml:mi>S</mml:mi>
<mml:msub>
<mml:mi>Q</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> represents lower quartile of RSPS, <inline-formula id="inf39">
<mml:math id="m47">
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents multiple of interquartile range of RSPS, <inline-formula id="inf40">
<mml:math id="m48">
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> represents the maximum value of the inner limit of RSPS, which is used as the threshold. If an object&#x2019;s RSPS is greater than the threshold <inline-formula id="inf41">
<mml:math id="m49">
<mml:mrow>
<mml:mi>&#x3c4;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, it is considered to be a CNV. After predicting CNVs, we further differentiate the types of CNV (gain and loss). If the RD of an object is greater than or equal to the average RD of all normal objects, it is regarded as a gain; if it is less, it is regarded as a loss.</p>
</sec>
</sec>
<sec sec-type="results|discussion" id="s3">
<title>3 Results and discussion</title>
<p>Along with the establishment of SPCNV, the design of a reasonable experimental scheme is crucial for verifying the effectiveness of the proposed method. In this study, the experimental component was divided into simulation and real data experiments. In the simulation data experiments, the performance of the proposed method was compared with four well-known similar methods (CNV-LOF, FREEC, CNVnator, and BIC-seq2) from five perspectives: recall, precision, F1-score, the number of gain and loss detections, and sensitivity of different size CNV detection. In order to verify the performance of the proposed method in real data applications, the above four comparison methods were also selected for comparison with the proposed method. The experimental data was a set of real human sequencing samples from the 1000 Genomes Project. Some previous studies have tested these samples and saved the test results to the Database of Genomic Variants (DGV), which was used as ground truth to calculate recall, precision, and F1-score for each method.</p>
<sec id="s3-1">
<title>3.1 Application of simulation data</title>
<p>IntSIM (<xref ref-type="bibr" rid="B34">Yuan et al., 2017</xref>) simulation software was adopted to generate the simulation data sets. Before using the software, the two key parameters of sample tumor purity (TP) and sequencing coverage (SC) were set from 0.2 to 0.8 and 5x, respectively. To ensure the reliability of the test results, 50 samples were generated under each set of configuration conditions and the average of which was used as the test result. There were six gains and eight losses embedded in each sample, whose lengths range from 10,000 to 50,000&#xa0;bp.</p>
<p>Based on the simulated datasets, the performance of SPCNV and four other alignment methods (CNV-LOF, FREEC, CNVnator, and BIC-seq2) were tested by calculating their recall, precision, and F1-scores. Recall is defined as the number of correctly detected CNV events divided by the total number of simulated CNV events, which can be calculated by the ground truth file (<xref ref-type="bibr" rid="B18">Magi et al., 2013</xref>). Precision is defined as the number of correctly detected CNV events divided by the total number of detected CNV events (<xref ref-type="bibr" rid="B18">Magi et al., 2013</xref>). The F1-score is defined as the harmonic mean of recall and precision. The experimental results of each method are depicted in <xref ref-type="fig" rid="F3">Figure 3</xref>, where the three-performance metrics (recall, precision, and F1-score) of each method are compared in the four simulation sample sets. According to the overall trend, the performance of the majority of methods improves with increasing tumor purity. For example, the recall of CNVnator is close to 0.2 when the tumor purity is equal to 0.2, and its recall exceeds 0.7 when the tumor purity is equal to 0.8. Correspondingly, its F1-score increases from 0.18 to 0.58. Among the five methods, SPCNV achieves the best F1-score in each dataset. BIC-seq2 obtains the lowest F1-score at a purity equal to 0.8, but its F1-scores are better than other three methods (CNV-LOF, FREEC, and CNVnator) at a purity equal to 0.6. The above situations indicate that BIC-seq2 can provide base level resolution and detect a large number of CNVs, but its precision is very low when detecting high-purity samples. BIC-seq2 is not suitable for the detection of high-purity samples. The F1-score of FREEC is better than other three methods (CNV-LOF, BIC-seq2, and CNVnator) when the tumor purity is equal to 0.8. When detecting low and medium purity samples, its performance is relatively balanced between recall and precision. When FREEC detects high-purity samples, its F1-score is superior to the other three comparison methods, but its recall is significantly higher than precision, which indicates that its performance in detecting high-purity samples is uneven. The F1-scores of CNV-LOF are better than other three methods (BIC-seq2, FREEC, and CNVnator) when the tumor purity is equal to 0.2 and 0.4, which indicates that it is suitable for detecting low and medium purity samples. Its advantage is to obtain high precision when detecting low and medium purity samples. In contrast, its recall rate is low. CNVnator obtains the lowest F1-scores at purities equal to 0.2, 0.4, and 0.6, because the precision of CNVnator is the lowest in all sample sets. CNVnator detects a large number of long CNVs, most of which are false-positive positions. In terms of recall, SPCNV achieves the best recall, except when the tumor purity is equal to 0.6. In terms of precision, SPCNV obtains the best precision in each sample set. Overall, the SPCNV gets the best trade-off between recall and precision in each simulation sample set.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>The performance of the five methods is compared in terms of recall, precision, and F1-score across four sets of simulation samples. Black curves indicate that the F1-score levels are harmonic means of recall and precision ranging from 0.1 to 0.9. The equations on the left and right sides of the comma represent the tumor purity (TP) and sequencing coverage (SC), respectively.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g003.tif"/>
</fig>
<p>Based on the above analysis and discussion, we further analyze the performance of each method in detecting gains and losses. <xref ref-type="fig" rid="F4">Figure 4</xref> details the performance of each method in detecting gains and losses in the four datasets. In general, the number of gains detected by each method is more than losses. The performance of SPCNV in detecting the gain and loss is relatively balanced compared to the other four methods. While CNV-LOF detects the most gains in each sample set, it obtains the least losses in the vast majority of cases. In most cases, FREEC detects far more gains than losses. CNVnator detects the least gains in each sample set and obtains fewer losses than most methods. BIC-seq2 detects the most losses in three sets of samples, indicating that it is suitable for detecting losses. In summary, SPCNV detects more gains and losses than most methods in each simulation sample set, which shows that its performance is robust in gain and loss detection.</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>The performance of the five methods is compared in terms of the number of detected gains and losses across four sets of simulation samples. The equations on the left and right sides of the comma represent the tumor purity (TP) and sequencing coverage (SC), respectively.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g004.tif"/>
</fig>
<p>As a supplement to the above experiments, we further analyzed the sensitivity of each method to detect CNV of different lengths. <xref ref-type="fig" rid="F5">Figure 5</xref> details the sensitivity of five methods to detect three CNV lengths (10&#xa0;kb, 20&#xa0;kb and 50&#xa0;kb) in four sets of samples. The sensitivity is defined as the ratio between the number of correctly detected CNVs and the number of simulated CNVs. The proposed method achieves the best sensitivity when the TP is equal to 0.2 and the CNV length is equal to 50&#xa0;kb, and ranks second when CNV sizes are equal to 10kb and 20&#xa0;kb. BIC-seq2 and CNV-LOF get the best sensitivity when the TP is equal to 0.2 and CNV sizes are equal to 10&#xa0;kb and 20&#xa0;kb. CNV-LOF achieves lower sensitivity at CNV sizes equal to 10&#xa0;kb and 50&#xa0;kb than other three comparison methods (SPCNV, FREEC and BIC-seq2). BIC-seq2 achieves lower sensitivity at CNV sizes equal to 20&#xa0;kb and 50&#xa0;kb than SPCNV. The proposed method gets the best sensitivity when the TP is equal to 0.4 and CNV sizes are equal to 20&#xa0;kb and 50&#xa0;kb, and ranks second when CNV size is equal to 10&#xa0;kb. CNV-LOF and BIC-seq2 get the best sensitivity when CNV size is equal to 10&#xa0;kb, and achieves lower sensitivity at CNV sizes equal to 20&#xa0;kb and 50&#xa0;kb than the proposed method. The proposed method gets the best sensitivity when the TP is equal to 0.6 and CNV size is equal to 20&#xa0;kb, and ranks second when CNV sizes are equal to 10&#xa0;kb and 50&#xa0;kb. At the same time, CNV-LOF and BIC-seq2 get the best sensitivity when CNV size is equal to 10&#xa0;kb and 50&#xa0;kb, respectively. CNV-LOF achieves lower sensitivity at CNV sizes equal to 20&#xa0;kb and 50&#xa0;kb than SPCNV. BIC-seq2 achieves lower sensitivity at CNV sizes equal to 10&#xa0;kb and 20&#xa0;kb than SPCNV. The proposed method gets the best sensitivity when the TP is equal to 0.8 and CNV size are equal to 20&#xa0;kb and 50&#xa0;kb, and ranks second when CNV size is equal to 10&#xa0;kb. CNV-LOF gets the best sensitivity when CNV size is equal to 10&#xa0;kb, and achieves the lowest sensitivity at CNV sizes equal to 20&#xa0;kb and 50&#xa0;kb. In general, the proposed method performs better than the other four comparison methods in detecting CNVS of different sizes.</p>
<fig id="F5" position="float">
<label>FIGURE 5</label>
<caption>
<p>The sensitivity of five methods at the three CNV length levels under four sets of simulation configurations. The equations on the left and right sides of the comma represent the tumor purity (TP) and sequencing coverage (SC), respectively.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g005.tif"/>
</fig>
</sec>
<sec id="s3-2">
<title>3.2 Application of real data</title>
<p>In order to verify the performance of the proposed method in real application scenarios, we use six real data samples (NA12878, NA12891, NA12892, NA19238, NA19239, and NA19240) from the 1,000 Genomes Project, which can be downloaded for free from <ext-link ext-link-type="uri" xlink:href="http://www.internationalgenome.org/">http://www.internationalgenome.org/</ext-link>. Some of the test results for these samples are recorded in the DGV, which is considered the ground truth for calculating the recall, precision, and F1-scores for each method. Similarly, we select the four methods (CNV-LOF, FREEC, CNVnator, and BIC-seq2) of the above experiments to compare with the proposed method. The experimental results are shown in <xref ref-type="fig" rid="F6">Figure 6</xref>. SPCNV obtains the highest F1-scores among the five samples and ranks second in the NA19238 sample. CNV-LOF obtains the highest F1-score in the NA19238 sample and ranks second in F1-score among the other five samples. BIC-seq2 does not detect the correct CNVs in NA12878, NA12891, and NA12892, and its F1-scores rank third in NA19238, NA19239, and NA19240. The performance of FREEC and CNVnator is relatively close, and their F1-scores are ranked third and fourth in NA12878, NA12891, and NA12892 and in NA12878, NA12891, and NA12892, respectively. In terms of recall, FREEC achieves the best recall three times, CNVnator has the best recall twice, and CNV-LOF achieves the best recall once. In terms of precision, SPCNV obtains the best precision in five of the six samples, and CNV-LOF has the best precision once. Overall, SPCNV has obvious advantages over the other four methods in terms of precision and F1-score.</p>
<fig id="F6" position="float">
<label>FIGURE 6</label>
<caption>
<p>The performance of the five methods is compared in terms of recall, precision, and F1-score across six real data samples. Black curves indicate F1-score levels that are harmonic means of recall and precision ranging from 0.1 to 0.9.</p>
</caption>
<graphic xlink:href="fgene-13-1084974-g006.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion and conclusion</title>
<p>In this work, a new method called SPCNV was proposed to detect CNVs using NGS data. SPCNV was developed based on the RD method and could be used to detect a single genome-wide sample. The proposed SPCNV method effectively removed abnormal bins, calibrated GC-content bias, and denoised and transformed the dimension of the read depth. Based on the preprocessed RD profile, it then computed the k nearest neighbors for each object. On this basis, we constructed the shortest path set, the shortest path relation set, and the shortest path cost set for each object, by which the mean shortest path cost was defined. At the same time, we computed the mean shortest path cost for each object and used the ratio between the mean shortest path cost of each object and the mean of the mean shortest path cost of its k nearest neighbors to construct a relative shortest path score. Based on the relative shortest path scores for each object, CNVs were predicted using boxplots. Both simulation and real data experiments were then carried out to verify the performance of the proposed method. In the simulation data experiments, the performance of the proposed method was evaluated from five aspects (recall, precision, F1-score, the number of gain and loss, and sensitivity of detection of CNV with different sizes). The experimental results showed that the proposed method achieved the best balance between recall and precision, gain and loss, as well as CNV of different sizes, respectively. In real data applications, the proposed method achieved the best F1-scores in most samples, indicating that its performance was reliable and effective in real application scenarios.</p>
<p>Traditional CNV detection methods generally assume in advance that the RDs obey a statistical model, use the model to calculate a <italic>p</italic>-value for each read depth, and use hypothesis testing methods or select a fixed threshold to predict CNVs. Compared with the traditional methods, the proposed method has three different characteristics, which are summarized as follows. 1) Compared with traditional density-based methods, using objects to construct shortest paths can effectively identify a small cluster of local variants, which have little difference but are isolated relative to all objects. 2) We treat a CNV as an outlier and effectively transform the outlier detection method into a CNV detection method. 3) Extracting the RD ratio and difference of the RD ratio provides a global and local perspective to capture the difference of the copy number corresponding to each RD, which is more conducive to the detection of local single isolated CNVs and local small clusters of CNVs.</p>
<p>Although the performance of the proposed method meets the detection needs to a certain extent, there are still some shortcomings that require improvement. At this stage, the resolution of the proposed method must be further improved. In the next step, we will extract the read depth and split reads to further enhance its resolution (<xref ref-type="bibr" rid="B32">Ye et al., 2009</xref>; <xref ref-type="bibr" rid="B12">Jiang et al., 2012</xref>). The selection of the number of nearest neighbors is a key step that can affect the accuracy of the detection results. In this study, the selection of this parameter is based on reference to previous studies (<xref ref-type="bibr" rid="B5">Breunig et al., 2000</xref>; <xref ref-type="bibr" rid="B13">Jin et al., 2006</xref>). While the performance of the method is good in most application scenarios, it may not be suitable for some individual cases. In future work, the selection method of this parameter will be improved to realize automatic optimization. In addition, the functions of the proposed method need to be further expanded for application to more scenarios. For example, analyzing the biological significance of CNVs and mapping oncogenes, which can provide strong support for targeted drug development and clinical treatment.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="http://www.internationalgenome.org/">http://www.internationalgenome.org/</ext-link>.</p>
</sec>
<sec id="s6">
<title>Author contributions</title>
<p>GL participated in the design of the algorithms, the entire framework of CNV detection, the design of experiments and wrote the manuscript. HY helped revise the manuscript and participated in the design of experiments. XY directed the whole work. All authors read the final manuscript and agreed on its contents for submission.</p>
</sec>
<sec sec-type="COI-statement" id="s7">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s8">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s9">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.1084974/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2022.1084974/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="Table1.DOCX" id="SM1" mimetype="application/DOCX" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Abyzov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Urban</surname>
<given-names>A. E.</given-names>
</name>
<name>
<surname>Snyder</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gerstein</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>CNVnator: An approach to discover, genotype, and characterize typical and atypical CNVs from family and population genome sequencing</article-title>. <source>Genome Res.</source> <volume>21</volume>, <fpage>974</fpage>&#x2013;<lpage>984</lpage>. <pub-id pub-id-type="doi">10.1101/gr.114876.110</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Adam</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>David</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Copy number variations and cancer</article-title>. <source>Genome Med.</source> <volume>1</volume>, <fpage>62</fpage>. <pub-id pub-id-type="doi">10.1186/gm62</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beroukhim</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mermel</surname>
<given-names>C. H.</given-names>
</name>
<name>
<surname>Porter</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Raychaudhuri</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Donovan</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>The landscape of somatic copy-number alteration across human cancers</article-title>. <source>Nature</source> <volume>463</volume>, <fpage>899</fpage>&#x2013;<lpage>905</lpage>. <pub-id pub-id-type="doi">10.1038/nature08822</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boeva</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Zinovyev</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Bleakley</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Vert</surname>
<given-names>J. P.</given-names>
</name>
<name>
<surname>Janoueix-Lerosey</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Delattre</surname>
<given-names>O.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Control-FREEC: A tool for assessing copy number and allelic content using next-generation sequencing data</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>268</fpage>&#x2013;<lpage>269</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btq635</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Breunig</surname>
<given-names>M. M.</given-names>
</name>
<name>
<surname>Kriegel</surname>
<given-names>H. P.</given-names>
</name>
<name>
<surname>Ng</surname>
<given-names>R. T.</given-names>
</name>
<name>
<surname>Sander</surname>
<given-names>J. r.</given-names>
</name>
</person-group> (<year>2000</year>). <article-title>Lof: Identifying density-based local outliers</article-title>. <source>Sigmod Rec.</source> <volume>29</volume>, <fpage>93</fpage>&#x2013;<lpage>104</lpage>. <pub-id pub-id-type="doi">10.1145/335191.335388</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Cao</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gelowani</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Xu</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>SeqCNV: A novel method for identification of copy number variations in targeted next-generation sequencing data</article-title>. <source>BMC Bioinforma.</source> <volume>18</volume>, <fpage>147</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-017-1566-3</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Condat</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>A direct algorithm for 1-D total variation denoising</article-title>. <source>IEEE Signal Process. Lett.</source> <volume>20</volume>, <fpage>1054</fpage>&#x2013;<lpage>1057</lpage>. <pub-id pub-id-type="doi">10.1109/lsp.2013.2278339</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dharanipragada</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Vogeti</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Parekh</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>iCopyDAV: Integrated platform for copy number variations-Detection, annotation and visualization</article-title>. <source>PLoS One</source> <volume>13</volume>, <fpage>e0195334</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0195334</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Duan</surname>
<given-names>J. b.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J. G.</given-names>
</name>
<name>
<surname>Deng</surname>
<given-names>H. W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y. P.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>CNV-TV: A robust method to discover copy number variation from short sequencing reads</article-title>. <source>BMC Bioinforma.</source> <volume>14</volume>, <fpage>150</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-14-150</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Freeman</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Perry</surname>
<given-names>G. H.</given-names>
</name>
<name>
<surname>Feuk</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Redon</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Mccarroll</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Altshuler</surname>
<given-names>D. M.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>Copy number variation: New insights in genome diversity</article-title>. <source>Genome Res.</source> <volume>16</volume>, <fpage>949</fpage>&#x2013;<lpage>961</lpage>. <pub-id pub-id-type="doi">10.1101/gr.3677206</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Fridley</surname>
<given-names>B. L.</given-names>
</name>
<name>
<surname>Chalise</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Tsai</surname>
<given-names>Y.-Y.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Vierkant</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Larson</surname>
<given-names>M. C.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Germline copy number variation and ovarian cancer survival</article-title>. <source>Front. Genet.</source> <volume>3</volume>, <fpage>142</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2012.00142</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Y. D.</given-names>
</name>
<name>
<surname>Brudno</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Prism: Pair-read informed split-read mapping for base-pair level detection of insertion, deletion and structural variants</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>2576</fpage>&#x2013;<lpage>2583</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts484</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jin</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Tung</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Han</surname>
<given-names>J. W.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Ranking outliers using symmetric neighborhood relationship</article-title>. <source>Adv. Knowl. Discov. Data Min.</source> <volume>3918</volume>, <fpage>577</fpage>&#x2013;<lpage>593</lpage>. <pub-id pub-id-type="doi">10.1007/11731139_68</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kumaran</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Cass</surname>
<given-names>C. E.</given-names>
</name>
<name>
<surname>Graham</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Mackey</surname>
<given-names>J. R.</given-names>
</name>
<name>
<surname>Hubaux</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Lam</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Germline copy number variations are associated with breast cancer risk and prognosis</article-title>. <source>Sci. Rep.</source> <volume>7</volume>, <fpage>14621</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-017-14799-7</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Durbin</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Fast and accurate long-read alignment with Burrows&#x2013;Wheeler transform</article-title>. <source>Bioinformatics</source> <volume>26</volume>, <fpage>589</fpage>&#x2013;<lpage>595</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp698</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Handsaker</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wysoker</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fennell</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ruan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Homer</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>The sequence alignment/map format and SAMtools</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2078</fpage>&#x2013;<lpage>2079</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp352</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>G. J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Yuan</surname>
<given-names>X. G.</given-names>
</name>
<name>
<surname>Wei</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Rkdoscnv: A local kernel density-based approach to the detection of copy number variations by using next-generation sequencing data</article-title>. <source>Front. Genet.</source> <volume>11</volume>, <fpage>569227</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2020.569227</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Magi</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tattini</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Cifola</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>D&#x2019;Aurizio</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Benelli</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mangano</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Excavator: Detecting copy number variants from whole-exome sequencing data</article-title>. <source>Genome Biol.</source> <volume>14</volume>, <fpage>R120</fpage>. <pub-id pub-id-type="doi">10.1186/gb-2013-14-10-r120</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McCarroll</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Altshuler</surname>
<given-names>D. M.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Copy-number variation and association studies of human disease</article-title>. <source>Nat. Genet.</source> <volume>39</volume>, <fpage>S37</fpage>&#x2013;<lpage>S42</lpage>. <pub-id pub-id-type="doi">10.1038/ng2080</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meyerson</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gabriel</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Getz</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Advances in understanding cancer genomes through second-generation sequencing</article-title>. <source>Nat. Rev. Genet.</source> <volume>11</volume>, <fpage>685</fpage>&#x2013;<lpage>696</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2841</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pinto</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Pagnamenta</surname>
<given-names>A. T.</given-names>
</name>
<name>
<surname>Klei</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Anney</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Merico</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Regan</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Functional impact of global rare copy number variation in autism spectrum disorders</article-title>. <source>Nature</source> <volume>466</volume>, <fpage>368</fpage>&#x2013;<lpage>372</lpage>. <pub-id pub-id-type="doi">10.1038/nature09146</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sebat</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lakshmi</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Troge</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Alexander</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lundin</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2004</year>). <article-title>Large-scale copy number polymorphism in the human genome</article-title>. <source>Science</source> <volume>305</volume>, <fpage>525</fpage>&#x2013;<lpage>528</lpage>. <pub-id pub-id-type="doi">10.1126/science.1098918</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sebat</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lakshmi</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Malhotra</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Troge</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lese-Martin</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Walsh</surname>
<given-names>T.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Strong association of de novo copy number mutations with autism</article-title>. <source>Science</source> <volume>316</volume>, <fpage>445</fpage>&#x2013;<lpage>449</lpage>. <pub-id pub-id-type="doi">10.1126/science.1138659</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sharp</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Locke</surname>
<given-names>D. P.</given-names>
</name>
<name>
<surname>McGrath</surname>
<given-names>S. D.</given-names>
</name>
<name>
<surname>Cheng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Bailey</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Vallente</surname>
<given-names>R. U.</given-names>
</name>
<etal/>
</person-group> (<year>2005</year>). <article-title>Segmental duplications and copy-number variation in the human genome</article-title>. <source>Am. J. Hum. Genet.</source> <volume>77</volume>, <fpage>78</fpage>&#x2013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1086/431652</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stefansson</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Rujescu</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Cichon</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Pietilainen</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Andres</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Stacy</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>Large recurrent microdeletions associated with schizophrenia</article-title>. <source>Nature</source> <volume>455</volume>, <fpage>232</fpage>&#x2013;<lpage>236</lpage>. <pub-id pub-id-type="doi">10.1038/nature07229</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Stone</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>O&#x27;Donovan</surname>
<given-names>M. C.</given-names>
</name>
<name>
<surname>Gurling</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Kirov</surname>
<given-names>G. K.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Rare chromosomal deletions and duplications increase risk of schizophrenia</article-title>. <source>Nature</source> <volume>455</volume>, <fpage>237</fpage>&#x2013;<lpage>241</lpage>. <pub-id pub-id-type="doi">10.1038/nature07239</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="confproc">
<person-group person-group-type="author">
<name>
<surname>Tang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Fu</surname>
<given-names>A. W.-c.</given-names>
</name>
<name>
<surname>Cheung</surname>
<given-names>D. W.</given-names>
</name>
</person-group> (<year>2002</year>). &#x201c;<article-title>Enhancing effectiveness of outlier detections for low density patterns</article-title>,&#x201d; in <conf-name>Proceeding of the Advances in Knowledge Discovery and Data Mining. 6th Pacific-Asia Conference, PAKDD</conf-name>, <conf-date>January 2002</conf-date>, <fpage>535</fpage>&#x2013;<lpage>548</lpage>. <comment>2336</comment>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tchatchou</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Burwinkel</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Chromosome copy number variation and breast cancer risk</article-title>. <source>Cytogenet. Genome Res.</source> <volume>123</volume>, <fpage>183</fpage>&#x2013;<lpage>187</lpage>. <pub-id pub-id-type="doi">10.1159/000184707</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Teo</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Pawitan</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ku</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Chia</surname>
<given-names>K. S.</given-names>
</name>
<name>
<surname>Salim</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Statistical challenges associated with detecting copy number variations with next-generation sequencing</article-title>. <source>Bioinformatics</source> <volume>28</volume>, <fpage>2711</fpage>&#x2013;<lpage>2718</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bts535</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Venkatraman</surname>
<given-names>E. S.</given-names>
</name>
<name>
<surname>Olshen</surname>
<given-names>A. B.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>A faster circular binary segmentation algorithm for the analysis of array CGH data</article-title>. <source>Bioinformatics</source> <volume>23</volume>, <fpage>657</fpage>&#x2013;<lpage>663</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btl646</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Xi</surname>
<given-names>R. B.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>Y. C.</given-names>
</name>
<name>
<surname>Kim</surname>
<given-names>T. M.</given-names>
</name>
<name>
<surname>Park</surname>
<given-names>P. J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Copy number analysis of whole-genome data using BIC-seq2 and its application to detection of cancer susceptibility variants</article-title>. <source>Nucleic Acids Res.</source> <volume>44</volume>, <fpage>6274</fpage>&#x2013;<lpage>6286</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkw491</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Schulz</surname>
<given-names>M. H.</given-names>
</name>
<name>
<surname>Long</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Apweiler</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ning</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Pindel: A pattern growth approach to detect break points of large deletions and medium sized insertions from paired-end short reads</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2865</fpage>&#x2013;<lpage>2871</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp394</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yoon</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Xuan</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Makarov</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Sebat</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Sensitive and accurate detection of copy number variants using read depth of coverage</article-title>. <source>Genome Res.</source> <volume>19</volume>, <fpage>1586</fpage>&#x2013;<lpage>1592</lpage>. <pub-id pub-id-type="doi">10.1101/gr.092981.109</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>X. G.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L. Y.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>IntSIM: An integrated simulator of next-generation sequencing data</article-title>. <source>IEEE Trans. Biomed. Eng.</source> <volume>64</volume>, <fpage>441</fpage>&#x2013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.1109/TBME.2016.2560939</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>X. G.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>J. Y.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Duan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Condel: Detecting copy number variation and genotyping deletion zygosity from single tumor samples using sequence data</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinforma.</source> <volume>17</volume>, <fpage>1141</fpage>&#x2013;<lpage>1153</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2018.2883333</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>X. G.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J. P.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Xi</surname>
<given-names>J. n.</given-names>
</name>
</person-group> (<year>2021a</year>). <article-title>A local outlier factor-based detection of copy number variations from NGS data</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinforma.</source> <volume>18</volume>, <fpage>1811</fpage>&#x2013;<lpage>1820</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2019.2961886</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yuan</surname>
<given-names>X. G.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Xi</surname>
<given-names>J. N.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Shang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<etal/>
</person-group> (<year>2021b</year>). <article-title>CNV_IFTV: An isolation forest and total variation-based detection of CNVs from short-read sequencing data</article-title>. <source>IEEE/ACM Trans. Comput. Biol. Bioinforma.</source> <volume>18</volume>, <fpage>539</fpage>&#x2013;<lpage>549</lpage>. <pub-id pub-id-type="doi">10.1109/TCBB.2019.2920889</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Jia</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Zhao</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Computational tools for copy number variation (CNV) detection using next-generation sequencing data: Features and perspectives</article-title>. <source>BMC Bioinforma.</source> <volume>14</volume>, <fpage>S1</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-14-S11-S1</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zijlstra</surname>
<given-names>W. P.</given-names>
</name>
<name>
<surname>van der Ark</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Sijtsma</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Outlier detection in test and questionnaire data</article-title>. <source>Multivar. Behav. Res.</source> <volume>42</volume>, <fpage>531</fpage>&#x2013;<lpage>555</lpage>. <pub-id pub-id-type="doi">10.1080/00273170701384340</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>