<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Pediatr.</journal-id>
<journal-title>Frontiers in Pediatrics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Pediatr.</abbrev-journal-title>
<issn pub-type="epub">2296-2360</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fped.2023.1203289</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Pediatrics</subject>
<subj-group>
<subject>Review</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Artificial intelligence-based approaches for the detection and prioritization of genomic mutations in congenital surgical diseases</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author"><name><surname>Lin</surname><given-names>Qiongfen</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/2276769/overview"/></contrib>
<contrib contrib-type="author"><name><surname>Tam</surname><given-names>Paul Kwong-Hang</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref><uri xlink:href="https://loop.frontiersin.org/people/117881/overview" /></contrib>
<contrib contrib-type="author" corresp="yes"><name><surname>Tang</surname><given-names>Clara Sze-Man</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
<xref ref-type="corresp" rid="cor1">&#x002A;</xref><uri xlink:href="https://loop.frontiersin.org/people/231340/overview" /></contrib>
</contrib-group>
<aff id="aff1"><label><sup>1</sup></label><addr-line>Department of Surgery, Li Ka Shing Faculty of Medicine</addr-line>, <institution>The University of Hong Kong</institution>, <addr-line>Hong Kong SAR</addr-line>, <country>China</country></aff>
<aff id="aff2"><label><sup>2</sup></label><addr-line>Faculty of Medicine</addr-line>, <institution>Macau University of Science and Technology</institution>, <addr-line>Macau, Macau SAR</addr-line>, <country>China</country></aff>
<aff id="aff3"><label><sup>3</sup></label><addr-line>Dr Li Dak-Sum Research Centree</addr-line>, <institution>The University of Hong Kong - Karolinska Institutet Collaboration in Regenerative Medicine</institution>, <addr-line>Hong Kong, Hong Kong SAR</addr-line>, <country>China</country></aff>
<author-notes>
<fn fn-type="edited-by"><p><bold>Edited by:</bold> Antonino Morabito, University of Florence, Italy</p></fn>
<fn fn-type="edited-by"><p><bold>Reviewed by:</bold> Jesper Eisfeldt, Karolinska University Hospital, Sweden Michele Callea, University of Florence, Italy</p></fn>
<corresp id="cor1"><label>&#x002A;</label><bold>Correspondence:</bold> Clara Sze-Man Tang <email>claratang@hku.hk</email></corresp>
</author-notes>
<pub-date pub-type="epub"><day>01</day><month>08</month><year>2023</year></pub-date>
<pub-date pub-type="collection"><year>2023</year></pub-date>
<volume>11</volume><elocation-id>1203289</elocation-id>
<history>
<date date-type="received"><day>10</day><month>04</month><year>2023</year></date>
<date date-type="accepted"><day>17</day><month>07</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Lin, Tam and Tang.</copyright-statement>
<copyright-year>2023</copyright-year><copyright-holder>Lin, Tam and Tang</copyright-holder><license license-type="open-access" xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the <ext-link ext-link-type="uri" xlink:href="http://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution License (CC BY)</ext-link>. The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>Genetic mutations are critical factors leading to congenital surgical diseases and can be identified through genomic analysis. Early and accurate identification of genetic mutations underlying these conditions is vital for clinical diagnosis and effective treatment. In recent years, artificial intelligence (AI) has been widely applied for analyzing genomic data in various clinical settings, including congenital surgical diseases. This review paper summarizes current state-of-the-art AI-based approaches used in genomic analysis and highlighted some successful applications that deepen our understanding of the etiology of several congenital surgical diseases. We focus on the AI methods designed for the detection of different variant types and the prioritization of deleterious variants located in different genomic regions, aiming to uncover susceptibility genomic mutations contributed to congenital surgical disorders.</p>
</abstract>
<kwd-group>
<kwd>artificial intelligence</kwd>
<kwd>congenital surgical diseases</kwd>
<kwd>variant detection</kwd>
<kwd>variant prioritization</kwd>
<kwd>bioinformatics</kwd>
</kwd-group><contract-num rid="cn001">T12C-714/14-R, T12-712/21-R</contract-num><contract-num rid="cn002">17113320, 17113420</contract-num><contract-num rid="cn003">PR-HKU-, 0819344, 09201436</contract-num><contract-sponsor id="cn001">Theme-Based Research Scheme</contract-sponsor><contract-sponsor id="cn002">General Research Fund</contract-sponsor><contract-sponsor id="cn003">Health and Medical Research Fund</contract-sponsor><counts>
<fig-count count="2"/>
<table-count count="1"/><equation-count count="0"/><ref-count count="45"/><page-count count="0"/><word-count count="0"/></counts><custom-meta-wrap><custom-meta><meta-name>section-at-acceptance</meta-name><meta-value>Pediatric Surgery</meta-value></custom-meta></custom-meta-wrap>
</article-meta>
</front>
<body><sec id="s1" sec-type="intro"><label>1.</label><title>Introduction</title>
<p>Congenital disorders, also known as congenital abnormalities or disabilities, are the leading causes of infant morbidity and mortality. Congenital surgical diseases refer to those medical conditions present at birth that require surgical intervention as the first-line treatment. Myriad factors, including genetic mutations, chromosomal abnormalities, and environmental factors such as toxins or virus infection can cause these conditions. Many of these congenital surgical diseases have been shown to have a strong genetic basis. For example, approximately 10&#x0025;&#x2013;30&#x0025; of patients with congenital heart disease (CHD), the most common congenital anomaly that affects around 1&#x0025; of newborns, may have an identified genetic cause (<xref ref-type="bibr" rid="B1">1</xref>, <xref ref-type="bibr" rid="B2">2</xref>). Chromosomal anomalies (e.g., trisomy 21 and 22q11.2 deletion) and mutations in genes such as <italic>GATA4</italic>, <italic>NOTCH1</italic>, <italic>NKX2-5</italic> and <italic>TBX1</italic> dysregulating cardiac morphogenesis and differentiation have been identified in individuals with CHD (<xref ref-type="bibr" rid="B2">2</xref>). In addition, common regulatory variants and rare mutations also predispose to an increased risk of less common surgical disorders, such as Hirschsprung disease and biliary atresia (<xref ref-type="bibr" rid="B3">3</xref>).</p>
<p>In the past decade, the advancement in next-generation sequencing (NGS) has revolutionized precision medicine, shifting the paradigm of genetic diagnosis toward big data analytics. Now, researchers are able to elucidate the genetic etiology of congenital diseases by analyzing massive omics data generated from DNA, RNA and epigenetic sequencing. Although genomic analysis has been confirmed to be a potent approach for identifying disease-causal variants, the detection and prioritization of these variants predisposed to diseases from a mass amount of data is still a barrier for researchers to tackle with. AI affords from the tremendous amount of data remains challenging. AI fills in this research gap by offering compelling solutions to big data genomic analysis in three major aspects: (i) detection of high-confidence genomic mutations from various genomic data; (ii) predicting the functional impact of these variants on protein structure or functions or regulatory elements; and (iii) prioritizing disease-causing variants in patients.</p>
<p>AI is a technic acted by machines to mimic human intelligence. In computer science, AI is defined as the study of &#x201C;intelligent agents&#x201D;. It can deal with complicated problems by intelligently searching through different relevant datasets, excavating the hidden patterns of the existing features, formulating prediction models, and giving the best solution (<xref ref-type="bibr" rid="B4">4</xref>).</p>
<p>There are multiple subfields of AI, including machine learning (ML), deep learning (DL), and natural language processing (<xref ref-type="fig" rid="F1">Figure&#x00A0;1</xref>). ML serves to build AI-driven applications using supervised or unsupervised learning methods (<xref ref-type="bibr" rid="B5">5</xref>, <xref ref-type="bibr" rid="B6">6</xref>). Supervised learning uses labeled training data to train the ML model by learning the patterns and relationships between the input features. The trained model is then used to predict the labels of the new unlabeled testing data. Supervised learning algorithms, including Random Forest, Na&#x00EF;ve Bayes, and Support Vector Machines (SVM), are mostly classification-and regression-based. The classification algorithm finds functions that help categorize the data into classes based on the input labels and is mostly applicable for binary or categorical data with discrete values. The regression algorithm predicts output labels based on the association between dependent and independent variables and is mostly used for predicting continuous data. On the other hand, unsupervised learning trains models with unlabeled data to explore hidden patterns in the input. Clustering and dimensionality reduction are common techniques for unsupervised learning methods, such as Hidden Markov models (HMMs) and k-means clustering. DL is the emerging machine learning subfield that trains models with massive data and various complex supervised or unsupervised algorithms. DL involves the use of multi-layer artificial neural networks to learn the complicated structures and patterns in the data. Among the DL algorithms, convolutional neural networks (CNN) and recurrent neural networks (RNN) are the most frequently used (<xref ref-type="bibr" rid="B7">7</xref>, <xref ref-type="bibr" rid="B8">8</xref>). For comprehensive information on the details of these ML algorithms, interested readers can refer to other reviews (<xref ref-type="bibr" rid="B5">5</xref>, <xref ref-type="bibr" rid="B6">6</xref>) that extend beyond the scope of genomic data applications discussed here.</p>
<fig id="F1" position="float"><label>Figure 1</label>
<caption><p>Schematic diagram of AI algorithms.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fped-11-1203289-g001.tif"/>
</fig>
<p>With the introduction of these advanced AI-based, especially DL-based, detection and prediction tools, countless disease-causing variants were identified and prioritized from big genomic data, dramatically enhancing our understanding of the etiology of numerous congenital surgical diseases and promoting the uptake of this new evidence into clinical practice. In this review, we will focus on how AI assists the genomic analysis of congenital surgical diseases by improving the performance of detecting and prioritizing candidate disease-causing mutations.</p>
</sec>
<sec id="s2"><label>2.</label><title>Application of AI models in the identification of genetic variants</title>
<p>Variant calling is the critical process to detect genetic variations from DNA sequencing data. Before this process, the alignment of sequencing reads to the reference genome is required, then genetic variants are detected by comparing the differences in base sequence between the aligned reads and the reference genome. To detect high-quality genomic variants sensitively and specifically, numerous tools have been developed using different AI models (Bayesian, Random Forest, CNN, etc.). The majority of these tools are established to identify single-nucleotide variations (SNVs), small insertions and deletions (indels), and copy number variants (CNVs), as these types of variation are the dominant sources of a genomic mutation linked to disease. Similarly, rare and novel CNVs can also be called from traditional SNP (single-nucleotide polymorphisms) array data using AI models trained on data with known copy number information.</p>
<sec id="s2a"><label>2.1.</label><title>Detection of SNVs and indels</title>
<p>As the major variant types, SNVs and indels could be detected by plenty of variant calling tools, among which Genome Analysis Toolkit (GATK), DeepVariants, and FreeBayes are frequently used. GATK is the most widely used programming framework for analyzing DNA sequencing data and for the discovery of SNVs and indels. It applies various machine learning methods like logistic regression, HMM, and Na&#x00EF;ve Bayes classification to reduce base errors and capture high-quality variants. For example, in the variant quality score recalibration (VQSR) process to filter low-quality variants, GATK trained on multiple variant annotations (e.g., genotype qualities, depth of coverage, mapping qualities, and local sequence context, etc.) of high-confident known variants like HapMap genotypes and Omni 2.5 genotypes for 1,000 Genomes samples (<xref ref-type="bibr" rid="B9">9</xref>). It then uses the trained model to assign a well-calibrated variant quality score to each variant in a callset and refines the callset to a desired high level of truth sensitivity (<xref ref-type="bibr" rid="B10">10</xref>&#x2013;<xref ref-type="bibr" rid="B12">12</xref>).</p>
<p>Unlike GATK, another deep learning-based variant caller called DeepVariants accurately identifies genetic variants using a single deep CNN model trained with known genotypes instead of the combination of multiple statistical models. Using the Inception architecture, the CNN model calculates the genotype likelihoods for each site using a pileup image of the reference genome and sequenced reads around each candidate variant (<xref ref-type="bibr" rid="B13">13</xref>). FreeBayes, a Bayesian variant calling tool, uses a haplotype-based method to read short haplotypes directly from sequencing data. It offers many advantages for variants detection compared to approaches that manipulate a single site simultaneously. To maintain semantic consistency between the candidate variants, the haplotype-based method will assess all categories of alleles in the same sequencing context simultaneously, improving the detection utilities and accuracy (<xref ref-type="bibr" rid="B14">14</xref>).</p>
</sec>
<sec id="s2b"><label>2.2.</label><title>Detection of CNVs</title>
<p>Another variant type, CNVs, is the variant that exhibits differences in the number of copies in specific DNA segments, specifically manifested as duplication or deletion of a particular size of DNA fragments (<xref ref-type="bibr" rid="B15">15</xref>). A popular CNV caller, PennCNV, uses HMM algorithm to detect CNVs from intensity data generated by high-resolution SNP arrays. The parameters of HMM models were first optimized by training using the Baum-Welch algorithm on large CNV regions from a large set of training samples. The optimized HMM then models the observed intensity data as a mixture of normal distributions, incorporated with the Log R Ratio (LRR) and B Allele Frequency (BAF) values for each SNP in the genome to predicted copy number states (e.g., 0, 1, 2, or more copies) at each genomic location. Next, the detected CNVs will be validated using the posterior probabilities of each copy number state calculated by the Bayesian algorithm with the use of pedigree information to obtain reliable CNVs (<xref ref-type="bibr" rid="B16">16</xref>). The same framework was further extended and adapted for calling CNVs from whole genome sequencing (WGS) data in PennCNV-Seq. Another DL-based tool, DeepCNV, aims for CNV validation instead of CNV calling. It attempts to replace human visual examination in order to reduce the false positive rate of CNVs, centering around the CNVs called by PennCNV. DeepCNV is constructed by a hybrid deep neural network architecture consisting of a deep convolutional neural network (CNN) and a deep fully connected neural network (DNN). It can deal with both image data and summary statistics output from PennCNV, using CNN and DNN algorithms respectively. This tool has completely changed the ability of CNV studies and can trim the raw CNV calls into reliable CNV sets with high effectiveness and efficiency (<xref ref-type="bibr" rid="B17">17</xref>).</p>
<p>In the new era of high-throughput technology, various tools emerged for identifying CNVs from NGS data using AI models. Based on a machine-learning approach, CN-Learn accurately detects high-confidence CNVs by aggregating multiple CNV-detected methods (CANOES, CODEX, CLAMMS, and XHMM) from exome sequencing data. Caller-specific and genomic features such as GC content, CNV concordance, and CNV size were obtained from multiple CNV callers and further used as the training dataset for a Random Forest classifier, eventually used to distinguish true or false positive calls for the identified CNVs (<xref ref-type="bibr" rid="B18">18</xref>). Similarly, CNV-JACG is developed with a random forest model for Judging the Accuracy of CNVs and Genotyping using paired-end WGS data. CNV-JACG is trained on 21 distinct features characterizing true CNV regions, including 13 features characterizing the breakpoints of CNVs, 6 features of the region encompassed by the CNV, and 2 features related to the variants called within the CNV region. After training, the model learns to determine true and false CNVs and make predictions on the input dataset, calling real CNVs (<xref ref-type="bibr" rid="B19">19</xref>).</p>
<p>CNVs have been reported to have a high impact on congenital surgical diseases. For example, it has been reported that around 3&#x0025;&#x2013;25&#x0025; of the CHD cases harbored rare pathogenic CNVs that could produce unproperly working proteins (<xref ref-type="bibr" rid="B2">2</xref>). To access the contribution of <italic>de novo</italic> CNVs in the pathogenesis of sporadic CHD, Glessner, J. T. et al. applied PennCNV and XHMM (exome hidden Markov model) for the detection of high-confident <italic>de novo</italic> CNVs from the genotyping array and whole exome sequencing (WES) data respectively (<xref ref-type="bibr" rid="B20">20</xref>). CNVs detected <italic>in silico</italic> were then validated experimentally using digital droplet PCR. Ultimately, they confirmed a significant increase in CNV burden in CHD cases compared with healthy controls (<xref ref-type="bibr" rid="B21">21</xref>).</p>
<p>Tetralogy of Fallot (TOF) is the most common subtype of CHD, characterized by pulmonary stenosis, ventricular septal defect, overriding aorta and hypertrophy of the right ventricle (<xref ref-type="bibr" rid="B22">22</xref>, <xref ref-type="bibr" rid="B23">23</xref>). A WGS study on 146 Chinese nonsyndromic TOF parent-offspring trios CNV-JACG for the identification of high-confidence CNVs (&#x003E;50&#x2005;bp) from the WGS data. The study identified 16 <italic>de novo</italic> CNVs in 14 TOF patients, accounting for 9.6&#x0025; in the Chinese TOF cohort, which is higher than that in the general population (<xref ref-type="bibr" rid="B24">24</xref>). CNV analysis on Hirschsprung disease (HSCR), also known as congenital intestinal aganglionosis, identified a novel candidate gene, NRG3, with an increased burden of intronic CNVs (both deletions and duplications) in patients. Furthermore, the CNV analysis also revealed the differential genetic architecture in relation to CNVs, such that syndromic HSCR was associated with longer CNVs whereas isolated HSCR were found to have an increased burden of shorter CNVs (<xref ref-type="bibr" rid="B25">25</xref>).</p>
<p>Biliary atresia (BA) is a rare pediatric hepatobiliary disorder with multifactorial etiology. It is characterized by progressive fibro-inflammatory obstruction of the bile duct. The exact cause of BA is still unknown, but it is thought to be caused by both genetic and environmental factors. Cheng et al. detected 29 BA-private CNVs from SNP array data of BA patients and controls using PennCNV, Birdseye and iPattern. By exploring the interconnectivity of CNVs, SNPs and genetic networks in BA patients, they observed a significant enrichment in the immune-inflammatory pathway for genes associated with these BA-associated CNVs (<xref ref-type="bibr" rid="B26">26</xref>).</p>
</sec>
</sec>
<sec id="s3"><label>3.</label><title>Application of AI models in variant prioritization</title>
<p>Generally, the critical process of genomic analysis includes variant detection and variant annotation. Variants could be annotated with multiple variant features, like their associated gene symbol, protein consequence of nucleotide change, allele frequency, etc., among which deleterious prediction is the important term. With the predicted deleterious score, one could easily prioritize potentially damaging causative variants, facilitating the clinical interpretation of variants and thus contributing significantly to the study of congenital diseases.</p>
<sec id="s3a"><label>3.1.</label><title>Prioritizing deleterious mutations in the coding region</title>
<p>Combined Annotation-Dependent Depletion (CADD) is the most widely used annotation tool to predict the deleteriousness of short variants (SNVs and indels) in genetic studies of both monogenic and complex diseases. It applies a machine learning model to aggregate diverse annotations, including evolutionary conservation metrics from other annotated tools (phastCons scores, GERP, and phyloP), regulatory information and functional prediction, into a single, comprehensive measure, including evolutionary conservation, regulatory information, functional prediction score for each variant. Using the SVM algorithm, the model is trained on a set of known pathogenic and benign variants, learning to discriminate between these two classes based on the input annotations with high precision and accuracy for all kinds of variants like missense, splice, and frameshift variants (<xref ref-type="bibr" rid="B27">27</xref>, <xref ref-type="bibr" rid="B28">28</xref>).</p>
<p>In contrast, another tool Rare Exome Variant Ensemble Learner (REVEL), is designed only to predict the pathogenicity of missense variants. Similar to CADD, REVEL is an ensemble method integrated with 13 other prediction tools: MutPred, FATHMM, VEST, PolyPhen, SIFT, PROVEAN, MutationAssessor, MutationTaster, LRT, GERP, SiPhy, phyloP, and phastCons. It is trained by Random Forest using a dataset of known pathogenic and rare neutral missense variants to predict the potential effect of the query variants. As reported, REVEL has better performance on pathogenicity prediction of missense variants than other ensemble methods: MetaSVM, MetaLR, KGGSeq, Condel, CADD, DANN, and Eigen and thus widely adopted for predicting <italic>in silico</italic> damaging effect (PP3) as supporting evidence of pathogenicity in ClinGen Expect specifications in variant interpretation (e.g., hearing loss Familial Hypercholesterolemia) (<xref ref-type="bibr" rid="B29">29</xref>&#x2013;<xref ref-type="bibr" rid="B31">31</xref>).</p>
<p>AI-based variant annotation has been instrumental in the genetic analysis of rare congenital surgical diseases. In a WGS study of a Chinese cohort with TOF, Tang et al. extracted potential rare damaging variants by the damaging Phred-scaled CADD scores; thereby identified 6 TOF patients with ultra-rare damaging variants in 3 known TOF genes (<italic>KDR</italic>, <italic>FLT4</italic> and <italic>NOTCH1</italic>). It also pointed out novel biological pathways and developmental hotspots relevant to the dysregulation of cardiac development in TOF through enrichment analysis (<xref ref-type="bibr" rid="B24">24</xref>). Page et al. called variants using GATK and defined likely pathogenic nonsynonymous variants with a scaled CADD score&#x2009;&#x2265;&#x2009;20, highlighting the increased burden of <italic>NOTCH1</italic> mutations in TOF (<xref ref-type="bibr" rid="B32">32</xref>). Likewise, a trio-based WES study on BA identified rare, deleterious <italic>de novo</italic> or biallelic variants in liver-expressed ciliary genes in 31.5&#x0025; (28/89) of the BA patients with the help of the CADD, SIFT and PolyPhen2. They found that these rare deleterious variants in liver-expressed ciliary genes were associated with a significant two-fold increased risk of BA, underlying the potential disease mechanism of BA led by the malformation and dysfunction of cilia (<xref ref-type="bibr" rid="B33">33</xref>).</p>
</sec>
<sec id="s3b"><label>3.2.</label><title>Prioritizing variants that may lead to alternative splicing</title>
<p>Alternative splicing is regulated by an extensive protein-RNA interaction network involving cis-elements within the pre-mRNA and trans-acting factors that bind to these cis-elements. It is a crucial regulator of gene expression, with around 15&#x0025; of disease-causal mutations predicted to alter mRNA splicing (<xref ref-type="bibr" rid="B34">34</xref>). Disruption of splicing (for example, exon skipping and intron retention) would result in aberrant proteins that don&#x0027;t work correctly. Nowadays, numerous tools have been developed to predict the effects of splice variants, emphasizing whether variants in the splice regions can potentially lead to the loss or gain of the splice donor or splice acceptor.</p>
<p>SpliceAI uses an ultra-deep CNN model to computationally predict the effects of genetic variants on splicing based on the sequence of the pre-mRNA transcript. SpliceAI trained on the dataset from GENCODE (an integrated annotation of gene features) and the RNA-seq data Genotype-Tissue Expression (GTEx). Training on the GTEx RNA-seq dataset conduces to enhance the sensitivity of splicing-altering variation detection, particularly for detecting deep intronic splicing variants. Given a genetic variation, SpliceAI generates a couple of scores for the effects on acceptor/donor gain and acceptor/donor gain (<xref ref-type="bibr" rid="B35">35</xref>). Similar to SpliceAI, MMSplice (modular modeling of splicing) is a neural network-based model to predict the effects of variants on exon skipping, splice site choice, splicing efficiency, and pathogenicity. It consists of six modules scoring sequences from different genomic regions, wherein the donor and acceptor modules are trained using GENCODE annotation features, while the exon modules (exon 5&#x2019; and exon 3&#x2019; modules) and intron modules (intron 5&#x2019; and intron 3&#x2019; modules) are trained using massively parallel reporter assays (MPRAs) experiment, based on different module architectures. These six modules are combined with a linear model to score the variant effects on exon skipping, alternative donor/acceptor site, and splicing efficiency separately. Furthermore, it integrated with a logistic regression model to predict variant pathogenicity. For each input variant, MMSplice would output several scores, including (1) a main score that exhibits the effect of the variant on the inclusion level, (2) a pathogenicity score that shows the potential pathogenic effect, (3) an efficiency score that demonstrates the variant effect on splicing efficiency of the exon and (4) several scores for the effects of the acceptor/donor/exon/intron according to the reference allele and alternative allele (<xref ref-type="bibr" rid="B36">36</xref>).</p>
<p>In the context of genomic analyses, tools like SpliceAI and MMSplice are typically employed not in isolation but as part of a more extensive set of methods to prioritize pathogenic variants with deleterious effects. Belbin et al. explored a cryptic splice variant in <italic>ABCB4</italic>, predicted to cause a splice acceptor loss by SpliceAI (score&#x2009;&#x003D;&#x2009;0.39) through an IBD-based (identity-by-descent) phenome-wide association study (PheWAS) analysis and fine-mapping. It was further validated to disrupt the splicing of the ABCB4 pre-mRNA <italic>in vitro</italic>, leading to the skip transcription of exon 23, thus resulting in liver disease (<xref ref-type="bibr" rid="B37">37</xref>). Given the complex genetic architecture of the congenital disease, most of the time, researchers may not only employ tools for the prediction of coding variants but also adopt other tools for the prediction of splicing variants or regulatory variants. For example, a study that concentrated on the detection of mosaic mutation implicated in CHD captured deleterious missense variant by REVEL (with a score&#x2009;&#x003E;&#x2009;0.5) and damaging splicing variants by SpliceAI (with a delta score&#x2009;&#x003E;&#x2009;0.5) (<xref ref-type="bibr" rid="B38">38</xref>). Therefore, researchers would annotate the splice variants together with other variants using some ensemble tools or databases. Take CADD-Splice (same as CADD v1.6) as an example, it integrated with several superior ML-based methods (including SpliceAI and MMSplice) to score the potential splicing effect led by the genetic variations (<xref ref-type="bibr" rid="B39">39</xref>). On the other hand, dbNSFP is a comprehensive database designed to annotate the functional impact of all SNPs in the human genome. It complied dozens of prediction scores from various tools, consisting of (i) functional prediction (from SIFT, Polyphen, CADD, etc.), (ii) conservation scores (from phyloP, phastCons, GERP&#x002B;&#x002B;, etc.), and (iii) many other variant annotations like allele frequency, gene information, protein information, splicing effect, regulatory elements, and gene-associated phenotype of mouse and zebrafish (<xref ref-type="bibr" rid="B40">40</xref>).</p>
</sec>
<sec id="s3c"><label>3.3.</label><title>Prioritizing potentially damaging regulatory variants</title>
<p>Historically, the majority of diseases&#x2019; pathogenic variants are detected in the protein-coding regions, although it only takes up around 2&#x0025; of human genomes. Nonetheless, disease-causing variations in the coding areas could only elucidate about 20&#x0025;&#x2013;50&#x0025; of the diseases&#x2019; etiology, indicating that rare noncoding variations may contribute substantially to disease risk (<xref ref-type="bibr" rid="B41">41</xref>). Unlike coding variants that may affect protein structure, function and folding, noncoding variants disrupting functional regulatory elements (e.g., enhancers, insulators, promoters, etc.) have the potential to dysregulate gene expression and thus contribute to genetic diseases (<xref ref-type="bibr" rid="B42">42</xref>). Deleteriousness prediction tools primarily trained with coding datasets, like CADD and REVEL, are insufficient to predict the pathogenicity of noncoding variation. Hence, other variation annotation tools specialized in predicting the regulatory effect of noncoding variants are needed.</p>
<p>DeepSEA is a deep learning model specialized in predicting the functional effects of noncoding mutations. It uses a multi-layer CNN architecture to decode the regulatory sequence from massive epigenomic profiles and predict the chromatin effects of the genomic mutations. DeepSEA takes a 1,000 base pairs (bp) DNA sequence centered on each variant as input and creates a couple of sequences harboring either the reference or alternative allele at the variant position. Then it calculates the chromatin effect size across each epigenomic feature for each reference and alternative allele, in which the absolute differences between wild-type and mutation could be obtained. Additionally, DeepSEA also takes evolutionary conservation into account and computes the conservation score for each variant using PhastCons, PhyloP and GERP&#x002B;&#x002B;. By incorporating the variant-phenotype information on human pathogenic variants from the Human Gene Mutation Database (HGMD), DeepSEA has the capacity to forecast the deleterious regulatory impacts that regulatory variations may have, thereby aiding in the prioritization of functional variations (<xref ref-type="bibr" rid="B43">43</xref>).</p>
<p>DeepSEA is a general deep learning model to predict the regulatory effects of noncoding variants for all kinds of diseases. HeartENN, on the other hand, is a heart-specific neural network built on top of DeepSEA to predict the epigenomic outcomes of variants in relation to heart diseases (like congenital heart disease) with a double number of convolution layers architecture (<xref ref-type="bibr" rid="B44">44</xref>). HeartENN is established with two neural network-based epigenomic effects models, one for predicting heart-specific human chromatin features (histone marks, transcription factors and DNase I accessibility) and the other for mice. To assess the utility of the HeartENN model, developers applied it to the WGS data from 749 CHD trios and 1,611 unaffected trios. They found that variants prioritized by HeartENN damaging score (scores &#x2265;0.1) exhibited significant enrichment of the known human CHD genes in CHD cases. Cooperating with a strategy focused on human fetal cardiac enhancers, they confirmed that genes enriched for noncoding DNVs in human fetal cardiac enhancers also have an excess burden on the noncoding DNVs with HeartENN scores &#x2265;0.1, suggesting the capability of the HeartENN in the prioritization of potentially disruptive regulatory noncoding DNVs implicated in CHD (<xref ref-type="bibr" rid="B44">44</xref>).</p>
<p>Multiscale Analysis of Regulatory Variants on the Epigenomic Landscape (MARVEL) is developed with a ML algorithm GLM-LARS (generalized linear model-based least angle regression) to prioritize phenotype-associated noncoding variants using WGS data and cell-type specific epigenomic profiles. It integrates gene annotation information, publicly available epigenetic data (e.g., enhancers, promoters, transcription factor motifs) from relevant tissues and the covariates of sample phenotypes to identify potential regulatory regions affected by the noncoding variants. The developers applied MARVEL to the WGS data of 431 short-segment Hirschsprung disease (S-HSCR) cases and 487 ethnically matched controls. Together with ChIP-seq and ATAC-seq data of the human pluripotent stem cell (hPSC)-derived enteric NC-like cells (hNC), they uncovered multiple novel genes implicated in S-HSCR by affecting neural crest migration and development (<xref ref-type="bibr" rid="B45">45</xref>).</p>
</sec>
</sec>
<sec id="s4"><label>4.</label><title>Current advances and challenges in variant interpretation</title>
<p>While AI-based tools have made significant contributions to the detection and prioritization of disease-causing variants (<xref ref-type="table" rid="T1">Table&#x00A0;1</xref>), a persistent challenge in genomic research lies in variant interpretation. In 2015, the ACMG/AMP published an authoritative guideline to standardize variant interpretation, which categorizes variants into five classes ranging from benign to pathogenic (<xref ref-type="bibr" rid="B30">30</xref>). Subsequently, multiple platforms were developed for automated variant interpretation based on the ACMG/AMP criteria, such as VarSome, VSClinical, and AION from Nostos genomics. However, although these platforms have facilitated the effective and efficient prioritization of pathogenic or likely pathogenic variants along with their supporting evidence, they still face challenges in interpreting variants of uncertain significance (VUS).</p>
<table-wrap id="T1" position="float"><label>Table 1</label>
<caption><p>AI-based tools utilized in the detection and prioritization of disease-causative variants.</p></caption>
<table frame="hsides" rules="groups">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th valign="top" align="left">Purpose</th>
<th valign="top" align="center">Variant type</th>
<th valign="top" align="center">Tools</th>
<th valign="top" align="center">Methods</th>
<th valign="top" align="center">Launch year</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left" rowspan="7">Variant calling</td>
<td valign="top" align="left" rowspan="3">SNVs/indels</td>
<td valign="top" align="left">GATK</td>
<td valign="top" align="left">HMM, Bayesian, etc.</td>
<td valign="top" align="center">2010</td>
</tr>
<tr>
<td valign="top" align="left">FreeBayes</td>
<td valign="top" align="left">Bayesian</td>
<td valign="top" align="center">2012</td>
</tr>
<tr>
<td valign="top" align="left">DeepVariants</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2018</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="4">CNVs</td>
<td valign="top" align="left">PennCNV</td>
<td valign="top" align="left">HMM</td>
<td valign="top" align="center">2007</td>
</tr>
<tr>
<td valign="top" align="left">CN-Learn</td>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">2019</td>
</tr>
<tr>
<td valign="top" align="left">CNV-JACG</td>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">2020</td>
</tr>
<tr>
<td valign="top" align="left">DeepCNV</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2021</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="7">Variant prioritizing</td>
<td valign="top" align="left" rowspan="2">Coding variants</td>
<td valign="top" align="left">CADD</td>
<td valign="top" align="left">SVM</td>
<td valign="top" align="center">2014</td>
</tr>
<tr>
<td valign="top" align="left">REVEL</td>
<td valign="top" align="left">Random Forest</td>
<td valign="top" align="center">2016</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="2">Splicing variants</td>
<td valign="top" align="left">SpliceAI</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2019</td>
</tr>
<tr>
<td valign="top" align="left">MMSplice</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2019</td>
</tr>
<tr>
<td valign="top" align="left" rowspan="3">Regulatory variants</td>
<td valign="top" align="left">DeepSEA</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2015</td>
</tr>
<tr>
<td valign="top" align="left">HeartENN</td>
<td valign="top" align="left">Deep CNN</td>
<td valign="top" align="center">2020</td>
</tr>
<tr>
<td valign="top" align="left">MARVEL</td>
<td valign="top" align="left">GLM-LARS</td>
<td valign="top" align="center">2020</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-fn1"><p>HMM, hidden markov model; CNN, convolutional neural networks; SVM, support vector machine; GLM-LARS, generalized linear model-based least angle regression.</p></fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s5" sec-type="conclusions"><label>5.</label><title>Conclusion</title>
<p>AI makes it possible to integrate and model vast amounts of genomic data quickly and accurately, facilitating the identification, annotation and prioritization of genetic mutations that contribute to disease development (<xref ref-type="fig" rid="F2">Figure&#x00A0;2</xref>). However, large amounts of diverse data are required to train the AI models. Small sample sizes and the lack of diversity in the data available for genomic analysis can limit the accuracy and reliability of the results generated by these models. Moreover, due to the potential variability in predicted outputs generated by distinct AI models, clinicians and researchers may encounter difficulties discerning the most precise outcome and interpreting the underlying pathomechanisms of congenital diseases. Overall, AI has been confirmed to be a powerful tool that revolutionizes disease-specific genomic analysis by providing speedy and precise insights into the complex relationship between genetics and disease development. Ultimately, these findings that traditional methods might have missed will lead to earlier diagnosis and better prognoses for patients with complex congenital disorders.</p>
<fig id="F2" position="float"><label>Figure 2</label>
<caption><p>Schematic diagram of the review. (<bold>A</bold>) Application of AI models in variant detection, including SNVs, indels and CNVs; (<bold>B</bold>) Application of AI models in the prioritization of disease-causing variants in different genomic regions (coding, splicing, and noncoding); (<bold>C</bold>) Application of AI models in the research of congenital surgical diseases.</p></caption>
<graphic xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="fped-11-1203289-g002.tif"/>
</fig>
</sec>
</body>
<back>
<sec id="s6" sec-type="author-contributions"><title>Author contributions</title>
<p>All authors listed have made a substantial, direct, and intellectual contribution to the work and approved it for publication.</p>
</sec>
<sec id="s7" sec-type="funding-information"><title>Funding</title>
<p>This study was supported by the Theme-based Research Scheme (T12C-714/14-R and T12-712/21-R), the General Research Fund (17113320 and 17113420 to CT), and the Health and Medical Research Fund (PR-HKU-1 to PT, 08193446 and 09201436 to CT).</p>
</sec>
<sec id="s8" sec-type="COI-statement"><title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s9" sec-type="disclaimer"><title>Publisher&#x0027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list><title>References</title>
<ref id="B1"><label>1.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pierpont</surname><given-names>ME</given-names></name><name><surname>Brueckner</surname><given-names>M</given-names></name><name><surname>Chung</surname><given-names>WK</given-names></name><name><surname>Garg</surname><given-names>V</given-names></name><name><surname>Lacro</surname><given-names>RV</given-names></name><name><surname>McGuire</surname><given-names>AL</given-names></name><etal/></person-group> <article-title>Genetic basis for congenital heart disease: revisited: a scientific statement from the American heart association</article-title>. <source>Circ</source>. (<year>2018</year>) <volume>138</volume>(<issue>21</issue>):<fpage>e653</fpage>&#x2013;<lpage>e711</lpage>. <pub-id pub-id-type="doi">10.1161/CIR.0000000000000606</pub-id></citation></ref>
<ref id="B2"><label>2.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nees</surname><given-names>SN</given-names></name><name><surname>Chung</surname><given-names>WK</given-names></name></person-group>. <article-title>Genetic basis of human congenital heart disease</article-title>. <source>Cold Spring Harbor Perspect Biol</source>. (<year>2020</year>) <volume>12</volume>(<issue>9</issue>):<fpage>a036749</fpage>. <pub-id pub-id-type="doi">10.1101/cshperspect.a036749</pub-id></citation></ref>
<ref id="B3"><label>3.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Negri</surname><given-names>E</given-names></name><name><surname>Coletta</surname><given-names>R</given-names></name><name><surname>Morabito</surname><given-names>A</given-names></name></person-group>. <article-title>Congenital short bowel syndrome: systematic review of a rare condition</article-title>. <source>J Pediatr Surg</source>. (<year>2020</year>) <volume>55</volume>(<issue>9</issue>):<fpage>1809</fpage>&#x2013;<lpage>14</lpage>. <pub-id pub-id-type="doi">10.1016/j.jpedsurg.2020.03.009</pub-id><pub-id pub-id-type="pmid">32278545</pub-id></citation></ref>
<ref id="B4"><label>4.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Poole</surname><given-names>DI</given-names></name><name><surname>Goebel</surname><given-names>RG</given-names></name><name><surname>Mackworth</surname><given-names>AK</given-names></name></person-group>. <source>Computational intelligence</source>. Vol. <volume>1</volume>. <publisher-loc>New York</publisher-loc>: <publisher-name>Oxford University Press</publisher-name> (<year>1998</year>).</citation></ref>
<ref id="B5"><label>5.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Azmi</surname><given-names>J</given-names></name><name><surname>Arif</surname><given-names>M</given-names></name><name><surname>Nafis</surname><given-names>MT</given-names></name><name><surname>Alam</surname><given-names>MA</given-names></name><name><surname>Tanweer</surname><given-names>S</given-names></name><name><surname>Wang</surname><given-names>G</given-names></name></person-group>. <article-title>A systematic review on machine learning approaches for cardiovascular disease prediction using medical big data</article-title>. <source>Med Eng Phys</source>. (<year>2022</year>) <volume>105</volume>:<fpage>103825</fpage>. <pub-id pub-id-type="doi">10.1016/j.medengphy.2022.103825</pub-id><pub-id pub-id-type="pmid">35781385</pub-id></citation></ref>
<ref id="B6"><label>6.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Galal</surname><given-names>A</given-names></name><name><surname>Talal</surname><given-names>M</given-names></name><name><surname>Moustafa</surname><given-names>A</given-names></name></person-group>. <article-title>Applications of machine learning in metabolomics: disease modeling and classification</article-title>. <source>Front Genet</source>. (<year>2022</year>) <volume>13</volume>:<fpage>1017340</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2022.1017340</pub-id><pub-id pub-id-type="pmid">36506316</pub-id></citation></ref>
<ref id="B7"><label>7.</label><citation citation-type="confproc"><person-group person-group-type="author"><name><surname>Ongsulee</surname><given-names>P</given-names></name></person-group>. <conf-name>Artificial intelligence, machine learning and deep learning</conf-name>. <conf-name>2017 15th international conference on ICT and knowledge engineering (ICT&#x0026;KE)</conf-name>; <conf-loc>Bangkok, Thailand</conf-loc> (<year>2017</year>).</citation></ref>
<ref id="B8"><label>8.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Schmidhuber</surname><given-names>J</given-names></name></person-group>. <article-title>Deep learning in neural networks: an overview</article-title>. <source>Neural Netw</source>. (<year>2015</year>) <volume>61</volume>:<fpage>85</fpage>&#x2013;<lpage>117</lpage>. <pub-id pub-id-type="doi">10.1016/j.neunet.2014.09.003</pub-id><pub-id pub-id-type="pmid">25462637</pub-id></citation></ref>
<ref id="B9"><label>9.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Consortium</surname><given-names>GP</given-names></name></person-group>. <article-title>A global reference for human genetic variation</article-title>. <source>Nature</source>. (<year>2015</year>) <volume>526</volume>(<issue>7571</issue>):<fpage>68</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1038/nature15393</pub-id><pub-id pub-id-type="pmid">26432245</pub-id></citation></ref>
<ref id="B10"><label>10.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>McKenna</surname><given-names>A</given-names></name><name><surname>Hanna</surname><given-names>M</given-names></name><name><surname>Banks</surname><given-names>E</given-names></name><name><surname>Sivachenko</surname><given-names>A</given-names></name><name><surname>Cibulskis</surname><given-names>K</given-names></name><name><surname>Kernytsky</surname><given-names>A</given-names></name><etal/></person-group> <article-title>The genome analysis toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data</article-title>. <source>Genome Res</source>. (<year>2010</year>) <volume>20</volume>(<issue>9</issue>):<fpage>1297</fpage>&#x2013;<lpage>303</lpage>. <pub-id pub-id-type="doi">10.1101/gr.107524.110</pub-id><pub-id pub-id-type="pmid">20644199</pub-id></citation></ref>
<ref id="B11"><label>11.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Van der Auwera</surname><given-names>GA</given-names></name><name><surname>Carneiro</surname><given-names>MO</given-names></name><name><surname>Hartl</surname><given-names>C</given-names></name><name><surname>Poplin</surname><given-names>R</given-names></name><name><surname>Del Angel</surname><given-names>G</given-names></name><name><surname>Levy-Moonshine</surname><given-names>A</given-names></name><etal/></person-group> <article-title>From FastQ data to high-confidence variant calls: the genome analysis toolkit best practices pipeline</article-title>. <source>Curr Protoc Bioinformatics</source>. (<year>2013</year>) <volume>43</volume>(<issue>1</issue>):<fpage>11.10.1</fpage>&#x2013;<lpage>11.10.33</lpage>. <pub-id pub-id-type="doi">10.1002/0471250953.bi1110s43</pub-id><pub-id pub-id-type="pmid">25431634</pub-id></citation></ref>
<ref id="B12"><label>12.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>DePristo</surname><given-names>MA</given-names></name><name><surname>Banks</surname><given-names>E</given-names></name><name><surname>Poplin</surname><given-names>R</given-names></name><name><surname>Garimella</surname><given-names>KV</given-names></name><name><surname>Maguire</surname><given-names>JR</given-names></name><name><surname>Hartl</surname><given-names>C</given-names></name><etal/></person-group> <article-title>A framework for variation discovery and genotyping using next-generation DNA sequencing data</article-title>. <source>Nat Genet</source>. (<year>2011</year>) <volume>43</volume>(<issue>5</issue>):<fpage>491</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1038/ng.806</pub-id><pub-id pub-id-type="pmid">21478889</pub-id></citation></ref>
<ref id="B13"><label>13.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Poplin</surname><given-names>R</given-names></name><name><surname>Chang</surname><given-names>PC</given-names></name><name><surname>Alexander</surname><given-names>D</given-names></name><name><surname>Schwartz</surname><given-names>S</given-names></name><name><surname>Colthurst</surname><given-names>T</given-names></name><name><surname>Ku</surname><given-names>A</given-names></name><etal/></person-group> <article-title>A universal SNP and small-indel variant caller using deep neural networks</article-title>. <source>Nat Biotechnol</source>. (<year>2018</year>) <volume>36</volume>(<issue>10</issue>):<fpage>983</fpage>&#x2013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.1038/nbt.4235</pub-id><pub-id pub-id-type="pmid">30247488</pub-id></citation></ref>
<ref id="B14"><label>14.</label><citation citation-type="other"><person-group person-group-type="author"><name><surname>Garrison</surname><given-names>E</given-names></name><name><surname>Marth</surname><given-names>G</given-names></name></person-group>. <comment>Haplotype-based variant detection from short-read sequencing. arXiv preprint arXiv:1207.3907</comment> (<year>2012</year>).</citation></ref>
<ref id="B15"><label>15.</label><citation citation-type="book"><person-group person-group-type="author"><name><surname>Mac&#x00E9;</surname><given-names>A</given-names></name><name><surname>Kutalik</surname><given-names>Z</given-names></name><name><surname>Valsesia</surname><given-names>A</given-names></name></person-group>. <article-title>Copy number variation</article-title>. <source>Methods Mol Biol</source>. (<year>2018</year>) <volume>1793</volume>:<fpage>231</fpage>&#x2013;<lpage>258</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-7868-7_14</pub-id></citation></ref>
<ref id="B16"><label>16.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wang</surname><given-names>K</given-names></name><name><surname>Li</surname><given-names>M</given-names></name><name><surname>Hadley</surname><given-names>D</given-names></name><name><surname>Liu</surname><given-names>R</given-names></name><name><surname>Glessner</surname><given-names>J</given-names></name><name><surname>Grant</surname><given-names>SF</given-names></name><etal/></person-group> <article-title>PennCNV: an integrated hidden markov model designed for high-resolution copy number variation detection in whole-genome SNP genotyping data</article-title>. <source>Genome Res</source>. (<year>2007</year>) <volume>17</volume>(<issue>11</issue>):<fpage>1665</fpage>&#x2013;<lpage>74</lpage>. <pub-id pub-id-type="doi">10.1101/gr.6861907</pub-id><pub-id pub-id-type="pmid">17921354</pub-id></citation></ref>
<ref id="B17"><label>17.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glessner</surname><given-names>JT</given-names></name><name><surname>Hou</surname><given-names>X</given-names></name><name><surname>Zhong</surname><given-names>C</given-names></name><name><surname>Zhang</surname><given-names>J</given-names></name><name><surname>Khan</surname><given-names>M</given-names></name><name><surname>Brand</surname><given-names>F</given-names></name><etal/></person-group> <article-title>DeepCNV: a deep learning approach for authenticating copy number variations</article-title>. <source>Brief Bioinform</source>. (<year>2021</year>) <volume>22</volume>(<issue>5</issue>):<fpage>bbaa381</fpage>. <pub-id pub-id-type="doi">10.1093/bib/bbaa381</pub-id><pub-id pub-id-type="pmid">33429424</pub-id></citation></ref>
<ref id="B18"><label>18.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pounraja</surname><given-names>VK</given-names></name><name><surname>Jayakar</surname><given-names>G</given-names></name><name><surname>Jensen</surname><given-names>M</given-names></name><name><surname>Kelkar</surname><given-names>N</given-names></name><name><surname>Girirajan</surname><given-names>S</given-names></name></person-group>. <article-title>A machine-learning approach for accurate detection of copy number variants from exome sequencing</article-title>. <source>Genome Res</source>. (<year>2019</year>) <volume>29</volume>(<issue>7</issue>):<fpage>1134</fpage>&#x2013;<lpage>43</lpage>. <pub-id pub-id-type="doi">10.1101/gr.245928.118</pub-id><pub-id pub-id-type="pmid">31171634</pub-id></citation></ref>
<ref id="B19"><label>19.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhuang</surname><given-names>X</given-names></name><name><surname>Ye</surname><given-names>R</given-names></name><name><surname>So</surname><given-names>MT</given-names></name><name><surname>Lam</surname><given-names>W-Y</given-names></name><name><surname>Karim</surname><given-names>A</given-names></name><name><surname>Yu</surname><given-names>M</given-names></name><etal/></person-group> <article-title>A random forest-based framework for genotyping and accuracy assessment of copy number variations</article-title>. <source>NAR Genomics Bioinform</source>. (<year>2020</year>) <volume>2</volume>(<issue>3</issue>):<fpage>lqaa071</fpage>. <pub-id pub-id-type="doi">10.1093/nargab/lqaa071</pub-id></citation></ref>
<ref id="B20"><label>20.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fromer</surname><given-names>M</given-names></name><name><surname>Moran</surname><given-names>JL</given-names></name><name><surname>Chambert</surname><given-names>K</given-names></name><name><surname>Banks</surname><given-names>E</given-names></name><name><surname>Bergen</surname><given-names>SE</given-names></name><name><surname>Ruderfer</surname><given-names>DM</given-names></name><etal/></person-group> <article-title>Discovery and statistical genotyping of copy-number variation from whole-exome sequencing depth</article-title>. <source>Am J Hum Genet</source>. (<year>2012</year>) <volume>91</volume>(<issue>4</issue>):<fpage>597</fpage>&#x2013;<lpage>607</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2012.08.005</pub-id><pub-id pub-id-type="pmid">23040492</pub-id></citation></ref>
<ref id="B21"><label>21.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Glessner</surname><given-names>JT</given-names></name><name><surname>Bick</surname><given-names>AG</given-names></name><name><surname>Ito</surname><given-names>K</given-names></name><name><surname>Homsy</surname><given-names>JG</given-names></name><name><surname>Rodriguez-Murillo</surname><given-names>L</given-names></name><name><surname>Fromer</surname><given-names>M</given-names></name><etal/></person-group> <article-title>Increased frequency of de novo copy number variants in congenital heart disease by integrative analysis of single nucleotide polymorphism array and exome sequence data</article-title>. <source>Circ Res</source>. (<year>2014</year>) <volume>115</volume>(<issue>10</issue>):<fpage>884</fpage>&#x2013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1161/CIRCRESAHA.115.304458</pub-id><pub-id pub-id-type="pmid">25205790</pub-id></citation></ref>
<ref id="B22"><label>22.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bailliard</surname><given-names>F</given-names></name><name><surname>Anderson</surname><given-names>RH</given-names></name></person-group>. <article-title>Tetralogy of fallot</article-title>. <source>Orphanet J Rare Dis</source>. (<year>2009</year>) <volume>4</volume>:<fpage>1</fpage>&#x2013;<lpage>10</lpage>. <pub-id pub-id-type="doi">10.1186/1750-1172-4-2</pub-id><pub-id pub-id-type="pmid">19133130</pub-id></citation></ref>
<ref id="B23"><label>23.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Apitz</surname><given-names>C</given-names></name><name><surname>Webb</surname><given-names>GD</given-names></name><name><surname>Redington</surname><given-names>AN</given-names></name></person-group>. <article-title>Tetralogy of fallot</article-title>. <source>Lancet</source>. (<year>2009</year>) <volume>374</volume>(<issue>9699</issue>):<fpage>1462</fpage>&#x2013;<lpage>71</lpage>. <pub-id pub-id-type="doi">10.1016/S0140-6736(09)60657-7</pub-id><pub-id pub-id-type="pmid">19683809</pub-id></citation></ref>
<ref id="B24"><label>24.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname><given-names>CSM</given-names></name><name><surname>Mononen</surname><given-names>M</given-names></name><name><surname>Lam</surname><given-names>WY</given-names></name><name><surname>Jin</surname><given-names>SC</given-names></name><name><surname>Zhuang</surname><given-names>X</given-names></name><name><surname>Garcia-Barcelo</surname><given-names>M-M</given-names></name><etal/></person-group> <article-title>Sequencing of a Chinese tetralogy of fallot cohort reveals clustering mutations in myogenic heart progenitors</article-title>. <source>JCI Insight</source>. (<year>2022</year>) <volume>7</volume>(<issue>2</issue>):<fpage>e152198</fpage>. <pub-id pub-id-type="doi">10.1172/jci.insight.152198</pub-id><pub-id pub-id-type="pmid">34905512</pub-id></citation></ref>
<ref id="B25"><label>25.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Tang</surname><given-names>CS-M</given-names></name><name><surname>Cheng</surname><given-names>G</given-names></name><name><surname>So</surname><given-names>MT</given-names></name><name><surname>Yip</surname><given-names>BHK</given-names></name><name><surname>Miao</surname><given-names>XP</given-names></name><name><surname>Wong</surname><given-names>EHM</given-names></name><etal/></person-group> <article-title>Genome-wide copy number analysis uncovers a new HSCR gene: NRG3</article-title>. <source>PLoS Genet</source>. (<year>2012</year>) <volume>8</volume>(<issue>5</issue>):<fpage>e1002687</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1002687</pub-id><pub-id pub-id-type="pmid">22589734</pub-id></citation></ref>
<ref id="B26"><label>26.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname><given-names>G</given-names></name><name><surname>Chung</surname><given-names>PHY</given-names></name><name><surname>Chan</surname><given-names>EKW</given-names></name><name><surname>So</surname><given-names>MT</given-names></name><name><surname>Sham</surname><given-names>PC</given-names></name><name><surname>Cherny</surname><given-names>SS</given-names></name><etal/></person-group> <article-title>Patient complexity and genotype-phenotype correlations in biliary atresia: a cross-sectional analysis</article-title>. <source>BMC Med Genomics</source>. (<year>2017</year>) <volume>10</volume>:<fpage>1</fpage>&#x2013;<lpage>9</lpage>. <pub-id pub-id-type="doi">10.1186/s12920-016-0237-y</pub-id><pub-id pub-id-type="pmid">28057009</pub-id></citation></ref>
<ref id="B27"><label>27.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kircher</surname><given-names>M</given-names></name><name><surname>Witten</surname><given-names>DM</given-names></name><name><surname>Jain</surname><given-names>P</given-names></name><name><surname>O&#x0027;roak</surname><given-names>BJ</given-names></name><name><surname>Cooper</surname><given-names>GM</given-names></name><name><surname>Shendure</surname><given-names>J</given-names></name><etal/></person-group> <article-title>A general framework for estimating the relative pathogenicity of human genetic variants</article-title>. <source>Nat Genet</source>. (<year>2014</year>) <volume>46</volume>(<issue>3</issue>):<fpage>310</fpage>&#x2013;<lpage>5</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2892</pub-id><pub-id pub-id-type="pmid">24487276</pub-id></citation></ref>
<ref id="B28"><label>28.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rentzsch</surname><given-names>P</given-names></name><name><surname>Witten</surname><given-names>D</given-names></name><name><surname>Cooper</surname><given-names>GM</given-names></name><name><surname>Shendure</surname><given-names>J</given-names></name><name><surname>Kircher</surname><given-names>M</given-names></name></person-group>. <article-title>CADD: predicting the deleteriousness of variants throughout the human genome</article-title>. <source>Nucleic Acids Res</source>. (<year>2019</year>) <volume>47</volume>(<issue>D1</issue>):<fpage>D886</fpage>&#x2013;<lpage>D894</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gky1016</pub-id><pub-id pub-id-type="pmid">30371827</pub-id></citation></ref>
<ref id="B29"><label>29.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ioannidis</surname><given-names>NM</given-names></name><name><surname>Rothstein</surname><given-names>JH</given-names></name><name><surname>Pejaver</surname><given-names>V</given-names></name><name><surname>Middha</surname><given-names>S</given-names></name><name><surname>McDonnell</surname><given-names>SK</given-names></name><name><surname>Baheti</surname><given-names>S</given-names></name><etal/></person-group> <source>Am J Hum Genet</source>. (<year>2016</year>) <volume>99</volume>(<issue>4</issue>):<fpage>877</fpage>&#x2013;<lpage>85</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2016.08.016</pub-id><pub-id pub-id-type="pmid">27666373</pub-id></citation></ref>
<ref id="B30"><label>30.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Richards</surname><given-names>S</given-names></name><name><surname>Aziz</surname><given-names>N</given-names></name><name><surname>Bale</surname><given-names>S</given-names></name><name><surname>Bick</surname><given-names>D</given-names></name><name><surname>Das</surname><given-names>S</given-names></name><name><surname>Gastier-Foster</surname><given-names>J</given-names></name><etal/></person-group> <article-title>Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American college of medical genetics and genomics and the association for molecular pathology</article-title>. <source>Genet Med</source>. (<year>2015</year>) <volume>17</volume>(<issue>5</issue>):<fpage>405</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1038/gim.2015.30</pub-id><pub-id pub-id-type="pmid">25741868</pub-id></citation></ref>
<ref id="B31"><label>31.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rivera-Mu&#x00F1;oz</surname><given-names>EA</given-names></name><name><surname>Milko</surname><given-names>LV</given-names></name><name><surname>Harrison</surname><given-names>SM</given-names></name><name><surname>Azzariti</surname><given-names>DR</given-names></name><name><surname>Kurtz</surname><given-names>CL</given-names></name><name><surname>Lee</surname><given-names>K</given-names></name><etal/></person-group> <article-title>Clingen variant curation expert panel experiences and standardized processes for disease and gene-level specification of the ACMG/AMP guidelines for sequence variant interpretation</article-title>. <source>Hum Mutat</source>. (<year>2018</year>) <volume>39</volume>(<issue>11</issue>):<fpage>1614</fpage>&#x2013;<lpage>22</lpage>. <pub-id pub-id-type="doi">10.1002/humu.23645</pub-id></citation></ref>
<ref id="B32"><label>32.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Page</surname><given-names>DJ</given-names></name><name><surname>Miossec</surname><given-names>MJ</given-names></name><name><surname>Williams</surname><given-names>SG</given-names></name><name><surname>Monaghan</surname><given-names>RM</given-names></name><name><surname>Fotiou</surname><given-names>E</given-names></name><name><surname>Cordell</surname><given-names>HJ</given-names></name><etal/></person-group> <article-title>Whole exome sequencing reveals the major genetic contributors to nonsyndromic tetralogy of fallot</article-title>. <source>Circ Res</source>. (<year>2019</year>) <volume>124</volume>(<issue>4</issue>):<fpage>553</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1161/CIRCRESAHA.118.313250</pub-id><pub-id pub-id-type="pmid">30582441</pub-id></citation></ref>
<ref id="B33"><label>33.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lam</surname><given-names>WY</given-names></name><name><surname>Tang</surname><given-names>CSM</given-names></name><name><surname>So</surname><given-names>MT</given-names></name><name><surname>Yue</surname><given-names>H</given-names></name><name><surname>Hsu</surname><given-names>JS</given-names></name><name><surname>Chung</surname><given-names>PH</given-names></name><etal/></person-group> <article-title>Identification of a wide spectrum of ciliary gene mutations in nonsyndromic biliary atresia patients implicates ciliary dysfunction as a novel disease mechanism</article-title>. <source>EBioMedicine</source>. (<year>2021</year>) <volume>71</volume>:<fpage>103530</fpage>. <pub-id pub-id-type="doi">10.1016/j.ebiom.2021.103530</pub-id><pub-id pub-id-type="pmid">34455394</pub-id></citation></ref>
<ref id="B34"><label>34.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Krawczak</surname><given-names>M</given-names></name><name><surname>Reiss</surname><given-names>J</given-names></name><name><surname>Cooper</surname><given-names>DN</given-names></name></person-group>. <article-title>The mutational spectrum of single base-pair substitutions in mRNA splice junctions of human genes: causes and consequences</article-title>. <source>Hum Genet</source>. (<year>1992</year>) <volume>90</volume>(<issue>1&#x2013;2</issue>):<fpage>41</fpage>&#x2013;<lpage>54</lpage>. <pub-id pub-id-type="doi">10.1007/BF00210743</pub-id><pub-id pub-id-type="pmid">1427786</pub-id></citation></ref>
<ref id="B35"><label>35.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jaganathan</surname><given-names>K</given-names></name><name><surname>Panagiotopoulou</surname><given-names>SK</given-names></name><name><surname>McRae</surname><given-names>JF</given-names></name><name><surname>Darbandi</surname><given-names>SF</given-names></name><name><surname>Knowles</surname><given-names>D</given-names></name><name><surname>Li</surname><given-names>YI</given-names></name><etal/></person-group> <article-title>Predicting splicing from primary sequence with deep learning</article-title>. <source>Cell</source>. (<year>2019</year>) <volume>176</volume>(<issue>3</issue>):<fpage>535</fpage>&#x2013;<lpage>548.e24</lpage>. <pub-id pub-id-type="doi">10.1016/j.cell.2018.12.015</pub-id><pub-id pub-id-type="pmid">30661751</pub-id></citation></ref>
<ref id="B36"><label>36.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Cheng</surname><given-names>J</given-names></name><name><surname>Nguyen</surname><given-names>TYD</given-names></name><name><surname>Cygan</surname><given-names>KJ</given-names></name><name><surname>&#x00C7;elik</surname><given-names>MH</given-names></name><name><surname>Fairbrother</surname><given-names>WG</given-names></name><name><surname>Avsec</surname><given-names>Z</given-names></name><etal/></person-group> <article-title>MMSplice: modular modeling improves the predictions of genetic variant effects on splicing</article-title>. <source>Genome Biol</source>. (<year>2019</year>) <volume>20</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>15</lpage>. <pub-id pub-id-type="doi">10.1186/s13059-019-1653-z</pub-id><pub-id pub-id-type="pmid">30606230</pub-id></citation></ref>
<ref id="B37"><label>37.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Belbin</surname><given-names>GM</given-names></name><name><surname>Rutledge</surname><given-names>S</given-names></name><name><surname>Dodatko</surname><given-names>T</given-names></name><name><surname>Cullina</surname><given-names>S</given-names></name><name><surname>Turchin</surname><given-names>MC</given-names></name><name><surname>Kohli</surname><given-names>S</given-names></name><etal/></person-group> <article-title>Leveraging health systems data to characterize a large effect variant conferring risk for liver disease in Puerto Ricans</article-title>. <source>Am J Hum Genet</source>. (<year>2021</year>) <volume>108</volume>(<issue>11</issue>):<fpage>2099</fpage>&#x2013;<lpage>111</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2021.09.016</pub-id><pub-id pub-id-type="pmid">34678161</pub-id></citation></ref>
<ref id="B38"><label>38.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Hsieh</surname><given-names>A</given-names></name><name><surname>Morton</surname><given-names>SU</given-names></name><name><surname>Willcox</surname><given-names>JA</given-names></name><name><surname>Gorham</surname><given-names>JM</given-names></name><name><surname>Tai</surname><given-names>AC</given-names></name><name><surname>Qi</surname><given-names>H</given-names></name><etal/></person-group> <article-title>EM-mosaic detects mosaic point mutations that contribute to congenital heart disease</article-title>. <source>Genome Med</source>. (<year>2020</year>) <volume>12</volume>:<fpage>1</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1186/s13073-020-00738-1</pub-id></citation></ref>
<ref id="B39"><label>39.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rentzsch</surname><given-names>P</given-names></name><name><surname>Schubach</surname><given-names>M</given-names></name><name><surname>Shendure</surname><given-names>J</given-names></name><name><surname>Kircher</surname><given-names>M</given-names></name></person-group>. <article-title>CADD-Splice&#x2014;improving genome-wide variant effect prediction using deep learning-derived splice scores</article-title>. <source>Genome Med</source>. (<year>2021</year>) <volume>13</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>12</lpage>. <pub-id pub-id-type="doi">10.1186/s13073-021-00835-9</pub-id><pub-id pub-id-type="pmid">33397400</pub-id></citation></ref>
<ref id="B40"><label>40.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname><given-names>X</given-names></name><name><surname>Li</surname><given-names>C</given-names></name><name><surname>Mou</surname><given-names>C</given-names></name><name><surname>Dong</surname><given-names>Y</given-names></name><name><surname>Tu</surname><given-names>Y</given-names></name></person-group>. <article-title>dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs</article-title>. <source>Genome Med</source>. (<year>2020</year>) <volume>12</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>8</lpage>. <pub-id pub-id-type="doi">10.1186/s13073-019-0693-z</pub-id></citation></ref>
<ref id="B41"><label>41.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zappala</surname><given-names>Z</given-names></name><name><surname>Montgomery</surname><given-names>SB</given-names></name></person-group>. <article-title>Non-coding loss-of-function variation in human genomes</article-title>. <source>Hum Hered</source>. (<year>2016</year>) <volume>81</volume>(<issue>2</issue>):<fpage>78</fpage>&#x2013;<lpage>87</lpage>. <pub-id pub-id-type="doi">10.1159/000447453</pub-id><pub-id pub-id-type="pmid">28076858</pub-id></citation></ref>
<ref id="B42"><label>42.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Perenthaler</surname><given-names>E</given-names></name><name><surname>Yousefi</surname><given-names>S</given-names></name><name><surname>Niggl</surname><given-names>E</given-names></name><name><surname>Barakat</surname><given-names>TS</given-names></name></person-group>. <article-title>Beyond the exome: the non-coding genome and enhancers in neurodevelopmental disorders and malformations of cortical development</article-title>. <source>Front Cell Neurosci</source>. (<year>2019</year>) <volume>13</volume>:<fpage>352</fpage>. <pub-id pub-id-type="doi">10.3389/fncel.2019.00352</pub-id><pub-id pub-id-type="pmid">31417368</pub-id></citation></ref>
<ref id="B43"><label>43.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname><given-names>J</given-names></name><name><surname>Troyanskaya</surname><given-names>OG</given-names></name></person-group>. <article-title>Predicting effects of noncoding variants with deep learning&#x2013;based sequence model</article-title>. <source>Nat Methods</source>. (<year>2015</year>) <volume>12</volume>(<issue>10</issue>):<fpage>931</fpage>&#x2013;<lpage>4</lpage>. <pub-id pub-id-type="doi">10.1038/nmeth.3547</pub-id><pub-id pub-id-type="pmid">26301843</pub-id></citation></ref>
<ref id="B44"><label>44.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Richter</surname><given-names>F</given-names></name><name><surname>Morton</surname><given-names>SU</given-names></name><name><surname>Kim</surname><given-names>SW</given-names></name><name><surname>Kitaygorodsky</surname><given-names>A</given-names></name><name><surname>Wasson</surname><given-names>LK</given-names></name><name><surname>Chen</surname><given-names>KM</given-names></name><etal/></person-group> <article-title>Genomic analyses implicate noncoding de novo variants in congenital heart disease</article-title>. <source>Nat Genet</source>. (<year>2020</year>) <volume>52</volume>(<issue>8</issue>):<fpage>769</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-020-0652-z</pub-id><pub-id pub-id-type="pmid">32601476</pub-id></citation></ref>
<ref id="B45"><label>45.</label><citation citation-type="journal"><person-group person-group-type="author"><name><surname>Fu</surname><given-names>AX</given-names></name><name><surname>Lui</surname><given-names>KC</given-names></name><name><surname>Tang</surname><given-names>CM</given-names></name><name><surname>Ng</surname><given-names>RK</given-names></name><name><surname>Lai</surname><given-names>FL</given-names></name><name><surname>Lau</surname><given-names>ST</given-names></name><etal/></person-group> <article-title>Whole-genome analysis of noncoding genetic variations identifies multiscale regulatory element perturbations associated with hirschsprung disease</article-title>. <source>Genome Res</source>. (<year>2020</year>) <volume>30</volume>(<issue>11</issue>):<fpage>1618</fpage>&#x2013;<lpage>32</lpage>. <pub-id pub-id-type="doi">10.1101/gr.264473.120</pub-id></citation></ref></ref-list>
</back>
</article>