<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">758366</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.758366</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Genetic Variation and the Distribution of Variant Types in the Horse</article-title>
<alt-title alt-title-type="left-running-head">Durward-Akhurst et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Genetic Variation in the Horse</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Durward-Akhurst</surname>
<given-names>S. A.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1439327/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Schaefer</surname>
<given-names>R. J.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Grantham</surname>
<given-names>B.</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1512319/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Carey</surname>
<given-names>W. K.</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Mickelson</surname>
<given-names>J. R.</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/566812/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>McCue</surname>
<given-names>M. E.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/436020/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Department of Veterinary Population Medicine, University of Minnesota, <addr-line>Minneapolis</addr-line>, <addr-line>MN</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Interval Bio LLC, <addr-line>Mountain View</addr-line>, <addr-line>CA</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Department of Veterinary and Biomedical Sciences, University of Minnesota, <addr-line>Minneapolis</addr-line>, <addr-line>MN</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/985272/overview">Jianye Ge</ext-link>, University of North Texas Health Science Center, United&#x20;States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/381884/overview">Shaojun Xie</ext-link>, Purdue University, United&#x20;States</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/760687/overview">Yi-Hsuan Pan</ext-link>, East China Normal University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1228664/overview">Meng Huang</ext-link>, University of Colorado Boulder, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: S. A. Durward-Akhurst, <email>durwa004@umn.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Evolutionary and Population Genetics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>02</day>
<month>12</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>758366</elocation-id>
<history>
<date date-type="received">
<day>16</day>
<month>08</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>10</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Durward-Akhurst, Schaefer, Grantham, Carey, Mickelson and McCue.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Durward-Akhurst, Schaefer, Grantham, Carey, Mickelson and McCue</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Genetic variation is a key contributor to health and disease. Understanding the link between an individual&#x2019;s genotype and the corresponding phenotype is a major goal of medical genetics. Whole genome sequencing (WGS) within and across populations enables highly efficient variant discovery and elucidation of the molecular nature of virtually all genetic variation. Here, we report the largest catalog of genetic variation for the horse, a species of importance as a model for human athletic and performance related traits, using WGS of 534 horses. We show the extent of agreement between two commonly used variant callers. In data from ten target breeds that represent major breed clusters in the domestic horse, we demonstrate the distribution of variants, their allele frequencies across breeds, and identify variants that are unique to a single breed. We investigate variants with no homozygotes that may be potential embryonic lethal variants, as well as variants present in all individuals that likely represent regions of the genome with errors, poor annotation or where the reference genome carries a variant. Finally, we show regions of the genome that have higher or lower levels of genetic variation compared to the genome average. This catalog can be used for variant prioritization for important equine diseases and traits, and to provide key information about regions of the genome where the assembly and/or annotation need to be improved.</p>
</abstract>
<kwd-group>
<kwd>genetic variation</kwd>
<kwd>whole genome sequence</kwd>
<kwd>variant discovery</kwd>
<kwd>equine</kwd>
<kwd>breed differences</kwd>
<kwd>genetics</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Genetic variation is a key contributor to health and disease, and understanding the link between an individual&#x2019;s genotype and the corresponding phenotype is a major goal of genetic research (<xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>). Whole genome sequencing (WGS) within and across populations enables highly efficient variant discovery and elucidation of the molecular nature of virtually all genetic variation, from single nucleotide polymorphisms (SNPs) to copy number variants (CNVs) and other large structural variants (<xref ref-type="bibr" rid="B52">Sudmant et&#x20;al., 2015</xref>). Large-scale studies of genetic variation in humans have dramatically improved our understanding of genetic variation across a species and within populations (<xref ref-type="bibr" rid="B17">Genomes Project et&#x20;al., 2012</xref>). Large-scale variant catalogs establish patterns of variation across the genome, including non-coding regions (<xref ref-type="bibr" rid="B40">Mu et&#x20;al., 2011</xref>), permitting elucidation of regional variability in mutation and recombination rates. Knowledge of the background genetic &#x201c;noise&#x201d; helps to decipher the link between genotype and phenotype&#x2014;allowing filtering of likely neutral variants from potentially deleterious variants in genomic regions of interest or within biologic candidate&#x20;genes.</p>
<p>The equine reference genome (<xref ref-type="bibr" rid="B55">Wade et&#x20;al., 2009</xref>) has provided a key basis for genetic investigations in horse populations (<xref ref-type="bibr" rid="B46">Rebolledo-Mendez et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B45">Raudsepp et&#x20;al., 2019</xref>). However, despite consistent progress in our understanding of genetic disease in the horse, disease-causing variants have been identified for less than 20% of the currently recognized genetic diseases (Online Mendelian Inheritance in Animals, OMIA). Similar to humans, many of the significant GWAS regions of interest found for equine traits have been in non-coding regions of the genome. The unknown function of these regions has been a barrier to pinpointing and confirming the causal variant, which is necessary to develop effective genetic tests or targeted treatments. A more complete account of &#x201c;normal&#x201d; genetic variation in the horse is critical to establishing the link between genotype and phenotype (<xref ref-type="bibr" rid="B59">Yngvadottir et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>).</p>
<p>Here we provide the first large-scale catalog of genetic variation in the horse derived from WGS, with the intent of describing variant numbers, types, allele frequencies, and genomic location distribution across the general population and within and across major breeds.</p>
</sec>
<sec id="s2">
<title>2 Materials and Methods</title>
<sec id="s2-1">
<title>2.1 Samples</title>
<p>Paired-end whole genome sequencing was performed on 534 horses using Illumina technology (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S1</xref>). Forty-four different breeds predominantly from North America and Central Europe were selected (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S2</xref>) based on availability of Illumina WGS from previous and ongoing studies (<xref ref-type="bibr" rid="B15">Finno et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B35">McCoy et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B43">Norton et&#x20;al., 2016</xref>, <xref ref-type="bibr" rid="B42">2019</xref>; <xref ref-type="bibr" rid="B50">Schultz 2016</xref>; <xref ref-type="bibr" rid="B4">Bellone et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B48">Schaefer et&#x20;al., 2017</xref>), publicly available data from the Sequence Read Archive (<xref ref-type="bibr" rid="B29">Leinonen et&#x20;al., 2011</xref>), and with the aim of collecting a minimum of 15 individuals per breed for 10 target breeds (Arabian, Belgian, Clydesdale, Icelandic, Morgan, Quarter Horse, Shetland, Standardbred, Thoroughbred, and Welsh Pony) that represent major groups of worldwide equine genetic diversity (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>).</p>
</sec>
<sec id="s2-2">
<title>2.2 Mapping and Variant Calling</title>
<p>Standard fastq quality control and trimming were performed using Fastqc 0.11.8 and Trimmomatic 0.38, respectively. Raw reads were aligned to the EquCab 3.0 reference horse assembly (<xref ref-type="bibr" rid="B25">Kalbfleisch et&#x20;al., 2018</xref>) and variants identified using a modified version of the genome analysis toolkit best practices (<xref ref-type="bibr" rid="B1">Van der Auwera et&#x20;al., 2013</xref>), modified to allow for joint variant calling by GATK haplotype caller (<xref ref-type="bibr" rid="B39">McKenna et&#x20;al., 2010</xref>) and BCFtools mpileup (<xref ref-type="bibr" rid="B33">Li H. et&#x20;al., 2009</xref>). In brief, reads were mapped to the EquCab 3.0 reference genome using Burrows-Wheeler Aligner (BWA) (<xref ref-type="bibr" rid="B32">Li and Durbin 2009</xref>). PCR duplicates were detected and removed using Picard tools (<ext-link ext-link-type="uri" xlink:href="https://broadinstitute.github.io/picard/">https://broadinstitute.github.io/picard/</ext-link>) version 2.18.27 and then indel realignment was performed using the Genome Analysis Toolkit (GATK) version 4.1.0.0 indel realigner and base-quality score recalibration was performed (<xref ref-type="bibr" rid="B9">DePristo et&#x20;al., 2011</xref>). Genome variant call format (genome VCF) files were produced for each individual horse, and then group variant calling was performed using GATK haplotype caller version 4.1.0.0 (<xref ref-type="bibr" rid="B39">McKenna et&#x20;al., 2010</xref>) and BCFtools mpileup version 1.9 (<xref ref-type="bibr" rid="B33">Li H. et&#x20;al., 2009</xref>) using default settings. Hard-filtering was performed using the GATK best practice guidelines (<xref ref-type="bibr" rid="B1">Van der Auwera et&#x20;al., 2013</xref>). To maximize the specificity of the variants, the intersection of the variants across both callers was obtained using GATK &#x201c;SelectVariants&#x201d; and was used for downstream analysis (<xref ref-type="bibr" rid="B1">Van der Auwera et&#x20;al., 2013</xref>).</p>
</sec>
<sec id="s2-3">
<title>2.3 Variant Analyses</title>
<p>Descriptive statistics for both variant callers and the intersection were created using BCFtools (<xref ref-type="bibr" rid="B31">Li 2011</xref>). Missingness for each horse and each variant site was calculated using VCFtools. The transition to transversion (TsTv) and heterozygous to non-reference homozygous (hetNRhom) ratios were calculated for each variant caller, across the population and by breed using BCFtools (<xref ref-type="bibr" rid="B31">Li 2011</xref>) and Python. Predicted functional effect for each variant in the intersect file was determined using SnpEff (<xref ref-type="bibr" rid="B7">Cingolani et&#x20;al., 2012b</xref>) with a custom dictionary based on the RefSeq version of EquCab 3.0 (<xref ref-type="bibr" rid="B48">Schaefer et&#x20;al., 2017</xref>). High, moderate, and low impact variants were extracted using SnpSift (<xref ref-type="bibr" rid="B6">Cingolani et&#x20;al., 2012a</xref>) and Python was used to manipulate VCF files. Python and BCFtools (<xref ref-type="bibr" rid="B33">Li H. et&#x20;al., 2009</xref>) were used to manipulate output&#x20;files.</p>
</sec>
<sec id="s2-4">
<title>2.4 Breed Differences Between Variants</title>
<p>Variants that were considered rare (&#x3c;3%) or common (&#x3e;10%) were extracted from each breed VCF file using Python and BCFtools (<xref ref-type="bibr" rid="B33">Li H. et&#x20;al., 2009</xref>). Variants that were rare in one breed and common in another, rare/common or common/rare in the breed/population, or uniquely present in one breed were selected for investigation. BCFtools view was used to extract variants that were only present in homozygous states, or were present in all genotyped individuals. Python was to extract variant details.</p>
</sec>
<sec id="s2-5">
<title>2.5 Regions of the Genome With High or Low Genetic Variation</title>
<p>BCFtools stats was used on 10&#xa0;kb regions across the genome. R was used to calculate the average genetic variation and to find regions with high (more than twice the average variation) and low (less than half the average variation) levels of genetic variation. BCFtools view was used to extract these regions from the intersect. Python was then used to extract variant details and additional analysis performed with&#x20;R.</p>
</sec>
<sec id="s2-6">
<title>2.6 Statistical Analyses</title>
<p>A Kruskal Wallis test and linear regression were used to determine if there were breed differences. The nonlinear relationship between depth of coverage and the number of variants identified was determined using R. Due to the association between depth of coverage and the number of variants detected, estimated marginal means (EMMEANS) (<xref ref-type="bibr" rid="B30">Lenth 2018</xref>) were calculated, with depth of coverage included in the regression models. T tests were used to compare variant types and allele frequencies between coding and non-coding variants, and to compare constraint metric scores between high and low variation regions. A chi-square test was used to compare the variant impact between the high and low variation regions. Confidence intervals (95%) were calculated for each breed. All statistical analyses were performed using R. Significance was set at <italic>p</italic>&#x20;&#x3c;&#x20;0.05.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Variant Discovery</title>
<p>WGS of 534 horses across 46 different breeds (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S2</xref>), was performed on Illumina platforms (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S1</xref>). This sample set included DNA from a minimum of 15 horses in each of 10 target breeds (Arabian, Belgian, Clydesdale, Icelandic Horse, Morgan, Quarter Horse, Shetland Pony, Standardbred, Thoroughbred, and Welsh Pony) that represent major breed clusters in the horse (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>), which were sequenced to a target depth of coverage (DOC) of 10X. Raw reads were mapped, quality control was performed, and variants were identified using a modified version of the genome analysis toolkit best practices pipeline (<xref ref-type="bibr" rid="B1">Van der Auwera et&#x20;al., 2013</xref>) (see <xref ref-type="sec" rid="s2">Section 2</xref>). 155,201,208,820 total reads from the 534 horses mapped uniquely to the EquCab 3.0 reference genome (<xref ref-type="bibr" rid="B25">Kalbfleisch et&#x20;al., 2018</xref>). Mean and median read length, uniquely mapped paired reads, and depth of coverage, and ranges for these values are provided in <xref ref-type="table" rid="T1">Table&#x20;1</xref>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Median, mean, and range of summary statistics of the mapping pipeline derived from WGS data from 534 horses.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">Median</th>
<th align="center">Mean</th>
<th align="center">Range</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Read length (bp)</td>
<td align="center">99.6</td>
<td align="center">114.5</td>
<td align="center">73.9&#x2013;234.2</td>
</tr>
<tr>
<td align="left">Uniquely mapped paired end reads</td>
<td align="center">240,563,458</td>
<td align="center">290,638,968</td>
<td align="center">17,030,804&#x2013;1,536,494,934</td>
</tr>
<tr>
<td align="left">Depth of coverage (X)</td>
<td align="center">9.2</td>
<td align="center">11.5</td>
<td align="center">1.4&#x2013;46.7</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>GATK Haplotype Caller and BCFtools identified 42,900,494 and 33,395,275 variants, respectively. To increase specificity of the identified variants, the intersect of both variant callers [31,140,769 variants (29,038,030 SNPs and 2,102,379 indels)] was used for downstream analysis (<xref ref-type="table" rid="T2">Table&#x20;2</xref>). On average, there were 2.27 variants (range 0.88&#x2013;3.12) per kb of sequence. The mean number of heterozygous sites per individual per kb was 1.54 (range 0.56&#x2013;2.63). There was a significant non-linear association between the number of variants identified and the depth of coverage [DOC (correlation estimate 0.62, <italic>p</italic>
<sub>adjusted</sub> 0.009)]. The distribution of variants by DOC was similar for each breed (<xref ref-type="fig" rid="F1">Figure&#x20;1</xref>), therefore, estimated marginal means (EMMEANs) (<xref ref-type="bibr" rid="B30">Lenth 2018</xref>) accounting for DOC were used for further analyses. The median (range) degree of missingness for each individual horse was 0.01 (0.001&#x2013;0.58) and for each variant site was 0.026 (0.000&#x2013;0.998) (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>). Missingness by individual was moderately negatively correlated with DOC: correlation &#x2212;0.45, 95% confidence interval &#x2212;0.511 to &#x2212;0.38 and <italic>p</italic>&#x20;&#x3c; 0.001 (<xref ref-type="fig" rid="F2">Figure&#x20;2</xref>).</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Number of variants, TsTv, and HetNRhom ratio from 534 WGS identified by each variant caller [GATK Haplotype Caller (HC) and BCFtools (BT), and the union and intersection of the variant callers].</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Variant Caller</th>
<th align="center">Variants</th>
<th align="center">SNPs</th>
<th align="center">INDELs</th>
<th align="center">MA sites</th>
<th align="center">MA&#x20;SNP sites</th>
<th align="center">TsTv ratio</th>
<th align="center">HetNRhom ratio</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">HC</td>
<td align="center">42,900,494</td>
<td align="center">38,205,867</td>
<td align="center">4,694,627</td>
<td align="center">2,974,935</td>
<td align="center">2,127,391</td>
<td align="center">1.87</td>
<td align="center">2.48</td>
</tr>
<tr>
<td align="left">BT</td>
<td align="center">33,395,275</td>
<td align="center">30,642,613</td>
<td align="center">2,752,662</td>
<td align="center">783,956</td>
<td align="center">295,703</td>
<td align="center">2.08</td>
<td align="center">2.18</td>
</tr>
<tr>
<td align="left">Union</td>
<td align="center">45,154,996</td>
<td align="center">439,810,450</td>
<td align="center">5,344,546</td>
<td align="center">3,298,397</td>
<td align="center">2,178,474</td>
<td align="center">1.54</td>
<td align="center">2.17</td>
</tr>
<tr>
<td align="left">Intersect</td>
<td align="center">31,140,769</td>
<td align="center">29,038,030</td>
<td align="center">2,102,379</td>
<td align="center">1,547,737</td>
<td align="center">990,223</td>
<td align="center">1.94</td>
<td align="center">2.24</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>MA, multiallelic; TsTv, transition/transversion; HetNRhom, heterozygous-non-reference homozygous.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Distribution of the average number of variants identified for each breed by DOC quantiles (Q1 &#x3d; 1.43&#x2013;7.14 X, Q2 &#x3d; 7.15&#x2013;9.16 X, Q3 &#x3d; 9.17&#x2013;14.6 X, Q4 &#x3d; 14.7&#x2013;46.7 X). The colored lines represent the 10 target horse breeds [Arabian, Belgian, Clydesdale, Icelandic horse, Morgan horse, Quarter Horse (QH), Shetland, Standardbred (STB), Thoroughbred (TB), and Welsh Pony (WP)] and the remaining horse breeds (Other).</p>
</caption>
<graphic xlink:href="fgene-12-758366-g001.tif"/>
</fig>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Missingness by individual <bold>(A)</bold>, by depth of coverage <bold>(B)</bold> and by chromosome <bold>(C)</bold>. The 10 target breeds and other breeds are represented in colors shown in the figure legend in <xref ref-type="fig" rid="F2">Figures 2A,B</xref>.</p>
</caption>
<graphic xlink:href="fgene-12-758366-g002.tif"/>
</fig>
<p>The transition to transversion (TsTv) ratio and the heterozygous to non-reference homozygous (hetNRhom) ratios from the variant callers intersect were 1.94 and 2.24, respectively (<xref ref-type="table" rid="T2">Table&#x20;2</xref>). There were significant but marginal breed differences in the TsTv and hetNRhom ratios (<italic>p</italic>&#x20;&#x3c; 0.001), with the highest TsTv ratio in Shetlands (1.95) and the lowest in Thoroughbreds (1.92), The same was true with the hetNRhom ratio, which was the highest in Thoroughbreds (3.21) and the lowest in Clydesdales (1.48) (<xref ref-type="table" rid="T3">Table&#x20;3</xref>). The majority (57%) of variants had a minor allele frequency (MAF) &#x3c; 5%, <xref ref-type="fig" rid="F3">Figure&#x20;3</xref>), with the mean MAF being 13.2% (0.09&#x2013;100%). In total, there were 2,481,075 (2,447,610 SNPs and 33,465 indels) singleton variants.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Estimated marginal mean of the transition to transversion (TsTv) and heterozygous to non-reference homozygous (hetNRhom) ratios accounting for depth of coverage.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">&#xa0;Breed</th>
<th align="center">TsTv ratio</th>
<th align="center">hetNRhom ratio</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Arabian</td>
<td align="char" char=".">1.93</td>
<td align="char" char=".">1.53</td>
</tr>
<tr>
<td align="left">Belgian</td>
<td align="char" char=".">1.94</td>
<td align="char" char=".">2.02</td>
</tr>
<tr>
<td align="left">Clydesdale</td>
<td align="char" char=".">1.93</td>
<td align="char" char=".">1.68</td>
</tr>
<tr>
<td align="left">Icelandic</td>
<td align="char" char=".">1.94</td>
<td align="char" char=".">1.89</td>
</tr>
<tr>
<td align="left">Morgan</td>
<td align="char" char=".">1.94</td>
<td align="char" char=".">2.26</td>
</tr>
<tr>
<td align="left">Quarter Horse</td>
<td align="char" char=".">1.93</td>
<td align="char" char=".">2.26</td>
</tr>
<tr>
<td align="left">Shetland</td>
<td align="char" char=".">1.95</td>
<td align="char" char=".">2.42</td>
</tr>
<tr>
<td align="left">Standardbred</td>
<td align="char" char=".">1.93</td>
<td align="char" char=".">2.03</td>
</tr>
<tr>
<td align="left">Thoroughbred</td>
<td align="char" char=".">1.92</td>
<td align="char" char=".">3.19</td>
</tr>
<tr>
<td align="left">Welsh Pony</td>
<td align="char" char=".">1.94</td>
<td align="char" char=".">2.27</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Minor allele frequency distribution of the variants.</p>
</caption>
<graphic xlink:href="fgene-12-758366-g003.tif"/>
</fig>
<p>Each individual horse had on average 5,580,202 variants (5,099,978 SNPs and 480,224 indels), with on average 1,805,127 in homozygous and 3,775,075 in heterozygous states (<xref ref-type="table" rid="T4">Table&#x20;4</xref>). There were also breed-specific differences in variant number and homozygous variant number per individual (<xref ref-type="table" rid="T4">Table&#x20;4</xref>). The EMMEAN variants per individual, accounting for DOC, was lowest in Thoroughbreds (5,000,516) and highest in Belgians (6,100,544). The EMMEAN homozygous variants per individual, accounting for DOC, was also lowest in Thoroughbreds (1,225,441) and highest in Belgians (2,325,470).</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>EMMEANs for number of variants within breeds accounting for DOC, standard error (SE), and 95% confidence intervals, with breed and DOC as predictor variables.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">&#xa0;Breed</th>
<th align="center">Variants included</th>
<th align="center">EMMEAN</th>
<th align="center">SE</th>
<th align="center">Lower confidence interval</th>
<th align="center">Upper confidence interval</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Average</td>
<td align="left">All</td>
<td align="center">5,580,202</td>
<td align="center">33,642</td>
<td align="center">5,514,039</td>
<td align="center">5,646,364</td>
</tr>
<tr>
<td align="left">Arabian</td>
<td align="left">All</td>
<td align="center">5,587,952</td>
<td align="center">54,820</td>
<td align="center">5,480,322</td>
<td align="center">5,695,582</td>
</tr>
<tr>
<td align="left">Belgian</td>
<td align="left">All</td>
<td align="center">6,100,544</td>
<td align="center">67,559</td>
<td align="center">5,967,903</td>
<td align="center">6,233,186</td>
</tr>
<tr>
<td align="left">Clydesdale</td>
<td align="left">All</td>
<td align="center">5,977,299</td>
<td align="center">69,379</td>
<td align="center">5,841,084</td>
<td align="center">6,113,515</td>
</tr>
<tr>
<td align="left">Icelandic</td>
<td align="left">All</td>
<td align="center">6,093,155</td>
<td align="center">72,659</td>
<td align="center">5,950,500</td>
<td align="center">6,235,810</td>
</tr>
<tr>
<td align="left">Morgan</td>
<td align="left">All</td>
<td align="center">5,602,287</td>
<td align="center">67,541</td>
<td align="center">5,469,681</td>
<td align="center">5,734,893</td>
</tr>
<tr>
<td align="left">Quarter Horse</td>
<td align="left">All</td>
<td align="center">5,473,401</td>
<td align="center">39,427</td>
<td align="center">5,395,992</td>
<td align="center">5,550,810</td>
</tr>
<tr>
<td align="left">Shetland</td>
<td align="left">All</td>
<td align="center">5,645,668</td>
<td align="center">44,377</td>
<td align="center">5,558,541</td>
<td align="center">5,732,794</td>
</tr>
<tr>
<td align="left">Standardbred</td>
<td align="left">All</td>
<td align="center">5,623,603</td>
<td align="center">44,981</td>
<td align="center">5,535,289</td>
<td align="center">5,711,917</td>
</tr>
<tr>
<td align="left">Thoroughbred</td>
<td align="left">All</td>
<td align="center">5,000,516</td>
<td align="center">43,061</td>
<td align="center">4,915,972</td>
<td align="center">5,085,060</td>
</tr>
<tr>
<td align="left">Welsh Pony</td>
<td align="left">All</td>
<td align="center">5,836,729</td>
<td align="center">67,893</td>
<td align="center">5,703,431</td>
<td align="center">5,970,028</td>
</tr>
<tr>
<td align="left">Average</td>
<td align="left">Homozygous</td>
<td align="center">1,805,127</td>
<td align="center">10,931</td>
<td align="center">1,783,628</td>
<td align="center">1,826,626</td>
</tr>
<tr>
<td align="left">Arabian</td>
<td align="left">Homozygous</td>
<td align="center">1,812,878</td>
<td align="center">54,820</td>
<td align="center">1,705,248</td>
<td align="center">1,920,508</td>
</tr>
<tr>
<td align="left">Belgian</td>
<td align="left">Homozygous</td>
<td align="center">2,325,470</td>
<td align="center">67,559</td>
<td align="center">2,192,828</td>
<td align="center">2,458,111</td>
</tr>
<tr>
<td align="left">Clydesdale</td>
<td align="left">Homozygous</td>
<td align="center">2,202,225</td>
<td align="center">69,379</td>
<td align="center">2,066,010</td>
<td align="center">2,338,440</td>
</tr>
<tr>
<td align="left">Icelandic</td>
<td align="left">Homozygous</td>
<td align="center">2,318,081</td>
<td align="center">72,659</td>
<td align="center">2,175,426</td>
<td align="center">2,460,736</td>
</tr>
<tr>
<td align="left">Morgan</td>
<td align="left">Homozygous</td>
<td align="center">1,827,212</td>
<td align="center">67,541</td>
<td align="center">1,694,607</td>
<td align="center">1,959,818</td>
</tr>
<tr>
<td align="left">Quarter Horse</td>
<td align="left">Homozygous</td>
<td align="center">1,698,327</td>
<td align="center">39,427</td>
<td align="center">1,620,918</td>
<td align="center">1,775,735</td>
</tr>
<tr>
<td align="left">Shetland</td>
<td align="left">Homozygous</td>
<td align="center">1,870,593</td>
<td align="center">44,377</td>
<td align="center">1,783,466</td>
<td align="center">1,957,720</td>
</tr>
<tr>
<td align="left">Standardbred</td>
<td align="left">Homozygous</td>
<td align="center">1,848,529</td>
<td align="center">44,981</td>
<td align="center">1,760,215</td>
<td align="center">1,936,843</td>
</tr>
<tr>
<td align="left">Thoroughbred</td>
<td align="left">Homozygous</td>
<td align="center">1,225,441</td>
<td align="center">43,061</td>
<td align="center">1,140,897</td>
<td align="center">1,309,985</td>
</tr>
<tr>
<td align="left">Welsh Pony</td>
<td align="left">Homozygous</td>
<td align="center">2,061,655</td>
<td align="center">67,893</td>
<td align="center">1,928,356</td>
<td align="center">2,194,953</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The top half of the table provides the EMMEAN for the total number of variants per individual (All) and the bottom half of the table provides the total number of homozygous variants present in each individual.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The EMMEAN variant number per breed was significantly correlated with one estimation of effective population size across breeds (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>) (<italic>p</italic>&#x20;&#x3d; 0.02, Pearson&#x2019;s correlation &#x3d; 0.83, 95% confidence interval 0.22&#x2013;0.97), but not a more recent estimate of effective population size across breeds using a higher marker density (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>) (<italic>p</italic>&#x20;&#x3d; 0.54, Pearson&#x2019;s correlation &#x3d; 0.26, 95% confidence interval &#x2212;0.54&#x2013;0.82). The EMMEAN homozygous variant number per breed was not significantly correlated with either estimate (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>) of effective population size (<italic>p</italic>&#x20;&#x3d; 0.07, Pearson&#x2019;s correlation &#x3d; 0.72, 95% confidence interval &#x2212;0.08&#x2013;0.95 and <italic>p</italic>&#x20;&#x3d; 1, Pearson&#x2019;s correlation &#x3d; &#x2212;0.001, 95% confidence interval &#x2212;0.70&#x2013;0.70, respectively).</p>
</sec>
<sec id="s3-2">
<title>3.2 Variants Shared Across Breeds</title>
<p>In total, 27,719,724 variants (25,386,978 SNPs and 2,332,746 indels) were shared by at least two breeds (shared variants, <xref ref-type="sec" rid="s11">Supplementary Table&#x20;S3</xref>). Only 2% (637,610) of these variants were genic (within 5,000&#xa0;bp of a gene). Genic variants included 8,427 high (4,934 SNPs and 3,493 indels), 199,243 moderate (195,630 SNPs and 3,613 indels), and 429,940 low (427,436 SNPs and 2,504 indels) impact variants with 6,697 predicted to cause loss of function (LOF). Genic variants were within or near 10,013 individual genes. The mean shared variant MAF was 0.07. The mean shared variants per breed was 13,802,105 (range 11,800,825 in Clydesdales to 15,724,634 in Quarter Horses).</p>
</sec>
<sec id="s3-3">
<title>3.3 Variants With Large Allele Frequency Discrepancies</title>
<p>10,633,492 variants (121,900 genic) were considered rare (MAF &#x3c;3%) in one breed and common (MAF &#x3e;10%) in at least one other breed (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S4</xref>). Each breed had on average 1,216,800 variants that had large allele frequency discrepancies with another breed. The fewest number of variant discrepancies were between the Quarter Horse and Thoroughbred (728,459 variants) and the highest number of variant discrepancies were between the Thoroughbred and the Shetland (2,028,658 variants).</p>
<p>4,876,293 variants (190,653 genic) were rare (MAF &#x3c;3%) in one breed and common (MAF &#x3e;10%, mean MAF 0.15) in the remainder of the study population (excluding individuals of that breed) (breed-specific rare variants, <xref ref-type="sec" rid="s11">Supplementary Table&#x20;S4</xref>). Each genic variant was present in or close to at least one of 10,361 genes. The mean number of breed-specific rare variants was 475,458. The fewest number of variant discrepancies was between the Quarter Horse and the population (8,248 variants) and the highest number of variant discrepancies was between the Thoroughbred and the population (913,739).</p>
<p>3,563,454 variants (62,803 coding) were common (MAF &#x3e;10%) in one breed and rare (MAF &#x3c;3%, mean MAF 0.02) in the remaining population (excluding individuals of that breed) (breed-specific common variants, <xref ref-type="sec" rid="s11">Supplementary Table&#x20;S4</xref>). On average, each of these variants was present in 1.13 breeds (range 1&#x2013;5) and each genic variant was shared by 1.11 breeds (range 1&#x2013;4). Genic variants were present in or close to at least one of 10,163 genes. On average, each breed had 444,464 variants that were common in the breed and rare in the population. The fewest number of variant discrepancies was between the Thoroughbred and the population (115,241 variants) and the highest number of variant discrepancies was between the Icelandic horse and the population (745,623).</p>
</sec>
<sec id="s3-4">
<title>3.4 Variants With No Homozygotes Present</title>
<p>2,889 variants (2,586 SNPs, 303 indels) were only present in a heterozygous state, with no homozygotes identified. Twenty-six of these variants were present within 14 different genes. 12/14 of these genes are uncharacterized or equine specific transcripts or olfactory related genes. Six were predicted to be high impact (four were predicted to be LOF variants), 10 were predicted to be moderate impact, and 10 were predicted to be low impact variants (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S5</xref>). The allele frequency was marginally but significantly different (<italic>p</italic>&#x20;&#x3d; 0.003) between non-genic (0.498) and genic (0.499) variants.</p>
</sec>
<sec id="s3-5">
<title>3.5 Variants Present in all Horses</title>
<p>114,733 variants (103,414 SNPs, 11,319 indels) were present in all 534 horses. Of these variants, 1,426 were present in genic regions of 504 genes. These variants were predicted to have a high (170 variants), moderate (644 variants) and low (612 variants) impact on phenotype, with 145 predicted to be LOF variants. Of these variants, 9,756 (9,351 SNPs, 405 indels) were homozygous in all horses. Ninety-two of these were present in genic regions affecting 58 genes. Eight were predicted to be high impact (four were predicted to be LOF), 46 were predicted to be moderate impact, and 38 were predicted to be low impact variants.</p>
</sec>
<sec id="s3-6">
<title>3.6 Regions of the Genome With High or Low Genetic Variation</title>
<p>The variant caller intersect file was first split by chromosome and then into 10,000&#xa0;bp windows (240,910 regions in total) to determine the average number of variants per 10&#xa0;Kb region across the genome. Regions with more than two times or less than half of the average number of variants were classified as regions of high or low variability, respectively. Each 10&#xa0;Kb region carried on average 122 variants (range 0&#x2013;3,143 variants), consisting of 114 SNPs (range 0&#x2013;3,075&#xa0;SNPs) and 8 indels (range 0&#x2013;121). There were 6,341 regions with more than double the mean variant number, including, on average 414 variants (396 SNPs and 18 indels) with a TsTv ratio of 1.76. There were 17,791 regions with less than half the average number of variants, including, on average 35 variants (32 SNPs and three indels) with a TsTv ratio of&#x20;1.59.</p>
<p>Highly variable regions contained 2,625,382 variants (<xref ref-type="sec" rid="s11">Supplementary Table&#x20;S6</xref>) with a mean MAF of 17% (range 0.01&#x2013;100%). The most common variant types in the high variability regions were intergenic (1,287,507) and intronic (574,441) (<xref ref-type="fig" rid="F4">Figure&#x20;4</xref>). The low variability regions contained 20,777 variants with a mean MAF of 11% (range 0.01&#x2013;100%). The most common variant types in the low variation regions were intronic (306,569) and intergenic (182,937) variants (<xref ref-type="fig" rid="F4">Figure&#x20;4</xref>). The variant impact was significantly different between regions of high and low variation [<italic>p</italic>&#x20;&#x3c; 0.001 (<xref ref-type="table" rid="T5">Table&#x20;5</xref>)].</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Percentage of coding variants for each type of variant called by SnpEff for low variation regions (orange) and high variation regions (teal).</p>
</caption>
<graphic xlink:href="fgene-12-758366-g004.tif"/>
</fig>
<table-wrap id="T5" position="float">
<label>TABLE 5</label>
<caption>
<p>Impact of variants identified in high and low variation regions.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th align="center">High</th>
<th align="center">Moderate</th>
<th align="center">Low</th>
<th align="center">Modifier</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">
<bold>High variation region</bold>
</td>
<td align="char" char=".">2061</td>
<td align="center">39,260</td>
<td align="center">48,234</td>
<td align="center">2,535,827</td>
</tr>
<tr>
<td align="left">
<bold>Low variation region</bold>
</td>
<td align="char" char=".">303</td>
<td align="center">6,293</td>
<td align="center">11,521</td>
<td align="center">602,660</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The allele frequency of the variants in the low variability regions was significantly less than the allele frequency of variants in high variability regions (<italic>p</italic>&#x20;&#x3c; 0.001, 95% confidence interval: 0.06&#x2013;0.06). For the genes with haploinsufficiency scores available, the mean score was lower (<italic>p</italic>&#x20;&#x3c; 0.001, 95% confidence interval: &#x2212;0.15 to &#x2212;0.10) in genes containing variants in the high variability regions (0.23) compared to genes containing variants in the low variability regions (0.33). The predicted loss of function tolerance score was not significantly different (<italic>p</italic>&#x20;&#x3d; 0.13, 95% confidence interval: &#x2212;0.07&#x2013;0.01) between genes containing variants in the high variability regions (0.33) compared to genes containing variants in the low variability regions (0.37).</p>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>This report comprises the first large-scale catalog of genetic variation developed for the horse, a species with potential as a translational model for many athletic phenotypes. We used WGS of 534 horses to determine overall genetic variation in the general equine population as well as 10 individual breeds, report variants with population and breed MAF discrepancies, identify variants with no homozygotes as well as variants that are present in all individuals, and in addition identify genomic regions with high or low genetic variation.</p>
<p>The significant association between the number of variants identified and the depth of coverage was unsurprising, as it has long been recognized that deeper coverage improves variant calling accuracy (<xref ref-type="bibr" rid="B34">Li R. et&#x20;al., 2009</xref>). The nonlinear association between depth of coverage and number of variants identified is particularly relevant to future genetic studies, as there is minimal to no gain in the numbers of variants detected by increasing depth of coverage over&#x20;10X.</p>
<p>We elected to use the intersect of two commonly used callers based on evidence from previous work showing that specificity of identified variants could be improved (<xref ref-type="bibr" rid="B14">Field et&#x20;al., 2015</xref>). This method identified 29,882,273 variants, which is higher than the 25,800,000 variants previously identified in a smaller cohort of 88 horses using only GATK haplotype caller (<xref ref-type="bibr" rid="B24">Jagannathan et&#x20;al., 2019b</xref>). This difference is likely due to the increased number of horses and inclusion of more genetically distinct breeds in our study, increasing our ability to identify rare variants down to a MAF across breeds of 0.0009 compared with 0.0057 in the earlier study. Additionally, we use the intersect of two variant callers (GATK HaplotypeCaller and BCFtools) to improve specificity of variants identified in the previous publication (<xref ref-type="bibr" rid="B24">Jagannathan et&#x20;al., 2019b</xref>). While this likely would have led to a reduced sensitivity, we were still able to identify an additional 4,082,273 variants.</p>
<p>The 1.54 heterozygous variants per kb of sequence is similar to that reported in cattle (1.44 per kb) (<xref ref-type="bibr" rid="B8">Daetwyler et&#x20;al., 2014</xref>), but is higher than reported in the Yoruba (1.03 per kb) and European human populations (0.68 per kb) (<xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>). The cause of this higher variation in the horse is likely related to the heterogeneity of this population, which included 46 different breeds, compared to three cattle breeds (<xref ref-type="bibr" rid="B8">Daetwyler et&#x20;al., 2014</xref>) and only a single population in the human studies (<xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>). Given the limited genetic diversity (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>) of the horse compared with human populations this is still somewhat somewhat surprising. However, a recent study of effective population size in the horse suggests that several breeds have larger effective population sizes (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>) than reported in human populations (<xref ref-type="bibr" rid="B54">Tenesa et&#x20;al., 2007</xref>), and we would therefore, expect to see increased heterozygosity. Another reason for the increased number of heterozygous variants in horses is likely to be related to errors in the reference genome which, unlike the human reference genome that is based on multiple individuals, is based on a single horse (<xref ref-type="bibr" rid="B25">Kalbfleisch et&#x20;al., 2018</xref>). In the original paper from the 1,000 human genomes consortium, it was concluded that a site where every individual was homozygous for an allele not present in the reference genome was a reference genome error. This accounted for &#x223C;1 error per 30&#xa0;kb of sequence. In this study, we identified 114,733 variants that were present in every individual in this population, which are presumed to be related to an error in the reference genome or a true rare variant that is present in the individual horse sequenced for the reference genome. The length of the RefSeq equine reference genome is 2,506.95 MB, therefore, we would expect to see &#x223C;1.37 errors per 30&#xa0;kb of sequence, which may partially explain the increased number of variants in horses compared with humans.</p>
<p>Unsurprisingly, the degree of missingness per individual was negatively correlated with depth of coverage. This was not a linear correlation however, and beyond 10X coverage, there was minimal improvement in the degree of missingness, suggesting that 10X coverage is a reasonable target for population scale sequencing projects. The average missingness per individual varied greatly. Most of the horses with missingness &#x3e;0.20 were horses that were not included in the breed analysis owing to being in the other breed category. Six of the Shetland ponies had missingness &#x3e;0.20 and these were ponies with a targeted depth of coverage of &#x223C; 6X. The degree of missingness across the 32 autosomes and X chromosomes also varied, with the highest degree of missingness on chromosomes 12 and X. This may be related to larger uncharacterized regions on the X chromosome compared to most autosomes.</p>
<p>The TsTv and hetNRhom ratios are frequently utilized for quality control of sequencing data (<xref ref-type="bibr" rid="B18">Guo et&#x20;al., 2013</xref>). The TsTv ratio here was similar (1.94) to reports of the expected TsTv ratio from genome sequencing data in humans of &#x223C;2 (<xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>), which is a good indicator of SNP quality (<xref ref-type="bibr" rid="B18">Guo et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B56">Wang et&#x20;al., 2015</xref>). The slight reduction in our study compared with human studies is likely related to decreased genetic diversity and smaller effective population sizes in horses (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>) compared with humans, as well as a lower quality reference genome (<xref ref-type="bibr" rid="B55">Wade et&#x20;al., 2009</xref>). The TsTv ratio varies across the genome, but in humans does not vary based on ancestry (<xref ref-type="bibr" rid="B56">Wang et&#x20;al., 2015</xref>). In the 10 target breeds, we did find that the TsTv ratio varied significantly but marginally by ancestry. This may be related to the varying depth of sequence coverage, which was not uniform across breeds. The hetNRhom ratio (2.24) was higher than the expected value of 2 based on Hardy-Weinberg equilibrium (<xref ref-type="bibr" rid="B18">Guo et&#x20;al., 2013</xref>) and there were breed differences. This is consistent with human (<xref ref-type="bibr" rid="B56">Wang et&#x20;al., 2015</xref>) and canine (<xref ref-type="bibr" rid="B23">Jagannathan et&#x20;al., 2019a</xref>) reports that the hetNRhom ratio varies by ancestry. In humans, the highest median hetNRhom ratio was 2.0 in African populations with the lowest ratio of 1.4 in Asian populations; none of the populations investigated had a median hetNRhom ratio &#x3e;2.0 (<xref ref-type="bibr" rid="B56">Wang et&#x20;al., 2015</xref>). However, at least one dog breed had hetNRhom ratio of 3.3 in the catalog of canine genetic variation (<xref ref-type="bibr" rid="B23">Jagannathan et&#x20;al., 2019a</xref>). This is likely related to the increased levels of inbreeding in horses compared with most human populations.</p>
<p>The variant totals differed by breed, which is consistent with reports in different cattle breeds (<xref ref-type="bibr" rid="B8">Daetwyler et&#x20;al., 2014</xref>) and regional human populations (<xref ref-type="bibr" rid="B16">Genomes Project et&#x20;al., 2010</xref>). Previously, this has been related to effective population size (<xref ref-type="bibr" rid="B8">Daetwyler et&#x20;al., 2014</xref>), and we did see a significant association (<italic>p</italic>&#x20;&#x3c; 0.0001) between the number of variants identified in each of the 10 target breeds and a report of effective population size using 54K SNP array data (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>). However, this association was not seen with a more recent estimate of effective population size that used imputed genome-wide SNP data (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>). This may partially be related to different breeds studied, as the <xref ref-type="bibr" rid="B3">Beeson et&#x20;al. (2019)</xref> paper did not include effective population size estimates for the Shetland and Clydesdale breeds. Additionally, we are not accounting for the degree of relatedness between the breeds studied and the reference genome. In this study, the Thoroughbred (which is the EquCab3 reference genome breed) has the fewest variants compared to the other breeds, consistent with both the reference genome being from a Thoroughbred and having the smallest effective population size (1,784) in Beeson et&#x20;al. (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>). However, the Quarter Horse, which has the largest population size (6,516) (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>), but is more related to the Thoroughbred (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>) than other breeds, has a number of variants that is closer to the median, consistent with its close relatedness to the Thoroughbred (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>), but inconsistent with the large effective population size (<xref ref-type="bibr" rid="B3">Beeson et&#x20;al., 2019</xref>). This would suggest, that while the number of variants does vary by breed, as seen in different human regional populations, the relationship between the horse breed and the reference genome appears to have had an effect in this population.</p>
<p>A large number of variants were shared by at least one other breed, which is not surprising given the close relatedness of the breeds investigated (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>). However, there were also multiple variants with large allele frequency discrepancies between breeds. We defined a minor allele frequency &#x3c;3% as rare and &#x3e;10% as common due to limitations in the number of horses in each breed investigated, rather than values used in most human studies (&#x3c;0.5% for rare variants and &#x3e;5% for common (<xref ref-type="bibr" rid="B5">Bomba et&#x20;al., 2017</xref>)). With the 534 horses here our power to detect variants present at a minor allele frequency of 3 and 0.5% in the population is 1.00 and 0.93, respectively. However, it is important to note that for the breed analysis with the least number of horses in the Clydesdales (19) our power to detect these allele frequencies is only 0.44 and 0.01, respectively. To have a power greater than 0.8 to detect all rare variants &#x3c;3% or &#x3c;0.5% allele frequency within a breed we would need to sequence 55 and 325 horses within that breed, respectively. We therefore, had 80% power to detect the rare variants (minor allele frequency 3%) in Quarter Horses, Shetlands and Thoroughbreds. The number of variants with marked population discrepancies was quite high, with &#x223C;35% of variants considered rare in one breed and common in another. The reason for large allele frequency differences between populations is thought to be related to genetic drift (<xref ref-type="bibr" rid="B21">Hofer et&#x20;al., 2009</xref>) in humans. However, given that different horse breeds have been selectively bred for different traits (<xref ref-type="bibr" rid="B2">Avila et&#x20;al., 2018</xref>), it is also likely that selection at least partially accounts for some of the variants with large allele frequency discrepancies between breeds. Only &#x223C;5% of variants were unique to a single breed, with about 10% of these being in coding regions. The &#x223C;5% of unique variants in horses is also lower than seen in cattle where &#x223C;31% of variants are unique to a single breed (<xref ref-type="bibr" rid="B8">Daetwyler et&#x20;al., 2014</xref>).</p>
<p>Variants with no homozygotes were explored to determine if there were variants present in the general population that could be embryonic lethal in homozygous form. 2,888/2,889 of these variants were present at a minor allele frequency in the general population greater than would be expected for a homozygous lethal disease (MAF &#x3e; 0.10). The one variant that was rare had a MAF of 0.03 and was an intergenic in frame deletion. Using the rules of Hardy-Weinberg equilibrium we would need to sequence 1,111 horses to identify just one homozygote, therefore using this dataset&#x20;alone, we cannot determine if this variant is embryonic lethal. Given that 96% of the variants with no homozygotes have an allele frequency around 0.50, it is highly unlikely that these variants are embryonic lethal; rather it is possible that these regions are related to mapping errors due to the presence of paralogs or pseudogenes, or the presence of structural variants. This is supported by the fact that all but two of the genes containing these variants were uncharacterized or equine only transcripts or olfactory receptor&#x20;genes.</p>
<p>Identifying regions of the genome with high or low variation in the general population is critically important for the investigation of possible disease-causing variants, as regions with high genetic variability are less likely to contain disease-causing variants for fully penetrant Mendelian diseases (<xref ref-type="bibr" rid="B27">Karczewski et&#x20;al., 2019</xref>). However, our analysis unexpectantly found that genes in high variability regions had lower haploinsufficiency scores, suggesting that damaging variants are less well tolerated (<xref ref-type="bibr" rid="B22">Huang et&#x20;al., 2010</xref>) than for genes present in the low variability regions. This is likely due to the large number of genes without haploinsufficiency scores (70% of high variation region genes and 23% of low variation region genes). 65% of genes in high and 18% of genes in low variation regions had the &#x201c;LOC&#x201d; designation that are genes unique to the horse or are only predicted to be equivalent to a human gene, and therefore would not be included in human databases of variant constraint. This would suggest that as expected, genes in the low variability regions are more similar to human genes than genes in the high variability regions.</p>
<p>Overall, this is the first large-scale catalog of genetic variation in the domestic horse and will be highly useful for evaluation of background genetic variation in any future genetic study. This catalog has paved the way for future investigation of the regions of the genome that are shared or have marked MAF discrepancies across breeds. Regions with marked discrepancies between breeds, or between a breed and the population, can then be interrogated to further look for signatures of selection both across breeds, as well as within breeds. Additionally, further investigation into which of the variants that are present in all individuals are related to poor genome annotation or instead are true variants in the reference genome is needed and will lead to improvement of the horse reference genome in the future. By improving knowledge of the poorly annotated regions of the horse genome we will be able to correct these for future versions of the reference genome. This will have benefits for future phenotype-causing variant identification studies. Given the relatively unique utility of the horse as a model for human athletic related traits (<xref ref-type="bibr" rid="B20">Hill et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B47">Rooney et&#x20;al., 2018</xref>) and diseases (<xref ref-type="bibr" rid="B57">Ward et&#x20;al., 2004</xref>; <xref ref-type="bibr" rid="B37">McCue et&#x20;al., 2008</xref>; <xref ref-type="bibr" rid="B38">McIlwraith et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B36">McCoy et&#x20;al., 2013</xref>; <xref ref-type="bibr" rid="B43">Norton et&#x20;al., 2016</xref>), an improved ability to identify phenotype-causing variants in the horse may shed light on analogous genetic diseases in humans. This is especially true for complex phenotypes such as athleticism, osteoarthritis (<xref ref-type="bibr" rid="B38">McIlwraith et&#x20;al., 2012</xref>), and exertional rhabdomyolysis (<xref ref-type="bibr" rid="B43">Norton et&#x20;al., 2016</xref>) where the limited genetic diversity in the horse (<xref ref-type="bibr" rid="B44">Petersen et&#x20;al., 2013</xref>) may further accelerate our ability to identify the&#x20;true phenotype-causing variants. This improvement in the reference genome combined with an improved understanding of the background genetic variation (<xref ref-type="bibr" rid="B58">Yang et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B13">Farwell et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B12">Ellingford et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B19">Hartmannov&#xe1; et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B26">K&#xe4;ns&#xe4;koski et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B28">Kojima et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B41">Noll et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B51">Smedley et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B10">Dolzhenko et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B11">Eldomery et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B49">Schneider et&#x20;al., 2017</xref>) in the horse should vastly increase the identification of phenotype-causing variants for important equine diseases.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The variant datasets presented in this study can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/bioproject/PRJEB47918">https://www.ncbi.nlm.nih.gov/bioproject/PRJEB47918</ext-link>.</p>
</sec>
<sec id="s6">
<title>Ethics Statement</title>
<p>Ethical review and approval was not required for this animal study because this study used samples previously collected by our lab and collaborators with institutional ethics review and approval and written consent from owners for the participation of their horses this study. The remainder of the samples were publicly available.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>SD-A was involved in the grant-writing, study design, data analysis, and manuscript preparation. MM and JM were responsible for the grant-writing and study design. They also supervised and provided expertise for the data analysis and manuscript preparation. RS, BG, and WC assisted with developing and running the mapping and variant calling pipelines. All authors approved the manuscript prior to submission to the journal.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This work was supported by: USDA NIFA-AFRI Project 2017-67015-26296: Tools to Link Phenotype to Genotype in the Horse, The American Quarter Horse Association, and a University of Minnesota Multistate grant. Salary support for SD-A was provided by an American College of Veterinary Internal Medicine Foundation fellowship, by a T32 Institutional Training Grant in Comparative Medicine and Pathology (5T320D010993-12), by the 2019 Elaine and Bertram Klein Development Award, and a Morris Animal Foundation Postdoctoral Research Fellowship (D20EQ-403).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>Authors BG and WC own and work for IntervalBio LLC, the computational company that was compensated to develop the mapping and variant calling pipeline.</p>
<p>The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors acknowledge the Minnesota Supercomputing Institute (MSI) at the University of Minnesota for providing resources that contributed to the research results reported within this paper. URL: <ext-link ext-link-type="uri" xlink:href="http://www.msi.umn.edu/">http://www.msi.umn.edu</ext-link>.</p>
</ack>
<sec id="s11">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.758366/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2021.758366/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.docx" id="SM1" mimetype="application/docx" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Auwera</surname>
<given-names>G. A.</given-names>
</name>
<name>
<surname>Carneiro</surname>
<given-names>M. O.</given-names>
</name>
<name>
<surname>Hartl</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Poplin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Del Angel</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Levy&#x2010;Moonshine</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>From FastQ Data to High&#x2010;Confidence Variant Calls: The Genome Analysis Toolkit Best Practices Pipeline</article-title>. <source>Curr. Protoc. Bioinformatics</source> <volume>43</volume>, <fpage>111</fpage>&#x2013;<lpage>1033</lpage>. <pub-id pub-id-type="doi">10.1002/0471250953.bi1110s43</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Avila</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Schaefer</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>McCue</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Genome-wide Signatures of Selection Reveal Genes Associated with Performance in American Quarter Horse Subpopulations</article-title>. <source>Front. Genet.</source> <volume>9</volume>. <pub-id pub-id-type="doi">10.3389/fgene.2018.00249</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Beeson</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>McCue</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Exploration of fine-scale Recombination Rate Variation in the Domestic Horse</article-title>. <source>Genome Res.</source> <volume>29</volume>, <fpage>1744</fpage>&#x2013;<lpage>1752</lpage>. <pub-id pub-id-type="doi">10.1101/gr.243311.118</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bellone</surname>
<given-names>R. R.</given-names>
</name>
<name>
<surname>Liu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Petersen</surname>
<given-names>J.&#x20;L.</given-names>
</name>
<name>
<surname>Mack</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Singer-Berk</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Dr&#xf6;gem&#xfc;ller</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>A Missense Mutation in Damage-specific DNA Binding Protein 2 Is a Genetic Risk Factor for Limbal Squamous Cell Carcinoma in Horses</article-title>. <source>Int. J.&#x20;Cancer</source> <volume>141</volume>, <fpage>342</fpage>&#x2013;<lpage>353</lpage>. <pub-id pub-id-type="doi">10.1002/ijc.30744</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bomba</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Walter</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Soranzo</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>The Impact of Rare and Low-Frequency Genetic Variants in Common Disease</article-title>. <source>Genome Biol.</source> <volume>18</volume>. <pub-id pub-id-type="doi">10.1186/s13059-017-1212-4</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cingolani</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Patel</surname>
<given-names>V. M.</given-names>
</name>
<name>
<surname>Coon</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Land</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Ruden</surname>
<given-names>D. M.</given-names>
</name>
<etal/>
</person-group> (<year>2012a</year>). <article-title>Using <italic>Drosophila melanogaster</italic> as a Model for Genotoxic Chemical Mutational Studies with a New Program, SnpSift</article-title>. <source>Front. Gene</source> <volume>3</volume>, <fpage>35</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2012.00035</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cingolani</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Platts</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Coon</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Nguyen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2012b</year>). <article-title>A Program for Annotating and Predicting the Effects of Single Nucleotide Polymorphisms, SnpEff</article-title>. <source>Fly</source> <volume>6</volume>, <fpage>80</fpage>&#x2013;<lpage>92</lpage>. <pub-id pub-id-type="doi">10.4161/fly.19695</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daetwyler</surname>
<given-names>H. D.</given-names>
</name>
<name>
<surname>Capitan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pausch</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Stothard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>van Binsbergen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Br&#xf8;ndum</surname>
<given-names>R. F.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Whole-genome Sequencing of 234 Bulls Facilitates Mapping of Monogenic and Complex Traits in Cattle</article-title>. <source>Nat. Genet.</source> <volume>46</volume>, <fpage>858</fpage>&#x2013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1038/ng.3034</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>DePristo</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Banks</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Poplin</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Garimella</surname>
<given-names>K. V.</given-names>
</name>
<name>
<surname>Maguire</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Hartl</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>A Framework for Variation Discovery and Genotyping Using Next-Generation DNA Sequencing Data</article-title>. <source>Nat. Genet.</source> <volume>43</volume>, <fpage>491</fpage>&#x2013;<lpage>498</lpage>. <pub-id pub-id-type="doi">10.1038/ng.806</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dolzhenko</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>van Vugt</surname>
<given-names>J.&#x20;J.&#x20;F. A.</given-names>
</name>
<name>
<surname>Shaw</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Bekritsky</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Van Blitterswijk</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Narzisi</surname>
<given-names>G.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Detection of Long Repeat Expansions from PCR-free Whole-Genome Sequence Data</article-title>. <source>Genome Res.</source> <volume>27</volume>, <fpage>1895</fpage>&#x2013;<lpage>1903</lpage>. <pub-id pub-id-type="doi">10.1101/gr.225672.117</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eldomery</surname>
<given-names>M. K.</given-names>
</name>
<name>
<surname>Coban-Akdemir</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Harel</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Rosenfeld</surname>
<given-names>J.&#x20;A.</given-names>
</name>
<name>
<surname>Gambin</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Stray-Pedersen</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Lessons Learned from Additional Research Analyses of Unsolved Clinical Exome Cases</article-title>. <source>Genome Med.</source> <volume>9</volume>. <pub-id pub-id-type="doi">10.1186/s13073-017-0412-6</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ellingford</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Barton</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bhaskar</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>S. G.</given-names>
</name>
<name>
<surname>Sergouniotis</surname>
<given-names>P. I.</given-names>
</name>
<name>
<surname>O&#x27;Sullivan</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Whole Genome Sequencing Increases Molecular Diagnostic Yield Compared with Current Diagnostic Testing for Inherited Retinal Disease</article-title>. <source>Ophthalmology</source> <volume>123</volume>, <fpage>1143</fpage>&#x2013;<lpage>1150</lpage>. <pub-id pub-id-type="doi">10.1016/j.ophtha.2016.01.009</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Farwell</surname>
<given-names>K. D.</given-names>
</name>
<name>
<surname>Shahmirzadi</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>El-Khechen</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Powis</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Chao</surname>
<given-names>E. C.</given-names>
</name>
<name>
<surname>Tippin Davis</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Enhanced Utility of Family-Centered Diagnostic Exome Sequencing with Inheritance Model-Based Analysis: Results from 500 Unselected Families with Undiagnosed Genetic Conditions</article-title>. <source>Genet. Med.</source> <volume>17</volume>, <fpage>578</fpage>&#x2013;<lpage>586</lpage>. <pub-id pub-id-type="doi">10.1038/gim.2014.154</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Field</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Cho</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Andrews</surname>
<given-names>T. D.</given-names>
</name>
<name>
<surname>Goodnow</surname>
<given-names>C. C.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Reliably Detecting Clinically Important Variants Requires Both Combined Variant Calls and Optimized Filtering Strategies</article-title>. <source>PLoS ONE</source> <volume>10</volume>, <fpage>e0143199</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0143199</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Finno</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Stevens</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Affolter</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Joshi</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Ramsay</surname>
<given-names>S.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>SERPINB11 Frameshift Variant Associated with Novel Hoof Specific Phenotype in Connemara Ponies</article-title>. <source>Plos Genet.</source> <volume>11</volume>, <fpage>e1005122</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1005122</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Genomes Project</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Abecasis</surname>
<given-names>G. R.</given-names>
</name>
<name>
<surname>Altshuler</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Auton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Brooks</surname>
<given-names>L. D.</given-names>
</name>
<name>
<surname>Durbin</surname>
<given-names>R. M.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>A Map of Human Genome Variation from Population-Scale Sequencing</article-title>. <source>Nature</source> <volume>467</volume>, <fpage>1061</fpage>&#x2013;<lpage>1073</lpage>. <pub-id pub-id-type="doi">10.1038/nature09534</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Genomes Project</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Abecasis</surname>
<given-names>G. R.</given-names>
</name>
<name>
<surname>Auton</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Brooks</surname>
<given-names>L. D.</given-names>
</name>
<name>
<surname>DePristo</surname>
<given-names>M. A.</given-names>
</name>
<name>
<surname>Durbin</surname>
<given-names>R. M.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>An Integrated Map of Genetic Variation from 1,092 Human Genomes</article-title>. <source>Nature</source> <volume>491</volume>, <fpage>56</fpage>&#x2013;<lpage>65</lpage>. <pub-id pub-id-type="doi">10.1038/nature11632</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Sheng</surname>
<given-names>Q.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Samuels</surname>
<given-names>D. C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Three-stage Quality Control Strategies for DNA Re-sequencing Data</article-title>. <source>Brief. Bioinform.</source> <volume>15</volume>, <fpage>879</fpage>&#x2013;<lpage>889</lpage>. <pub-id pub-id-type="doi">10.1093/bib/bbt069</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hartmannov&#xe1;</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Piherov&#xe1;</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Tauchmannov&#xe1;</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kidd</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Acott</surname>
<given-names>P. D.</given-names>
</name>
<name>
<surname>Crocker</surname>
<given-names>J.&#x20;F. S.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Acadian Variant of Fanconi Syndrome Is Caused by Mitochondrial Respiratory Chain Complex I Deficiency Due to a Non-coding Mutation in Complex I Assembly Factor NDUFAF6</article-title>. <source>Hum. Mol. Genet.</source> <volume>25</volume>, <fpage>4062</fpage>&#x2013;<lpage>4079</lpage>. <pub-id pub-id-type="doi">10.1093/hmg/ddw245</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hill</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>McGivney</surname>
<given-names>B. A.</given-names>
</name>
<name>
<surname>Gu</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Whiston</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Machugh</surname>
<given-names>D. E.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>A Genome-wide SNP-Association Study Confirms a Sequence Variant (g.66493737C&#x3e;T) in the Equine Myostatin (MSTN) Gene as the Most Powerful Predictor of Optimum Racing Distance for Thoroughbred Racehorses</article-title>. <source>BMC genomics</source> <volume>11</volume>, <fpage>552</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-11-552</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hofer</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ray</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wegmann</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Excoffier</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Large Allele Frequency Differences between Human continental Groups Are More Likely to Have Occurred by Drift during Range Expansions Than by Selection</article-title>. <source>Ann. Hum. Genet.</source> <volume>73</volume>, <fpage>95</fpage>&#x2013;<lpage>108</lpage>. <pub-id pub-id-type="doi">10.1111/j.1469-1809.2008.00489.x</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Marcotte</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Hurles</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Characterising and Predicting Haploinsufficiency in the Human Genome</article-title>. <source>Plos Genet.</source> <volume>6</volume>, <fpage>e1001154</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1001154</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jagannathan</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Dr&#xf6;gem&#xfc;ller</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Leeb</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Aguirre</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Andr&#xe9;</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Bannasch</surname>
<given-names>D.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). <article-title>A Comprehensive Biomedical Variant Catalogue Based on Whole Genome Sequences of 582 Dogs and Eight Wolves</article-title>. <source>Anim. Genet.</source> <volume>50</volume>, <fpage>695</fpage>&#x2013;<lpage>704</lpage>. <pub-id pub-id-type="doi">10.1111/age.12834</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jagannathan</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Gerber</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Rieder</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Tetens</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Thaller</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Dr&#xf6;gem&#xfc;ller</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2019b</year>). <article-title>Comprehensive Characterization of Horse Genome Variation by Whole-Genome Sequencing of 88 Horses</article-title>. <source>Anim. Genet.</source> <volume>50</volume>, <fpage>74</fpage>&#x2013;<lpage>77</lpage>. <pub-id pub-id-type="doi">10.1111/age.12753</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Kalbfleisch</surname>
<given-names>T. S.</given-names>
</name>
<name>
<surname>Rice</surname>
<given-names>E. S.</given-names>
</name>
<name>
<surname>DePriest</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Walenz</surname>
<given-names>B. P.</given-names>
</name>
<name>
<surname>Hestand</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Vermeesch</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <source>EquCab3, an Updated Reference Genome for the Domestic Horse</source>. <comment>bioRxiv</comment>. <pub-id pub-id-type="doi">10.1101/306928</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>K&#xe4;ns&#xe4;koski</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>J&#xe4;&#xe4;skel&#xe4;inen</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>J&#xe4;&#xe4;skel&#xe4;inen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tommiska</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Saarinen</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lehtonen</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Complete Androgen Insensitivity Syndrome Caused by a Deep Intronic Pseudoexon-Activating Mutation in the Androgen Receptor Gene</article-title>. <source>Sci. Rep.</source> <volume>6</volume>. <pub-id pub-id-type="doi">10.1038/srep32819</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Karczewski</surname>
<given-names>K. J.</given-names>
</name>
<name>
<surname>Francioli</surname>
<given-names>L. C.</given-names>
</name>
<name>
<surname>Tiao</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Cummings</surname>
<given-names>B. B.</given-names>
</name>
<name>
<surname>Alf&#xf6;ldi</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>Q.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Variation across 141,456 Human Exomes and Genomes Reveals the Spectrum of Loss-Of-Function Intolerance across Human Protein-Coding Genes</article-title>. <source>bioRxiv</source>. <pub-id pub-id-type="doi">10.1101/531210</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kojima</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kawai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Misawa</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Mimori</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Nagasaki</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>STR-realigner: A Realignment Method for Short Tandem Repeat Regions</article-title>. <source>BMC Genomics</source> <volume>17</volume>. <pub-id pub-id-type="doi">10.1186/s12864-016-3294-x</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Leinonen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sugawara</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Shumway</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>The Sequence Read Archive</article-title>. <source>Nucleic Acids Res.</source> <volume>39</volume>, <fpage>D19</fpage>&#x2013;<lpage>D21</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkq1019</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Lenth</surname>
<given-names>R. V.</given-names>
</name>
</person-group> (<year>2018</year>). <source>Emmeans: Estimated Marginal Means, Aka Least-Squares Means</source>. </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>A Statistical Framework for SNP Calling, Mutation Discovery, Association Mapping and Population Genetical Parameter Estimation from Sequencing Data</article-title>. <source>Bioinformatics</source> <volume>27</volume>, <fpage>2987</fpage>&#x2013;<lpage>2993</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btr509</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Durbin</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Fast and Accurate Short Read Alignment with Burrows-Wheeler Transform</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>1754</fpage>&#x2013;<lpage>1760</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp324</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Handsaker</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wysoker</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Fennell</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Ruan</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Homer</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2009a</year>). <article-title>The Sequence Alignment/Map Format and SAMtools</article-title>. <source>Bioinformatics</source> <volume>25</volume>, <fpage>2078</fpage>&#x2013;<lpage>2079</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btp352</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Fang</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kristiansen</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2009b</year>). <article-title>SNP Detection for Massively Parallel Whole-Genome Resequencing</article-title>. <source>Genome Res.</source> <volume>19</volume>, <fpage>1124</fpage>&#x2013;<lpage>1132</lpage>. <pub-id pub-id-type="doi">10.1101/gr.088013.108</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McCoy</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Beeson</surname>
<given-names>S. K.</given-names>
</name>
<name>
<surname>Splan</surname>
<given-names>R. K.</given-names>
</name>
<name>
<surname>Lykkjen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ralston</surname>
<given-names>S. L.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Identification and Validation of Risk Loci for Osteochondrosis in Standardbreds</article-title>. <source>BMC genomics</source> <volume>17</volume>, <fpage>41</fpage>&#x2013;<lpage>0162385</lpage>. <pub-id pub-id-type="doi">10.1186/s12864-016-2385-z</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McCoy</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Toth</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Dolvik</surname>
<given-names>N. I.</given-names>
</name>
<name>
<surname>Ekman</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ellermann</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Olstad</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Articular Osteochondrosis: A Comparison of Naturally-Occurring Human and Animal Disease</article-title>, <source>Osteoarthritis Cartilage</source> <volume>21</volume>, <fpage>1638</fpage>&#x2013;<lpage>1647</lpage>. <pub-id pub-id-type="doi">10.1016/j.joca.2013.08.011</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McCue</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Valberg</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>M. B.</given-names>
</name>
<name>
<surname>Wade</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>DiMauro</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Akman</surname>
<given-names>H. O.</given-names>
</name>
<etal/>
</person-group> (<year>2008</year>). <article-title>Glycogen Synthase (GYS1) Mutation Causes a Novel Skeletal Muscle Glycogenosis</article-title>. <source>Genomics</source> <volume>91</volume>, <fpage>458</fpage>&#x2013;<lpage>466</lpage>. <pub-id pub-id-type="doi">10.1016/j.ygeno.2008.01.011</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McIlwraith</surname>
<given-names>C. W.</given-names>
</name>
<name>
<surname>Frisbie</surname>
<given-names>D. D.</given-names>
</name>
<name>
<surname>Kawcak</surname>
<given-names>C. E.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>The Horse as a Model of Naturally Occurring Osteoarthritis</article-title>. <source>Bone Jt. Res.</source> <volume>1</volume>, <fpage>297</fpage>&#x2013;<lpage>309</lpage>. <pub-id pub-id-type="doi">10.1302/2046-3758.111.2000132</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>McKenna</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hanna</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Banks</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Sivachenko</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cibulskis</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Kernytsky</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>The Genome Analysis Toolkit: a MapReduce Framework for Analyzing Next-Generation DNA Sequencing Data</article-title>. <source>Genome Res.</source> <volume>20</volume>, <fpage>1297</fpage>&#x2013;<lpage>1303</lpage>. <pub-id pub-id-type="doi">10.1101/gr.107524.110</pub-id> </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mu</surname>
<given-names>X. J.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Z. J.</given-names>
</name>
<name>
<surname>Kong</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Lam</surname>
<given-names>H. Y. K.</given-names>
</name>
<name>
<surname>Gerstein</surname>
<given-names>M. B.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Analysis of Genomic Variation in Non-coding Elements Using Population-Scale Sequencing Data from the 1000 Genomes Project</article-title>. <source>Nucleic Acids Res.</source> <volume>39</volume>, <fpage>7058</fpage>&#x2013;<lpage>7076</lpage>. <pub-id pub-id-type="doi">10.1093/nar/gkr342</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Noll</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Smith</surname>
<given-names>L. D.</given-names>
</name>
<name>
<surname>Yoo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Fiedler</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Cooley</surname>
<given-names>L. D.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Clinical Detection of Deletion Structural Variants in Whole-Genome Sequences</article-title>. <source>Npj&#x20;Genomic Med.</source> <volume>1</volume>. <pub-id pub-id-type="doi">10.1038/npjgenmed.2016.26</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Norton</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Avila</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Schultz</surname>
<given-names>N. E.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Geor</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>McCue</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Evaluation of an HMGA2 Variant for Pleiotropic Effects on Height and Metabolic Traits in Ponies</article-title>. <source>J.&#x20;Vet. Intern. Med.</source> <volume>33</volume>, <fpage>942</fpage>&#x2013;<lpage>952</lpage>. <pub-id pub-id-type="doi">10.1111/jvim.15403</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Norton</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Binns</surname>
<given-names>M. M.</given-names>
</name>
<name>
<surname>Blott</surname>
<given-names>S. C.</given-names>
</name>
<name>
<surname>Caputo</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Isgren</surname>
<given-names>C. M.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>Heritability of Recurrent Exertional Rhabdomyolysis in Standardbred and Thoroughbred Racehorses Derived from SNP Genotyping Data</article-title>. <source>Jhered</source> <volume>107</volume>, <fpage>537</fpage>&#x2013;<lpage>543</lpage>. <pub-id pub-id-type="doi">10.1093/jhered/esw042</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Petersen</surname>
<given-names>J.&#x20;L.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.&#x20;R.</given-names>
</name>
<name>
<surname>Cothran</surname>
<given-names>E. G.</given-names>
</name>
<name>
<surname>Andersson</surname>
<given-names>L. S.</given-names>
</name>
<name>
<surname>Axelsson</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bailey</surname>
<given-names>E.</given-names>
</name>
<etal/>
</person-group> (<year>2013</year>). <article-title>Genetic Diversity in the Modern Horse Illustrated from Genome-wide SNP Data</article-title>. <source>PloS one</source> <volume>8</volume>, <fpage>e54997</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0054997</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Raudsepp</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Finno</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Bellone</surname>
<given-names>R. R.</given-names>
</name>
<name>
<surname>Petersen</surname>
<given-names>J.&#x20;L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Ten Years of the Horse Reference Genome: Insights into Equine Biology, Domestication and Population Dynamics in the post&#x2010;genome Era</article-title>. <source>Anim. Genet.</source> <volume>50</volume>, <fpage>569</fpage>&#x2013;<lpage>597</lpage>. <pub-id pub-id-type="doi">10.1111/age.12857</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rebolledo-Mendez</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Hestand</surname>
<given-names>M. S.</given-names>
</name>
<name>
<surname>Coleman</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Zeng</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Orlando</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>MacLeod</surname>
<given-names>J.&#x20;N.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Comparison of the Equine Reference Sequence with its Sanger Source Data and New Illumina Reads</article-title>. <source>PLoS One</source> <volume>10</volume>, <fpage>e0126852</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0126852</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Rooney</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>E. W.</given-names>
</name>
<name>
<surname>Kelly</surname>
<given-names>V. P.</given-names>
</name>
<name>
<surname>Porter</surname>
<given-names>R. K.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>The &#x201c;Speed Gene&#x201d; Effect of Myostatin Arises in Thoroughbred Horses Due to a Promoter Proximal SINE Insertion</article-title>. <source>PLoS ONE</source> <volume>13</volume>, <fpage>e0205664</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0205664</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schaefer</surname>
<given-names>R. J.</given-names>
</name>
<name>
<surname>Schubert</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Bailey</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bannasch</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Barrey</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Bar-Gal</surname>
<given-names>G. K.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Developing a 670k Genotyping Array to Tag &#x223C;2M SNPs across 24 Horse Breeds</article-title>. <source>BMC Genomics</source> <volume>18</volume>, <fpage>565</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-017-3943-8</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schneider</surname>
<given-names>V. A.</given-names>
</name>
<name>
<surname>Graves-Lindsay</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Howe</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bouk</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>H.-C.</given-names>
</name>
<name>
<surname>Kitts</surname>
<given-names>P. A.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Evaluation of GRCh38 and De Novo Haploid Genome Assemblies Demonstrates the Enduring Quality of the Reference Assembly</article-title>. <source>Genome Res.</source> <volume>27</volume>, <fpage>849</fpage>&#x2013;<lpage>864</lpage>. <pub-id pub-id-type="doi">10.1101/gr.213611.116</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Schultz</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2016</year>). <source>Characterization of Equine Metabolic Syndrome and Mapping of Candidate Genetic Loci</source>. </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Smedley</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Schubach</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jacobsen</surname>
<given-names>J.&#x20;O. B.</given-names>
</name>
<name>
<surname>K&#xf6;hler</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zemojtel</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Spielmann</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2016</year>). <article-title>A Whole-Genome Analysis Framework for Effective Identification of Pathogenic Regulatory Variants in Mendelian Disease</article-title>. <source>Am. J.&#x20;Hum. Genet.</source> <volume>99</volume>, <fpage>595</fpage>&#x2013;<lpage>606</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2016.07.005</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sudmant</surname>
<given-names>P. H.</given-names>
</name>
<name>
<surname>Rausch</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Gardner</surname>
<given-names>E. J.</given-names>
</name>
<name>
<surname>Handsaker</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Abyzov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Huddleston</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>An Integrated Map of Structural Variation in 2,504 Human Genomes</article-title>. <source>Nature</source> <volume>526</volume>, <fpage>75</fpage>&#x2013;<lpage>81</lpage>. <pub-id pub-id-type="doi">10.1038/nature15394</pub-id> </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tenesa</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Navarro</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Hayes</surname>
<given-names>B. J.</given-names>
</name>
<name>
<surname>Duffy</surname>
<given-names>D. L.</given-names>
</name>
<name>
<surname>Clarke</surname>
<given-names>G. M.</given-names>
</name>
<name>
<surname>Goddard</surname>
<given-names>M. E.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Recent Human Effective Population Size Estimated from Linkage Disequilibrium</article-title>. <source>Genome Res.</source> <volume>17</volume>, <fpage>520</fpage>&#x2013;<lpage>526</lpage>. <pub-id pub-id-type="doi">10.1101/gr.6023607</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wade</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Giulotto</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Sigurdsson</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Zoli</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gnerre</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Imsland</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Genome Sequence, Comparative Analysis, and Population Genetics of the Domestic Horse</article-title>. <source>Science</source> <volume>326</volume>, <fpage>865</fpage>&#x2013;<lpage>867</lpage>. <pub-id pub-id-type="doi">10.1126/science.1178158</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Raskin</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Samuels</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Shyr</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Guo</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Genome Measures Used for Quality Control Are Dependent on Gene Function and Ancestry</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>318</fpage>&#x2013;<lpage>323</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu668</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ward</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Valberg</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Adelson</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Abbey</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Binns</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Mickelson</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Glycogen Branching Enzyme (GBE1) Mutation Causing Equine Glycogen Storage Disease IV</article-title>. <source>Mamm. Genome</source> <volume>15</volume>, <fpage>570</fpage>&#x2013;<lpage>577</lpage>. <pub-id pub-id-type="doi">10.1007/s00335-004-2369-1</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Muzny</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Niu</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Person</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>Y.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Molecular Findings Among Patients Referred for Clinical Whole-Exome Sequencing</article-title>. <source>Jama</source> <volume>312</volume>, <fpage>1870</fpage>. <pub-id pub-id-type="doi">10.1001/jama.2014.14601</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yngvadottir</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Xue</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Searle</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hunt</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Delgado</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Morrison</surname>
<given-names>J.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>A Genome-wide Survey of the Prevalence and Evolutionary Forces Acting on Human Nonsense SNPs</article-title>. <source>Am. J.&#x20;Hum. Genet.</source> <volume>84</volume>, <fpage>224</fpage>&#x2013;<lpage>234</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2009.01.008</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>