<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">740340</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.740340</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Optimizing Sequencing Resources in Genotyped Livestock Populations Using Linear Programming</article-title>
<alt-title alt-title-type="left-running-head">Cheng et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Sequencing Choices by Linear Programming</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Cheng</surname>
<given-names>Hao</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/774102/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Xu</surname>
<given-names>Keyu</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Li</surname>
<given-names>Jinghui</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Abraham</surname>
<given-names>Kuruvilla Joseph</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1402623/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Department of Animal Science, University of California, Davis, <addr-line>Davis</addr-line>, <addr-line>CA</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Department of Economics, FEARP, University of S&#xe3;o-Paulo, <addr-line>Ribeir&#xe3;o Preto</addr-line>, <country>Brazil</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Department of Computer Science-ICMC, University of S&#xe3;o Paulo, <addr-line>S&#xe3;o Carlos</addr-line>, <country>Brazil</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/82675/overview">Haja N. Kadarmideen</ext-link>, Synomics Limited, Denmark</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1239457/overview">Roger Ros-Freixedes</ext-link>, Universitat de Lleida, Spain</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/220613/overview">Gregor Gorjanc</ext-link>, University of Edinburgh, United&#x20;Kingdom</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Kuruvilla Joseph Abraham, <email>abraham@fmrp.usp.br</email>; Hao Cheng, <email>qtlcheng@ucdavis.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Livestock Genomics, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>22</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>740340</elocation-id>
<history>
<date date-type="received">
<day>12</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Cheng, Xu, Li and Abraham.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Cheng, Xu, Li and Abraham</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>Low-cost genome-wide single-nucleotide polymorphisms (SNPs) are routinely used in animal breeding programs. Compared to SNP arrays, the use of whole-genome sequence data generated by the next-generation sequencing technologies (NGS) has great potential in livestock populations. However, sequencing a large number of animals to exploit the full potential of whole-genome sequence data is not feasible. Thus, novel strategies are required for the allocation of sequencing resources in genotyped livestock populations such that the entire population can be imputed, maximizing the efficiency of whole genome sequencing budgets. We present two applications of linear programming for the efficient allocation of sequencing resources. The first application is to identify the minimum number of animals for sequencing subject to the criterion that each haplotype in the population is contained in at least one of the animals selected for sequencing. The second application is the selection of animals whose haplotypes include the largest possible proportion of common haplotypes present in the population, assuming a limited sequencing budget. Both applications are available in an open source program LPChoose. In both applications, LPChoose has similar or better performance than some other methods suggesting that linear programming methods offer great potential for the efficient allocation of sequencing resources. The utility of these methods can be increased through the development of improved heuristics.</p>
</abstract>
<kwd-group>
<kwd>dairy cattle</kwd>
<kwd>sequencing</kwd>
<kwd>linear programming</kwd>
<kwd>haplotypes</kwd>
<kwd>selection</kwd>
</kwd-group>
<contract-sponsor id="cn001">Coordena&#xe7;&#xe3;o de Aperfei&#xe7;oamento de Pessoal de N&#xed;vel Superior<named-content content-type="fundref-id">10.13039/501100002322</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>The discovery of genome-wide single-nucleotide polymorphisms (SNPs) and effective ways to assay them has revolutionized genetic analyses of quantitative traits in animal breeding (<xref ref-type="bibr" rid="B21">VanRaden, 2008</xref>; <xref ref-type="bibr" rid="B11">Hayes et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B22">VanRaden et&#x20;al., 2009</xref>; <xref ref-type="bibr" rid="B10">Habier et&#x20;al., 2010</xref>; <xref ref-type="bibr" rid="B24">Wolc et&#x20;al., 2012</xref>). In conventional breeding programs, low-cost SNP array data are routinely used in genomic selection to estimate breeding values. Compared to SNP arrays, the use of whole-genome sequence data generated by next-generation sequencing technologies (NGS) has great potential in livestock populations for causal mutation detection (<xref ref-type="bibr" rid="B5">Daetwyler et&#x20;al., 2014</xref>) in genome-wide association studies and more stable or accurate prediction of breeding values in genomic selection (<xref ref-type="bibr" rid="B15">Meuwissen and Goddard, 2010</xref>).</p>
<p>As sequencing is expensive, there is considerable interest in extracting as much information as possible by sequencing a limited number of animals and then imputing to other animals. Some strategies for allocating limited sequencing resources in genotyped livestock populations require that certain key individuals be sequenced at high coverage (<xref ref-type="bibr" rid="B7">Druet et&#x20;al., 2014</xref>) while others consider sequencing a large number of animals at lower coverage (<xref ref-type="bibr" rid="B20">VanRaden et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B13">Li et&#x20;al., 2011</xref>). Sequencing a large number of animals at low coverage offers certain advantages; in <xref ref-type="bibr" rid="B20">VanRaden et&#x20;al. (2015)</xref>, the authors comment on the improvement in genotype imputation attained through increasing the number of animals sequenced at low coverage, and in <xref ref-type="bibr" rid="B13">Li et&#x20;al. (2011)</xref>, the authors point out the benefits of sequencing many individuals at low coverage in complex trait association studies. The methods we present in this paper are better suited for selecting animals to be sequenced when selecting a large number of animals at low coverage.</p>
<p>The choice of which animals to sequence has been studied from the perspective of maximizing the efficiency of whole genome sequencing experiments. To this end, in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>, the authors discuss the problem of finding a subset of animals of limited size which covers all of the haplotypes present in a given population with a frequency above a predefined threshold. The selection of animals for sequencing based on haplotypes and their frequencies has also been extensively discussed in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref>. The methods we will present for the selection of animals assume that a low-density haplotype library has been constructed, and that low-density haplotype information on all animals is available. We will not make use of pedigree information (<xref ref-type="bibr" rid="B18">Ros-Freixedes et&#x20;al., 2020</xref>), or discuss overall imputation accuracy (<xref ref-type="bibr" rid="B9">Gonen et&#x20;al., 2017</xref>; <xref ref-type="bibr" rid="B17">Ros-Freixedes et&#x20;al., 2017</xref>). We will however assume that in order to recover through imputation a given haplotype in any animal, at least one animal containing that haplotype must be sequenced. This requirement imposes a number of constraints to be simultaneously satisfied.</p>
<p>Selecting the best set of animals in a population capable of satisfying certain constraints is an optimization problem in a space whose dimension is equal to the total number of animals in the population. In practical applications the optimization must be performed for tens of thousands of animals subject to tens of thousands of constraints. Despite the large dimension of the search space, and the large number of constraints, the optimization problem we are trying to solve here can be addressed through the use of a well-established set of techniques known as linear programming [for a review see (<xref ref-type="bibr" rid="B14">Luenberger and Ye, 2015</xref>)]. Linear programming has been previously applied in animal and plant breeding to optimize breeding decisions (<xref ref-type="bibr" rid="B6">Diaz et&#x20;al., 1999</xref>; <xref ref-type="bibr" rid="B16">Moeinizade et&#x20;al., 2019</xref>), but not to the allocation of sequencing resources. In this paper, we will study the application of linear programming to the allocation of sequencing resources to address two different questions which may arise in breeding programs.</p>
<p>The first question is to determine the minimum number of animals needed to permit sequence imputation into all other members of the population without loss of haplotype diversity, and then to identify these animals. If a subset of animals which contains all the haplotypes in the population can be identified, then the animals in this subset can be considered to be a starting point for sequencing to eventually permit imputation into the entire population. In order to minimize overall costs, we will attempt to find the smallest subset of animals with this property.</p>
<p>In practice, however, even after the smallest subset of such animals has been identified, it may not be possible to sequence all the animals identified in this manner due to budget constraints. It may then become necessary to choose a limited subset of animals carrying haplotypes which may be considered to be representative of the population on the basis of being more frequent. Identifying this limited subset of animals is the second question which will be addressed in this paper. The objectives of this paper are to study the use of linear programming models to address all of the questions described above, and to compare the performance of linear programming methods with previously published methods.</p>
<p>We will describe in detail the application of our linear programming based method, LPChoose, and also compare the numerical results obtained from LPChoose to several approaches including IWS (<xref ref-type="bibr" rid="B2">Bickhart et&#x20;al., 2016</xref>; <xref ref-type="bibr" rid="B4">Butty et&#x20;al., 2019</xref>), AHAP (<xref ref-type="bibr" rid="B7">Druet et&#x20;al., 2014</xref>; <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al., 2016</xref>) and AlphaSeqOpt (<xref ref-type="bibr" rid="B9">Gonen et&#x20;al., 2017</xref>). This numerical comparison between the results obtained by different methods is completely based on observed haplotype frequencies, we do not make use of ancestral haplotype frequencies or haplotypes frequencies from databases in our selection criteria for animals to sequence. We will also explain how some key criteria of the HSH and GDI schemes in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref> and the IWS scheme in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> for the selection of individuals can be understood in the language of linear programming. Two important questions related to imputation, haplotype phasing and resolution, which are addressed in <xref ref-type="bibr" rid="B9">Gonen et&#x20;al. (2017)</xref>, and in <xref ref-type="bibr" rid="B17">Ros-Freixedes et&#x20;al. (2017)</xref>, will not be discussed&#x20;here.</p>
</sec>
<sec id="s2">
<title>2 Methods</title>
<p>We now establish some notation which will be used throughout this paper. For each animal we associate an indicator variable <italic>x</italic>
<sub>
<italic>j</italic>
</sub> such that <italic>x</italic>
<sub>
<italic>j</italic>
</sub> &#x2208; (0, 1) and <italic>j</italic>&#x20;&#x2208; (1, 2, &#x2026;, <italic>n</italic>) where <italic>n</italic> is the total number of animals. <italic>x</italic>
<sub>
<italic>j</italic>
</sub> &#x3d; 1(0) means that the animal <italic>j</italic> is selected (not selected) for sequencing. We next introduce binary constants <italic>a</italic>
<sub>
<italic>ij</italic>
</sub> where <italic>i</italic>&#x20;&#x2208; (1, 2, &#x2026;, <italic>p</italic>) and <italic>p</italic> is the total number of unique haplotypes. <italic>a</italic>
<sub>
<italic>ij</italic>
</sub> &#x3d; 1(0) if animal <italic>j</italic> carries haplotype <italic>i</italic> (or not). Since we are solving systems of inequalities in which the variables are binary valued, we actually work within the framework of integer linear programming, which is more restricted than linear programming.</p>
<p>The first application of integer linear programming we consider is to identify the minimum number of animals for sequencing while meeting the criteria that each haplotype is contained in at least one of the animals selected for sequencing. If haplotype <italic>i</italic> is carried in at least one animal in the population, then we require that <inline-formula id="inf1">
<mml:math id="m1">
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>. This condition can be satisfied only if <italic>x</italic>
<sub>
<italic>j</italic>
</sub> &#x3d; 1 for at least one value of 1 &#x2264; <italic>j</italic>&#x20;&#x2264; <italic>n</italic>. A similar constraint must hold separately for all <italic>i</italic>&#x20;&#x2208; 1, 2, &#x2026;,&#x20;<italic>p</italic>.</p>
<p>In order to minimize the number of animals to be sequenced, we additionally require <inline-formula id="inf2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, where <italic>z</italic>
<sub>1</sub> is the number of selected animals, to be as small as possible. These inequalities can be collectively written as.<disp-formula id="e1">
<mml:math id="m3">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>minimize</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mn>0,1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m4">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>subject&#x2009;to&#x2009;the&#x2009;constraints</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2,3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mspace width="0.17em"/>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(2)</label>
</disp-formula>
<xref ref-type="disp-formula" rid="e1">Equation 1</xref> is the objective function to minimize the total number of selected animals. <xref ref-type="disp-formula" rid="e2">Equation 2</xref> is the set of constraints which ensures that each haplotype is present in at least one of the animals selected for sequencing.</p>
<p>The objective function and the constraints can also be written in matrix notation as follows.<disp-formula id="e3">
<mml:math id="m5">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>minimize</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mo>,</mml:mo>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(3)</label>
</disp-formula>
<disp-formula id="e4">
<mml:math id="m6">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>subject&#x2009;to</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mi mathvariant="bold">A</mml:mi>
<mml:mi mathvariant="bold-italic">x</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(4)</label>
</disp-formula>where <bold>A</bold> is the matrix of values <italic>a</italic>
<sub>
<italic>ij</italic>
</sub> and <bold>
<italic>x</italic>
</bold> denotes the vector of <italic>x</italic>
<sub>
<italic>i</italic>
</sub>&#x20;values. This system of inequalities can be solved by integer&#x20;linear programming. The solution for <italic>x</italic> will contain&#x20;some values equal to zero and others equal to one. The <italic>x</italic>
<sub>
<italic>j</italic>
</sub> which are equal to 1 correspond to the animals which should be sequenced.</p>
<p>The second question, to select a fixed number of animals with most common haplotypes, can also be addressed using integer linear programming. In order to prioritize animals with more frequent haplotypes we define a vector with elements <italic>h</italic>
<sub>
<italic>i</italic>
</sub> with 1 &#x2264; <italic>i</italic>&#x20;&#x2264; <italic>p</italic> whose <italic>i</italic>th element is the frequency of the <italic>i</italic>th haplotype. The vector <bold>h</bold> is then used to define another vector <bold>c</bold> &#x3d; <bold>h</bold>
<sup>
<italic>t</italic>
</sup>
<bold>A</bold>. The number of elements of <bold>c</bold> is equal to the number of animals. Values of <bold>c</bold> are larger for animals which carry many more frequent haplotypes. Thus the values of <bold>c</bold> will be used as a guide to select animals when sequencing resources are limited. In addition, it is important to ensure that the same haplotype is not sequenced in a large number of different animals so as to maximize the haplotype diversity in the animals selected for sequencing. All these different requirements can be summarized in the following set of inequalities.<disp-formula id="e5">
<mml:math id="m7">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>maximize</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>c</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mspace width="1em"/>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2208;</mml:mo>
<mml:mn>0,1</mml:mn>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(5)</label>
</disp-formula>
<disp-formula id="e6">
<mml:math id="m8">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mtext>subject&#x2009;to</mml:mtext>
<mml:mspace width="1em"/>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(6)</label>
</disp-formula>
<disp-formula id="e7">
<mml:math id="m9">
<mml:mtable class="eqnarray">
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
<mml:mi>a</mml:mi>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mspace width="1em"/>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1,2,3</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#x22ef;</mml:mo>
<mml:mspace width="0.17em"/>
<mml:mo>,</mml:mo>
<mml:mi>p</mml:mi>
</mml:mtd>
<mml:mtd columnalign="left"/>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(7)</label>
</disp-formula>
<xref ref-type="disp-formula" rid="e5">Equation 5</xref> is the objective function to maximize the number of more frequent haplotypes represented by the selected animals. <xref ref-type="disp-formula" rid="e6">Equation 6</xref> is the constraint to set the maximum number of animals to be sequenced to be equal to <italic>n</italic>
<sub>
<italic>max</italic>
</sub>. <xref ref-type="disp-formula" rid="e7">Equation 7</xref> is the constraint to ensure that each haplotype covered is at the most covered by <italic>r</italic>
<sub>
<italic>max</italic>
</sub> animals, where <italic>r</italic>
<sub>
<italic>max</italic>
</sub> should be a positive integer &#x2265; 1 and is chosen to be considerably smaller than <italic>n</italic>
<sub>
<italic>max</italic>
</sub>. Ideally, the objective function should be maximized with <italic>r</italic>
<sub>
<italic>max</italic>
</sub> chosen to be small and the maximization performed in one step. This would permit the optimal selection of multiple animals without introducing any approximations. In our data analysis, for reasons which will soon be discussed, we will introduce a number of additional approximations in the solution to <xref ref-type="disp-formula" rid="e5">Eqs 5</xref>&#x2013;<xref ref-type="disp-formula" rid="e7">7</xref>.</p>
<p>The form of <xref ref-type="disp-formula" rid="e5">Eqs 5</xref>&#x2013;<xref ref-type="disp-formula" rid="e7">7</xref> are generic in a linear programming context, and in a broader context the coefficients <italic>c</italic>
<sub>
<italic>j</italic>
</sub> in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> may be freely chosen, as long as the coefficients are positive. With the coefficients chosen as described in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref>, it is possible to prioritize the selection of animals in the IWS scheme (<xref ref-type="bibr" rid="B2">Bickhart et&#x20;al., 2016</xref>). The highly segregating haplotype (HSH) scheme mentioned in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref> uses coefficients similar to the ones we suggest, but with some additional multiplicative factors which are introduced to avoid including the same haplotype in multiple animals. In our framework, this is achieved by the choice of <italic>r</italic>
<sub>
<italic>max</italic>
</sub>. One important issue addressed in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>, how to select the smallest number of samples needed to sequence all haplotypes above a given frequency threshold, can be addressed within the framework of linear programming by solving <xref ref-type="disp-formula" rid="e3">Eqs 3</xref>, <xref ref-type="disp-formula" rid="e4">4</xref> after discarding the rows corresponding to the low frequency haplotypes. If no rows, or very few rows are discarded in <xref ref-type="disp-formula" rid="e3">Eqs 3</xref>, <xref ref-type="disp-formula" rid="e4">4</xref>, then animals containing rare haplotypes can be targeted, which is one of the objectives of the optimized GDI method discussed in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref>. Thus many of the main key ideas for the prioritization of animals in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> and <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref> can be incorporated into a Linear Programming framework.</p>
<p>Linear programming problems described above cannot be solved analytically, however branch and bound methods (<xref ref-type="bibr" rid="B12">Land and Doig, 1960</xref>) are guaranteed to converge to the global optimum. In practice, for reasons which will be discussed later, the convergence can be very slow without the use of approximations. It turns out that for the first application <xref ref-type="disp-formula" rid="e1">Eqs 1</xref>, <xref ref-type="disp-formula" rid="e2">2</xref> can be solved rapidly and exactly. No approximations are needed and the identification of which animals to sequence with a view to include all haplotypes is done in LPChoose in a single step. In <xref ref-type="bibr" rid="B9">Gonen et&#x20;al. (2017)</xref> the authors also comment on the possibility of determining which animals are needed to cover all haplotypes, but through the solution of a system of equations in multiple stages.</p>
<p>In the case of the second application, it turns out that convergence in LPChoose is very slow if exact solutions are sought. In order to facilitate convergence in the second application we will set the value of <italic>r</italic>
<sub>
<italic>max</italic>
</sub> to 2 and instead of selecting all <italic>n</italic>
<sub>
<italic>max</italic>
</sub> animals at once, we will first select 2 animals (i.e.,&#x20;<italic>n</italic>
<sub>
<italic>max</italic>
</sub> &#x3d; 2). Once the first two animals have been identified, these animals and all the haplotypes present in them are removed. Then the optimization in <xref ref-type="disp-formula" rid="e5">Eqs 5</xref>&#x2013;<xref ref-type="disp-formula" rid="e7">7</xref> is repeated but with the objective function in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> suitably modified to reflect the absence of the first two animals and also the haplotypes that they carry. This procedure can be carried out until the desired number of animals has been found. This approach amounts to breaking up the larger optimization, which cannot be solved exactly in reasonable time, into a smaller number of problems each of which can be solved exactly and quickly by linear programming, and then combining the results.</p>
<p>In <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>, both the AHAP selection schemes, AHAP1 and AHAP2, use the same weights as LPChoose, those in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref>, but consider only homozygous haplotypes. Furthermore, in the AHAP1 scheme, there is no updating of weights to take into account animals already selected. We will present results with no updating of weights, as in AHAP1, but using both homozygotes and heterozygotes. Our results for AHAP1 will use <italic>n</italic>
<sub>
<italic>max</italic>
</sub> &#x3d; 1 repeatedly until the desired number of animals is obtained. Even though the AHAP2 scheme in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> uses the same weights and updating as LPChoose, there is an important difference between LPChoose and AHAP2 as implemented in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>. LPChoose is able to select multiple animals at a time in contrast to the AHAP2 scheme in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>.</p>
<p>In the results we present for IWS, we will prioritize the animals using the IWS weighting scheme in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> and <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref>, and select a single animal at a time, in order to facilitate a comparison with the IWS implementation in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>. Multiple rounds of selection are performed until the desired number of animals is obtained.</p>
<p>In order to compare the different methods described, genotype data for five different scenarios were simulated following the general scheme for simulating artificial populations in <xref ref-type="bibr" rid="B9">Gonen et&#x20;al. (2017)</xref>. The number of generations considered in these 5 scenarios are 5, 10, 15, 30, or 50. These scenarios resemble modern cattle populations. The genome of 10,000 segregating loci on 10 chromosomes was simulated using the &#x201c;cattle genome&#x201d; option in AlphaSimR (<xref ref-type="bibr" rid="B8">Gaynor et&#x20;al., 2020</xref>). A quantitative trait controlled by 150 QTL of effects sampled from standard normal distributions distributed equally on 10 chromosomes was simulated. First, founders of 1,000 cattle of equal sex ratio were generated. At each generation, the best 25 males were selected as sires on the basis of their highest breeding value and mated to all 500 females as dams to produce next generations with 1,000 cattle of equal sex ratio. In our analysis, 10 replicates were simulated for each of the five scenarios. From these simulations, all individuals had haplotypes for 10,000 SNPs distributed equally across the 10 chromosomes. In each population, haplotype blocks of length 100 SNPs were obtained across 10 chromosomes. A mismatch of up to 10% was used to ensure that haplotypes with small differences were considered as identical.</p>
<p>In the results, the selection of a fixed number animals will also be made after the exclusion of haplotypes with a frequency of less than 1%, similar to the strategy adopted in <xref ref-type="bibr" rid="B19">Su et&#x20;al. (2014)</xref>. For the rest of the manuscript, haplotypes which are retained after this exclusion will be referred to as common haplotypes. The results in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> are also based on the exclusion of low frequency haplotypes but with a more drastic restriction on haplotype frequency than in <xref ref-type="bibr" rid="B19">Su et&#x20;al. (2014)</xref>. A brief description of the simulation scenarios is presented in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. We will also briefly discuss results in which the number of animals in <xref ref-type="table" rid="T1">Table&#x20;1</xref> remains the same but the number of haplotypes is considerably increased. An open-source, publicly available Julia package called LPChoose (<ext-link ext-link-type="uri" xlink:href="https://github.com/reworkhow/LPChoose.jl">https://github.com/reworkhow/LPChoose.jl</ext-link>) which makes use of the GNU Linear Programming Kit (GLPK) has been developed. For the sake of the reproducibility of our results, we make no use whatsoever of any proprietary solvers for linear programming.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Simulation scenarios.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Scenario</th>
<th align="center">Total number of animals</th>
<th align="center">Number of unique haplotypes</th>
<th align="center">Number of unique common haplotypes</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="center">6,000</td>
<td align="center">27,399&#x20;&#xb1; 332</td>
<td align="center">2,638&#x20;&#xb1; 42</td>
</tr>
<tr>
<td align="left">2</td>
<td align="center">11,000</td>
<td align="center">37,922&#x20;&#xb1; 758</td>
<td align="center">2,376&#x20;&#xb1; 50</td>
</tr>
<tr>
<td align="left">3</td>
<td align="center">16,000</td>
<td align="center">44,591&#x20;&#xb1; 826</td>
<td align="center">2,216&#x20;&#xb1; 49</td>
</tr>
<tr>
<td align="left">4</td>
<td align="center">31,000</td>
<td align="center">56,036&#x20;&#xb1; 1,313</td>
<td align="center">1,959&#x20;&#xb1; 54</td>
</tr>
<tr>
<td align="left">5</td>
<td align="center">51,000</td>
<td align="center">61,418&#x20;&#xb1; 2,023</td>
<td align="center">1,815&#x20;&#xb1; 44</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<title>3 Results</title>
<p>The minimum numbers of animals identified in LPChoose, AlphaSeqOpt, IWS, and AHAP to cover all unique haploptypes in five populations are shown in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. The results for AlphaSeqOpt in the third column of <xref ref-type="table" rid="T2">Table&#x20;2</xref> were obtained from repeated runs of AlphaSeqOpt with varying numbers of focal animals as defined in <xref ref-type="bibr" rid="B9">Gonen et&#x20;al. (2017)</xref> to be selected. In the case of IWS and AHAP, animals were selected one at a time until all haplotypes were covered. For these methods, the selection of successive animals was done using the weights discussed earlier. The minimum numbers of animals identified by LPChoose were consistently smaller than those from AlphaSeqOpt, IWS, and AHAP. The results of <xref ref-type="table" rid="T2">Table&#x20;2</xref> suggest that the difference between LPChoose and the other methods in the smallest number of animals increases as the size of the population increases.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Minimum number of animals identified representing all haplotypes in the population.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Scenario</th>
<th colspan="4" align="center">Minimum number of animals (mean &#xb1; SD)</th>
</tr>
<tr>
<th align="center">LPChoose</th>
<th align="center">AlphaSeqOpt</th>
<th align="center">IWS</th>
<th align="center">AHAP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="center">4,086&#x20;&#xb1; 13</td>
<td align="center">4,988&#x20;&#xb1; 30</td>
<td align="center">4,154&#x20;&#xb1; 30</td>
<td align="center">5,998&#x20;&#xb1; 2</td>
</tr>
<tr>
<td align="left">2</td>
<td align="center">6,663&#x20;&#xb1; 18</td>
<td align="center">8,563&#x20;&#xb1; 51</td>
<td align="center">6,802&#x20;&#xb1; 84</td>
<td align="center">10,999&#x20;&#xb1; 2</td>
</tr>
<tr>
<td align="left">3</td>
<td align="center">8,648&#x20;&#xb1; 34</td>
<td align="center">11,563&#x20;&#xb1; 132</td>
<td align="center">8,821&#x20;&#xb1; 75</td>
<td align="center">15,998&#x20;&#xb1; 3</td>
</tr>
<tr>
<td align="left">4</td>
<td align="center">12,229&#x20;&#xb1; 86</td>
<td align="center">17,506&#x20;&#xb1; 275</td>
<td align="center">12,738&#x20;&#xb1; 225</td>
<td align="center">30,997&#x20;&#xb1; 6</td>
</tr>
<tr>
<td align="left">5</td>
<td align="center">14,241&#x20;&#xb1; 119</td>
<td align="center">21,098&#x20;&#xb1; 494</td>
<td align="center">15,028&#x20;&#xb1; 356</td>
<td align="center">50,999&#x20;&#xb1; 1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We emphasize that the results in column 2 of <xref ref-type="table" rid="T2">Table&#x20;2</xref> obtained from LPChoose are obtained without any approximations, and are thus equal to the theoretically lowest values attainable for this kind of problem. Furthermore, when LPChoose is used to select the minimum number of animals to cover all haplotypes, some haplotypes may be covered by multiple animals. This introduces a certain level of redundancy in sequencing which may be desirable when certain haplotypes should be covered more frequently for facilitating imputation (<xref ref-type="bibr" rid="B17">Ros-Freixedes et&#x20;al., 2017</xref>).</p>
<p>LPChoose, AlphaSeqOpt, IWS, and AHAP were also used to select a fixed number of animals (100 animals) whose haplotypes represent the maximum proportions of the total haplotypes in the population. The results obtained are shown in <xref ref-type="table" rid="T3">Tables 3</xref>, <xref ref-type="table" rid="T4">4</xref>. In <xref ref-type="table" rid="T3">Table&#x20;3</xref>, the proportion of all haplotypes (with no lower limit on haplotype frequencies) represented by 100 selected animals were compared. In this scenario IWS performed best in terms of haplotype coverage, followed by LPChoose. However, independent of the method only a small proportion (0.1&#x2013;0.3) of all haplotypes were covered.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Proportion of all haplotypes represented by 100 selected animals in the population.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Scenario</th>
<th colspan="4" align="center">Proportion of all haplotypes (mean &#xb1; SD)</th>
</tr>
<tr>
<th align="center">LPChoose</th>
<th align="center">AlphaSeqOpt</th>
<th align="center">IWS</th>
<th align="center">AHAP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="char" char="plusmn">0.2218&#x20;&#xb1; 0.0036</td>
<td align="char" char="plusmn">0.2052&#x20;&#xb1; 0.0048</td>
<td align="char" char="plusmn">0.2504&#x20;&#xb1; 0.0032</td>
<td align="char" char="plusmn">0.1441&#x20;&#xb1; 0.0041</td>
</tr>
<tr>
<td align="left">2</td>
<td align="char" char="plusmn">0.1700&#x20;&#xb1; 0.0020</td>
<td align="char" char="plusmn">0.1528&#x20;&#xb1; 0.0048</td>
<td align="char" char="plusmn">0.1951&#x20;&#xb1; 0.0024</td>
<td align="char" char="plusmn">0.1007&#x20;&#xb1; 0.0024</td>
</tr>
<tr>
<td align="left">3</td>
<td align="char" char="plusmn">0.1498&#x20;&#xb1; 0.0026</td>
<td align="char" char="plusmn">0.1319&#x20;&#xb1; 0.0047</td>
<td align="char" char="plusmn">0.1735&#x20;&#xb1; 0.0009</td>
<td align="char" char="plusmn">0.0805&#x20;&#xb1; 0.0029</td>
</tr>
<tr>
<td align="left">4</td>
<td align="char" char="plusmn">0.1249&#x20;&#xb1; 0.0032</td>
<td align="char" char="plusmn">0.1119&#x20;&#xb1; 0.0037</td>
<td align="char" char="plusmn">0.1464&#x20;&#xb1; 0.0022</td>
<td align="char" char="plusmn">0.0584&#x20;&#xb1; 0.0024</td>
</tr>
<tr>
<td align="left">5</td>
<td align="char" char="plusmn">0.1145&#x20;&#xb1; 0.0027</td>
<td align="char" char="plusmn">0.1084&#x20;&#xb1; 0.0036</td>
<td align="char" char="plusmn">0.1370&#x20;&#xb1; 0.0023</td>
<td align="char" char="plusmn">0.0435&#x20;&#xb1; 0.0025</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Proportion of common haplotypes represented by 100 selected animals in the population.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Scenario</th>
<th colspan="4" align="center">Proportion of common haplotypes (mean &#xb1; SD)</th>
</tr>
<tr>
<th align="center">LPChoose</th>
<th align="center">AlphaSeqOpt</th>
<th align="center">IWS</th>
<th align="center">AHAP</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">1</td>
<td align="char" char="plusmn">0.9934&#x20;&#xb1; 0.0017</td>
<td align="char" char="plusmn">0.9239&#x20;&#xb1; 0.0032</td>
<td align="char" char="plusmn">0.9917&#x20;&#xb1; 0.0017</td>
<td align="char" char="plusmn">0.8520&#x20;&#xb1; 0.0159</td>
</tr>
<tr>
<td align="left">2</td>
<td align="char" char="plusmn">0.9911&#x20;&#xb1; 0.0028</td>
<td align="char" char="plusmn">0.9041&#x20;&#xb1; 0.0071</td>
<td align="char" char="plusmn">0.9899&#x20;&#xb1; 0.0017</td>
<td align="char" char="plusmn">0.8288&#x20;&#xb1; 0.0142</td>
</tr>
<tr>
<td align="left">3</td>
<td align="char" char="plusmn">0.9921&#x20;&#xb1; 0.0018</td>
<td align="char" char="plusmn">0.8766&#x20;&#xb1; 0.0171</td>
<td align="char" char="plusmn">0.9875&#x20;&#xb1; 0.0032</td>
<td align="char" char="plusmn">0.7625&#x20;&#xb1; 0.0257</td>
</tr>
<tr>
<td align="left">4</td>
<td align="char" char="plusmn">0.9899&#x20;&#xb1; 0.0021</td>
<td align="char" char="plusmn">0.8444&#x20;&#xb1; 0.0109</td>
<td align="char" char="plusmn">0.9780&#x20;&#xb1; 0.0030</td>
<td align="char" char="plusmn">0.7004&#x20;&#xb1; 0.0247</td>
</tr>
<tr>
<td align="left">5</td>
<td align="char" char="plusmn">0.9993&#x20;&#xb1; 0.0006</td>
<td align="char" char="plusmn">0.8335&#x20;&#xb1; 0.0088</td>
<td align="char" char="plusmn">0.9600&#x20;&#xb1; 0.0061</td>
<td align="char" char="plusmn">0.6287&#x20;&#xb1; 0.0336</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In <xref ref-type="table" rid="T4">Table&#x20;4</xref>, the proportion of common haplotypes represented by 100 selected animals were compared. This scenario is similar to that discussed in <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>. LPChoose performed better than the other three methods consistently across different populations, and the proportions of common haplotypes identified in LPChoose were usually higher than 0.99, followed by IWS with slightly lower proportions.</p>
<p>The better results obtained by using the IWS weighting scheme in <xref ref-type="table" rid="T3">Table&#x20;3</xref> are to be expected since IWS preferentially selects animals with low-frequency haplotypes by assigning larger weights to individuals containing rare haplotypes (<xref ref-type="bibr" rid="B2">Bickhart et&#x20;al., 2016</xref>), unlike LPChoose where more common haplotypes are preferred. The prioritization of animals in IWS, however, can be easily accommodated in a Linear programming framework by changing the weights in <xref ref-type="disp-formula" rid="e5">Eq.&#x20;5</xref>.</p>
<p>All the results discussed so far were obtained in 10&#xa0;min or less of execution time. If the number of haplotypes is increased five fold compared with our original simulations keeping the number of animals unchanged, there is a five fold increase in running times for LPChoose in application 1 and a ten fold increase for application 2. Our earlier conclusions about the relative merits of the different methods considered remain unchanged; for the analyses in <xref ref-type="table" rid="T2">Table&#x20;2</xref> and in <xref ref-type="table" rid="T4">Table&#x20;4</xref>, LPChoose still performs&#x20;best.</p>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>In this paper, linear programming methods are used to answer two questions for the allocation of sequencing resources, to identify the minimum number of animals whose haplotypes represent all the haplotypes in the population, and to choose for sequencing a fixed number of animals whose haplotypes represent the maximum proportion of the common haplotypes in the population. The results from <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T4">4</xref> suggest that linear programming methods, which permit the selection of more than one animal at a time, may still offer a number of advantages in comparison with other methods which rely on the selection of a single animal at a time. It is noteworthy that these improvements were obtained using relatively straightforward approximations, and a publicly available implementation of linear programming.</p>
<p>We make no explicit use of pedigree information, which may be important in deciding which animals to sequence (<xref ref-type="bibr" rid="B3">Boichard, 2002</xref>; <xref ref-type="bibr" rid="B18">Ros-Freixedes et&#x20;al., 2020</xref>). If the animals to be sequenced should also be selected based on positions in the pedigree, then additional requirements based on pedigree structure can be incorporated in a Linear Programming framework.</p>
<p>The expected accuracy of genotype imputation has also been used as a criterion to determine the best animals to be sequenced (<xref ref-type="bibr" rid="B18">Ros-Freixedes et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B25">Yu et&#x20;al., 2014</xref>). Genotype imputation accuracy is expected to improve when a larger number of animals are sequenced at lower coverage as opposed to a smaller number of animals at high coverage (keeping the product of read depth and number of sequenced animals fixed) (<xref ref-type="bibr" rid="B20">VanRaden et&#x20;al., 2015</xref>). Our results suggest that for a fixed number of animals, LPChoose covers more haplotypes than competing methods, suggesting that selecting animals through Linear Programming could permit a higher level of genotype imputation accuracy.</p>
<p>In <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref>, the authors point out that the accuracy of imputation of rare variants is affected by the lower limit put on haplotype frequencies and suggest the importance of including additional animals to improve the accuracy of imputation of rare variants. In <xref ref-type="bibr" rid="B1">Bhati et&#x20;al. (2020)</xref> as well, the authors comment on how the inclusion of rare variants is affected by the choice of animals sequenced. The linear programming methods we present here can be adapted to the selection of multiple animals with low frequency haplotypes by modifying the right hand sides of <xref ref-type="disp-formula" rid="e6">Eqs 6</xref>, <xref ref-type="disp-formula" rid="e7">7</xref> to increase the values of <italic>r</italic>
<sub>
<italic>max</italic>
</sub> in <xref ref-type="disp-formula" rid="e7">Eq. 7</xref> for those rows corresponding to rare haplotypes. The selection of multiple animals with low frequency haplotypes can also be achieved in the framework of Method 1, by increasing the right hand side of <xref ref-type="disp-formula" rid="e2">Eq. 2</xref> for those rows corresponding to low frequency haplotypes.</p>
<p>Furthermore, if additional animals with previously known sequence information are to be included in the set of animals to be sequenced, these animals can be included through additional constraints in <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>. Thereafter, any additional selection can be carried out while maximizing the information present in the animals included by requirement.</p>
<p>There are some unavoidable limitations associated with the linear programming methods which necessitate the use of approximations discussed earlier. To understand the necessity of approximations it is useful to use graph theory to rephrase the first problem we address as that of finding a minimum set cover on a bipartite graph <italic>via</italic> integer linear programming (<xref ref-type="bibr" rid="B23">Vazirani, 2003</xref>). The problem of finding a suitable subset of animals with common haplotypes, can also be rephrased as a set cover problem along with additional weighting factors. As there is no known polynomial time algorithm for solving the set cover problem approximations are unavoidable in large data sets. Greedy approximations which select a single element (or animal in the context of sequencing) at a time are relatively straightforward to implement, but frequently do not lead to the true optimum (<xref ref-type="bibr" rid="B23">Vazirani, 2003</xref>). Hence methods which rely on the selection of more than one animal at a time, such as those we have presented here, may lead to improved results.</p>
<p>To conclude, we have illustrated the use of linear programming for optimizing the allocation of sequencing resources. Our results suggest how linear programming can be used to extend and improve the approximations used in <xref ref-type="bibr" rid="B4">Butty et&#x20;al. (2019)</xref> and <xref ref-type="bibr" rid="B2">Bickhart et&#x20;al. (2016)</xref> to address the important questions discussed in these papers. It is encouraging that the superior results found by LPChoose in <xref ref-type="table" rid="T2">Tables 2</xref>, <xref ref-type="table" rid="T4">4</xref> were obtained without using proprietary linear programming solvers, and applying straightforward heuristics when necessary. The use of more sophisticated heuristics in conjunction with proprietary solvers could lead to additional improvements in both efficiency and lowered running&#x20;time.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/Supplementary Material, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>KJA and HC conceived the study. KJA, HC, KX, and JL contributed to the development of the methodology. KJA, HC, KX, and JL developed the LPChoose package. KX and JL performed the simulation and analysis. HC and KJA wrote the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>This study was financed in part by the Coordena&#xe7;&#xe3;o de Aperfei&#xe7;oamento de Pessoal de N&#xed;vel Superior Brasil (CAPES)&#x20;- Finance Code&#x20;001.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bhati</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kadri</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Crysnanto</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Pausch</surname>
<given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Assessing Genomic Diversity and Signatures of Selection in Original Braunvieh Cattle Using Whole-Genome Sequencing Data</article-title>. <source>BMC Genomics</source> <volume>21</volume>, <fpage>27</fpage>. <pub-id pub-id-type="doi">10.1186/s12864-020-6446-y</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bickhart</surname>
<given-names>D. M.</given-names>
</name>
<name>
<surname>Hutchison</surname>
<given-names>J.&#x20;L.</given-names>
</name>
<name>
<surname>Null</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>VanRaden</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Cole</surname>
<given-names>J.&#x20;B.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Reducing Animal Sequencing Redundancy by Preferentially Selecting Animals with Low-Frequency Haplotypes</article-title>. <source>J.&#x20;Dairy Sci.</source> <volume>99</volume>, <fpage>5526</fpage>&#x2013;<lpage>5534</lpage>. <pub-id pub-id-type="doi">10.3168/jds.2015-10347</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Boichard</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2002</year>). &#x201c;<article-title>Pedig: A Fortran Package for Pedigree Analysis Suited for Large Populations</article-title>,&#x201d; in <conf-name>Presented at 7th World Congress on Genetics Applied to Livestock Production</conf-name>, <conf-loc>Montpellier</conf-loc>, <conf-date>August 2002</conf-date>. </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Butty</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Sargolzaei</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Miglior</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Stothard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Schenkel</surname>
<given-names>F. S.</given-names>
</name>
<name>
<surname>Gredler-Grandl</surname>
<given-names>B.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Optimizing Selection of the Reference Population for Genotype Imputation from Array to Sequence Variants</article-title>. <source>Front. Genet.</source> <volume>10</volume>, <fpage>510</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2019.00510</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daetwyler</surname>
<given-names>H. D.</given-names>
</name>
<name>
<surname>Capitan</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Pausch</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Stothard</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>van Binsbergen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Br&#xf8;ndum</surname>
<given-names>R. F.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Whole-genome Sequencing of 234 Bulls Facilitates Mapping of Monogenic and Complex Traits in Cattle</article-title>. <source>Nat. Genet.</source> <volume>46</volume>, <fpage>858</fpage>&#x2013;<lpage>865</lpage>. <pub-id pub-id-type="doi">10.1038/ng.3034</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Diaz</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Toro</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Rekaya</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Comparison of Restricted Selection Strategies: an Application to Selection of cashmere Goats</article-title>. <source>Livestock Prod. Sci.</source> <volume>60</volume>, <fpage>89</fpage>&#x2013;<lpage>99</lpage>. </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Druet</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Macleod</surname>
<given-names>I. M.</given-names>
</name>
<name>
<surname>Hayes</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Toward Genomic Prediction from Whole-Genome Sequence Data: Impact of Sequencing Design on Genotype Imputation and Accuracy of Predictions</article-title>. <source>Heredity</source> <volume>112</volume>, <fpage>39</fpage>&#x2013;<lpage>47</lpage>. <pub-id pub-id-type="doi">10.1038/hdy.2013.13</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gaynor</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Gorjanc</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hickey</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>AlphaSimR: an R Package for Breeding Program Simulations</article-title>. <source>G3 Genes Genomes Genet.</source> <volume>11</volume>, <fpage>jkaa017</fpage>. <pub-id pub-id-type="doi">10.1093/g3journal/jkaa017</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gonen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ros-Freixedes</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Battagin</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Gorjanc</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hickey</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Method for the Allocation of Sequencing Resources in Genotyped Livestock Populations</article-title>. <source>Genet. Sel. Evol.</source> <volume>49</volume>, <fpage>47</fpage>. <pub-id pub-id-type="doi">10.1186/s12711-017-0322-5</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Habier</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tetens</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Seefried</surname>
<given-names>F.-R.</given-names>
</name>
<name>
<surname>Lichtner</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Thaller</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>The Impact of Genetic Relationship Information on Genomic Breeding Values in German holstein Cattle</article-title>. <source>Genet. Sel. Evol.</source> <volume>42</volume>, <fpage>5</fpage>. <pub-id pub-id-type="doi">10.1186/1297-9686-42-5</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hayes</surname>
<given-names>B. J.</given-names>
</name>
<name>
<surname>Bowman</surname>
<given-names>P. J.</given-names>
</name>
<name>
<surname>Chamberlain</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Verbyla</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Goddard</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Accuracy of Genomic Breeding Values in Multi-Breed Dairy Cattle Populations</article-title>. <source>Genet. Sel. Evol.</source> <volume>41</volume>, <fpage>51</fpage>. <pub-id pub-id-type="doi">10.1186/1297-9686-41-51</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Land</surname>
<given-names>A. H.</given-names>
</name>
<name>
<surname>Doig</surname>
<given-names>A. G.</given-names>
</name>
</person-group> (<year>1960</year>). <article-title>An Automatic Method of Solving Discrete Programming Problems</article-title>. <source>Econometrica</source> <volume>28</volume>, <fpage>497</fpage>&#x2013;<lpage>520</lpage>. <pub-id pub-id-type="doi">10.2307/1910129</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Sidore</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Kang</surname>
<given-names>H. M.</given-names>
</name>
<name>
<surname>Boehnke</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Abecasis</surname>
<given-names>G. R.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Low-coverage Sequencing: Implications for Design of Complex Trait Association Studies</article-title>. <source>Genome Res.</source> <volume>21</volume>, <fpage>940</fpage>&#x2013;<lpage>951</lpage>. <pub-id pub-id-type="doi">10.1101/gr.117259.110</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Luenberger</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Ye</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2015</year>). <source>Linear and Nonlinear Programming</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meuwissen</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Goddard</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Accurate Prediction of Genetic Values for Complex Traits by Whole-Genome Resequencing</article-title>. <source>Genetics</source> <volume>185</volume>, <fpage>623</fpage>&#x2013;<lpage>631</lpage>. <pub-id pub-id-type="doi">10.1534/genetics.110.116590</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moeinizade</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Hu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Schnable</surname>
<given-names>P. S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Optimizing Selection and Mating in Genomic Selection with a Look-Ahead Approach: An Operations Research Framework</article-title>. <source>G3 Genes Genomes Genet.</source> <volume>9</volume>, <fpage>2123</fpage>&#x2013;<lpage>2133</lpage>. <pub-id pub-id-type="doi">10.1534/g3.118.200842</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ros-Freixedes</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Gonen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Gorjanc</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Hickey</surname>
<given-names>J.&#x20;M.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>A Method for Allocating Low-Coverage Sequencing Resources by Targeting Haplotypes rather Than Individuals</article-title>. <source>Genet. Sel. Evol.</source> <volume>49</volume>, <fpage>78</fpage>. <pub-id pub-id-type="doi">10.1186/s12711-017-0353-y</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ros-Freixedes</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Whalen</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Gorjanc</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Milaham</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Hickey</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Evaluation of Sequencing Strategies for Whole-Genome Imputation with Hybrid Peeling</article-title>. <source>Genet. Selection Evol.</source> <volume>52</volume>, <fpage>18</fpage>. <pub-id pub-id-type="doi">10.1186/s12711-020-00537-7</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Su</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Koltes</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Saatchi</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Fernando</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Garrick</surname>
<given-names>D.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Characterizing Haplotype Diversity in Ten Us Beef Cattle Breeds</article-title>. <comment>Animal Industry Report</comment>. <comment>AS 660, ASL R2846</comment>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanRaden</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>O&#x27;Connell</surname>
<given-names>J.&#x20;R.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Fast Imputation Using Medium or Low-Coverage Sequence Data</article-title>. <source>BMC Genet.</source> <volume>16</volume>, <fpage>82</fpage>. <pub-id pub-id-type="doi">10.1186/s12863-015-0243-7</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanRaden</surname>
<given-names>P. M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Efficient Methods to Compute Genomic Predictions</article-title>. <source>J.&#x20;Dairy Sci.</source> <volume>91</volume>, <fpage>4414</fpage>&#x2013;<lpage>4423</lpage>. <pub-id pub-id-type="doi">10.3168/jds.2007-0980</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanRaden</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Van Tassell</surname>
<given-names>C. P.</given-names>
</name>
<name>
<surname>Wiggans</surname>
<given-names>G. R.</given-names>
</name>
<name>
<surname>Sonstegard</surname>
<given-names>T. S.</given-names>
</name>
<name>
<surname>Schnabel</surname>
<given-names>R. D.</given-names>
</name>
<name>
<surname>Taylor</surname>
<given-names>J.&#x20;F.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Invited Review: Reliability of Genomic Predictions for north American holstein Bulls</article-title>. <source>J.&#x20;Dairy Sci.</source> <volume>92</volume>, <fpage>16</fpage>&#x2013;<lpage>24</lpage>. <pub-id pub-id-type="doi">10.3168/jds.2008-1514</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Vazirani</surname>
<given-names>V.</given-names>
</name>
</person-group> (<year>2003</year>). <source>Approximation Algorithms</source>. <publisher-name>Springer Verlag Berlin Heidelberg GmbH</publisher-name>. </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolc</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Arango</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Settar</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Fulton</surname>
<given-names>J.&#x20;E.</given-names>
</name>
<name>
<surname>O&#x2019;Sullivan</surname>
<given-names>N. P.</given-names>
</name>
<name>
<surname>Preisinger</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Genome-wide Association Analysis and Genetic Architecture of Egg Weight and Egg Uniformity in Layer Chickens</article-title>. <source>Anim. Genet.</source> <volume>43</volume> (<issue>Suppl. 1</issue>), <fpage>87</fpage>&#x2013;<lpage>96</lpage>. <pub-id pub-id-type="doi">10.1111/j.1365-2052.2012.02381.x</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Wooliams</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Meuwisssen</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Prioritizing Animals for Dense Genotyping in Order to Impute Missing Genotypes of Sparsely Genotyped Animals</article-title>. <source>Genet. Selection Evol.</source> <volume>46</volume>, <fpage>46</fpage>. <pub-id pub-id-type="doi">10.1186/1297-9686-46-46</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>