<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="2.3" xml:lang="EN">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Plant Sci.</journal-id>
<journal-title>Frontiers in Plant Science</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Plant Sci.</abbrev-journal-title>
<issn pub-type="epub">1664-462X</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/fpls.2023.1217589</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Plant Science</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Haplotype blocks for genomic prediction: a comparative evaluation in multiple crop datasets</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Weber</surname>
<given-names>Sven E.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="author-notes" rid="fn001">
<sup>*</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1543887"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Frisch</surname>
<given-names>Matthias</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/629711"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Snowdon</surname>
<given-names>Rod J.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/114974"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Voss-Fels</surname>
<given-names>Kai P.</given-names>
</name>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/457846"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Plant Breeding, Justus Liebig University</institution>, <addr-line>Giessen</addr-line>, <country>Germany</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Department of Biometry and Population Genetics, Justus Liebig University</institution>, <addr-line>Giessen</addr-line>, <country>Germany</country>
</aff>
<aff id="aff3">
<sup>3</sup>
<institution>Institute for Grapevine Breeding, Hochschule Geisenheim University</institution>, <addr-line>Geisenheim</addr-line>, <country>Germany</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>Edited by: Lewis Lukens, University of Guelph, Canada</p>
</fn>
<fn fn-type="edited-by">
<p>Reviewed by: Jos&#xe9; Marcelo Soriano Viana, Universidade Federal de Vi&#xe7;osa, Brazil; Valerio Hoyos-Villegas, McGill University, Canada</p>
</fn>
<fn fn-type="corresp" id="fn001">
<p>*Correspondence: Sven E. Weber, <email xlink:href="mailto:Sven.E.Weber@agrar.uni-giessen.de">Sven.E.Weber@agrar.uni-giessen.de</email>
</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>09</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>14</volume>
<elocation-id>1217589</elocation-id>
<history>
<date date-type="received">
<day>05</day>
<month>05</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>08</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2023 Weber, Frisch, Snowdon and Voss-Fels</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Weber, Frisch, Snowdon and Voss-Fels</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>In modern plant breeding, genomic selection is becoming the gold standard for selection of superior genotypes. The basis for genomic prediction models is a set of phenotyped lines along with their genotypic profile. With high marker density and linkage disequilibrium (LD) between markers, genotype data in breeding populations tends to exhibit considerable redundancy. Therefore, interest is growing in the use of haplotype blocks to overcome redundancy by summarizing co-inherited features. Moreover, haplotype blocks can help to capture local epistasis caused by interacting loci. Here, we compared genomic prediction methods that either used single SNPs or haplotype blocks with regards to their prediction accuracy for important traits in crop datasets. We used four published datasets from canola, maize, wheat and soybean. Different approaches to construct haplotype blocks were compared, including blocks based on LD, physical distance, number of adjacent markers and the algorithms implemented in the software &#x201c;<italic>Haploview</italic>&#x201d; and <italic>&#x201c;HaploBlocker&#x201d;</italic>. The tested prediction methods included Genomic Best Linear Unbiased Prediction (GBLUP), Extended GBLUP to account for additive by additive epistasis (EGBLUP), Bayesian LASSO and Reproducing Kernel Hilbert Space (RKHS) regression. We found improved prediction accuracy in some traits when using haplotype blocks compared to SNP-based predictions, however the magnitude of improvement was very trait- and model-specific. Especially in settings with low marker density, haplotype blocks can improve genomic prediction accuracy. In most cases, physically large haplotype blocks yielded a strong decrease in prediction accuracy. Especially when prediction accuracy varies greatly across different prediction models, prediction based on haplotype blocks can improve prediction accuracy of underperforming models. However, there is no &#x201c;best&#x201d; method to build haplotype blocks, since prediction accuracy varied considerably across methods and traits. Hence, criteria used to define haplotype blocks should not be viewed as fixed biological parameters, but rather as hyperparameters that need to be adjusted for every dataset.</p>
</abstract>
<kwd-group>
<kwd>genomic selection</kwd>
<kwd>SNP markers</kwd>
<kwd>haploblocks</kwd>
<kwd>haplotype blocks</kwd>
<kwd>genomic prediction</kwd>
</kwd-group>
<contract-sponsor id="cn001">Bundesministerium f&#xfc;r Bildung und Forschung<named-content content-type="fundref-id">10.13039/501100002347</named-content>
</contract-sponsor>
<counts>
<fig-count count="2"/>
<table-count count="2"/>
<equation-count count="6"/>
<ref-count count="119"/>
<page-count count="17"/>
<word-count count="10838"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-in-acceptance</meta-name>
<meta-value>Plant Breeding</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec id="s1" sec-type="intro">
<label>1</label>
<title>Introduction</title>
<p>Genomic prediction has greatly improved animal and plant breeding (<xref ref-type="bibr" rid="B45">Hickey et&#xa0;al., 2017</xref>) and has the potential to improve genetic gain even in crops with complex genomes (<xref ref-type="bibr" rid="B103">Voss-Fels et&#xa0;al., 2021</xref>). In the past, predictions based on linear mixed models used relatedness to borrow information on target phenotypes of relatives. <xref ref-type="bibr" rid="B42">Henderson (1975)</xref> derived this relationship from pedigrees <italic>via</italic> the numerator relationship matrix with the expectation that each parent contributes exactly 50% of its genome to its offspring. With the advance of sequencing technology nowadays, genomic data is used to replace the pedigree relationship with realized relationships calculated from dense marker maps. Furthermore, with the inclusion of genetic markers, information about linkage disequilibrium and cosegregation is available for genomic prediction (<xref ref-type="bibr" rid="B37">Habier et&#xa0;al., 2013</xref>). Today, individuals in breeding populations of major crops can be sequenced with high quality at low costs, enabling the identification of millions of genome-wide single nucleotide polymorphism (SNP) markers that can be easily screened in large populations using high-throughput genotyping technologies. Together with phenotype measurements, genome-wide marker profiles can be used to predict breeding values of non-phenotyped individuals (<xref ref-type="bibr" rid="B59">Lande and Thompson, 1990</xref>; <xref ref-type="bibr" rid="B7">Bernardo, 1994</xref>; <xref ref-type="bibr" rid="B75">Meuwissen et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B97">VanRaden, 2008</xref>). This can assist breeders in the accurate identification of superior genotypes within their breeding material without the need for additional phenotyping. Moreover, it can facilitate the decision-making process for selecting which genotypes should undergo phenotyping, leading to reduced phenotyping costs and improved accuracy in estimating breeding values. Hence, genomic selection has the potential to considerably increase genetic gain and profit in many crops (<xref ref-type="bibr" rid="B103">Voss-Fels et&#xa0;al., 2021</xref>).</p>
<p>There are a variety of statistical methods for genome-based predictions (e.g. <xref ref-type="bibr" rid="B97">VanRaden, 2008</xref>; <xref ref-type="bibr" rid="B27">de los Campos et&#xa0;al., 2009</xref>; <xref ref-type="bibr" rid="B115">Zhang et&#xa0;al., 2010</xref>; <xref ref-type="bibr" rid="B34">Gianola, 2013</xref>; <xref ref-type="bibr" rid="B49">Hofheinz and Frisch, 2014</xref>; <xref ref-type="bibr" rid="B108">Werner et&#xa0;al., 2018a</xref>; <xref ref-type="bibr" rid="B76">Millet et&#xa0;al., 2019</xref>), differing in their assumptions of variance components, marker effects or marker modes of action. Examples for genomic prediction models are ridge regression BLUP, GBLUP (<xref ref-type="bibr" rid="B7">Bernardo, 1994</xref>; <xref ref-type="bibr" rid="B75">Meuwissen et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B97">VanRaden, 2008</xref>), Reproducing Kernel Hilbert Space Regression (RKHS) (<xref ref-type="bibr" rid="B27">de los Campos et&#xa0;al., 2009</xref>), as well as Bayesian models like Bayesian LASSO (<xref ref-type="bibr" rid="B80">Park and Casella, 2008</xref>) or Bayesian ridge regression (<xref ref-type="bibr" rid="B81">P&#xe9;rez and de los Campos, 2014</xref>).</p>
<p>However, biallelic SNPs are sometimes unable to identify all variants and allelic combinations of genes that contribute to a particular trait, since most genes carry multiple sequence polymorphisms. Furthermore, accurate genomic prediction is often obtained based on close relatives (<xref ref-type="bibr" rid="B97">VanRaden, 2008</xref>; <xref ref-type="bibr" rid="B40">Hayes et&#xa0;al., 2009</xref>) while this accuracy decreases as the validation individuals get more unrelated (<xref ref-type="bibr" rid="B38">Habier et&#xa0;al., 2010</xref>; <xref ref-type="bibr" rid="B110">Wolc et&#xa0;al., 2011</xref>). This implies that SNPs are not necessarily in LD with causal QTL and the prediction accuracy is at least partly driven by implicitly capturing relationship among individuals. Hence, one strategy to improve predictions is increasing marker density. With the advance of whole genome sequencing technologies, increasingly large and dense marker datasets can today be generated for most major crops (<xref ref-type="bibr" rid="B31">Edwards and Batley, 2010</xref>; <xref ref-type="bibr" rid="B114">Yu et&#xa0;al., 2011</xref>). However, increasing marker density does not consistently improve prediction accuracies (<xref ref-type="bibr" rid="B90">Solberg et&#xa0;al., 2008</xref>; <xref ref-type="bibr" rid="B30">Druet et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B41">Hayes et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B79">Norman et&#xa0;al., 2018</xref>) and often improvements are only observed following pre-selection of markers (<xref ref-type="bibr" rid="B95">van Binsbergen et&#xa0;al., 2015</xref>; <xref ref-type="bibr" rid="B78">Ni et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B84">Raymond et&#xa0;al., 2018</xref>). Furthermore, prediction accuracy is influenced by trait heritability (<xref ref-type="bibr" rid="B117">Zhang et&#xa0;al., 2017</xref>) and the number of genotypes with phenotypic records available for genomic selection. Hence, another approach to enhance prediction accuracy is by increasing the number of phenotyped lines used for model training (<xref ref-type="bibr" rid="B98">VanRaden et&#xa0;al., 2009</xref>; <xref ref-type="bibr" rid="B15">Combs and Bernardo, 2013</xref>). However, due to the high costs associated with phenotyping, this may not always be feasible, particularly when sparse testing methods (<xref ref-type="bibr" rid="B52">Jarquin et&#xa0;al., 2020</xref>; <xref ref-type="bibr" rid="B18">Crespo-Herrera et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B1">Atanda et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B94">Terraillon et&#xa0;al., 2022</xref>) are not applicable. Hence, one strategy to address low prediction accuracy could be to identify more informative variants for predictions without necessarily increasing the marker density <italic>per se</italic>.</p>
<p>Loci along the genome are usually inherited in a block-like structure, with only few recombination hotspots (<xref ref-type="bibr" rid="B25">Daly et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B54">Jeffreys et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B85">Reich et&#xa0;al., 2001</xref>) defining the so-called haplotype blocks. There are several ways to define a haplotype block, for example as a fixed window of adjacent markers, as a fixed window of adjacent base pairs, or based on a statistical measure of LD. While the first two are straightforward and simple, they may not represent haplotype blocks in a true biological sense. More sophisticated approaches may model the true haplotype blocks better. Commonly, LD based measures like D&#xb4; or <italic>r<sup>2</sup>
</italic> are used for construction of haplotype blocks (<xref ref-type="bibr" rid="B29">Devlin and Risch, 1995</xref>). Furthermore, prior information of interaction between adjacent markers may help model local epistasis (<xref ref-type="bibr" rid="B67">Liu et&#xa0;al., 2019</xref>), however, difficulties in computing higher order interactions limits the size of haplotype blocks of that type. Haplotype blocks are assumed to be in higher linkage disequilibrium with QTL, and it was proven that haplotype blocks are able to capture local epistasis of markers in close proximity (<xref ref-type="bibr" rid="B56">Jiang et&#xa0;al., 2018</xref>). Furthermore, it has been suggested that the problem of apparent or phantom epistasis, which occurs between markers and QTL in incomplete LD, can be overcome with haplotype blocks (<xref ref-type="bibr" rid="B111">Wood et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B28">de los Campos et&#xa0;al., 2019</xref>). Hence it can be assumed, that haplotype blocks may improve genomic prediction.</p>
<p>In genomic selection, there is evidence that markers grouped to haplotype blocks can improve genomic prediction (<xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B56">Jiang et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B2">Ballesta et&#xa0;al., 2019</xref>), while other studies delivered evidence against improving predictions (<xref ref-type="bibr" rid="B90">Solberg et&#xa0;al., 2008</xref>). Even with the methods described above for construction of haplotype blocks, it is always necessary to set appropriate hyperparameters like window size or an LD threshold to define block boundaries. Most previous studies in this area investigated a small range of LD thresholds, adjacent markers or window sizes in association studies and genomic prediction (<xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B44">Hess et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B71">Maldonado et&#xa0;al., 2019</xref>). However, in terms of genomic prediction for plant breeding the huge variety of options and hyperparameters possible to construct haplotype blocks were not assessed in detail. Hence, the present study sought to investigate the following questions: 1.) How does the method of building haplotype blocks and its parameters affect the number of haplotypes? 2.) Are haplotype block predictions different from SNP predictions in terms of prediction accuracy? 3.) Is there a preferable haplotype construction method to improve genomic prediction?</p>
<p>These questions were addressed by employing various methods for constructing haplotypes, which are commonly discussed in the literature. The methods range from simple approaches such as marker adjacency (<xref ref-type="bibr" rid="B99">Villumsen and Janss, 2009</xref>; <xref ref-type="bibr" rid="B100">Villumsen et&#xa0;al., 2009</xref>; <xref ref-type="bibr" rid="B56">Jiang et&#xa0;al., 2018</xref>; <xref ref-type="bibr" rid="B66">Liang et&#xa0;al., 2020</xref>) and physical distances (<xref ref-type="bibr" rid="B44">Hess et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B66">Liang et&#xa0;al., 2020</xref>) to more sophisticated methods based on LD thresholds (<xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B102">Voss-Fels et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B6">Bayer et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B64">Li et&#xa0;al., 2022</xref>) the confidence intervals of <italic>D&#xb4;</italic> method described by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref>, the <italic>Four-gamete Rule</italic> method described by <xref ref-type="bibr" rid="B105">Wang et&#xa0;al. (2002)</xref> the <italic>Solid Spine of LD</italic> method (<xref ref-type="bibr" rid="B4">Barrett et&#xa0;al., 2005</xref>) and <italic>&#x201c;HaploBlocker&#x201d;</italic> (<xref ref-type="bibr" rid="B83">Pook et&#xa0;al., 2019</xref>), using four example datasets from canola, maize, wheat and soybean. To assess prediction accuracy, genomic prediction was performed using GBLUP, Bayesian LASSO, EGBLUP and RKHS models.</p>
</sec>
<sec id="s2" sec-type="materials|methods">
<label>2</label>
<title>Materials and methods</title>
<sec id="s2_1">
<label>2.1</label>
<title>Datasets</title>
<p>The datasets examined in this study are all publicly available. The canola dataset is from a spring-type canola hybrid breeding program (<xref ref-type="bibr" rid="B51">Jan et&#xa0;al., 2016</xref>). Briefly, 475 double haploid (DH) pollinators were crossed with two male sterile lines to create 950 F<sub>1</sub> test hybrids. The hybrids where subsequently tested for seed yield, flowering time, field emergence, lodging, oil yield and glucosinolate content. For 910 test hybrids the complete phenotypic records were available, and all parental lines were genotyped with the Illumina <italic>Brassica</italic> 60k SNP array (<xref ref-type="bibr" rid="B14">Clarke et&#xa0;al., 2016</xref>). The maize dataset is derived from 847 test hybrids from a diverse dent nested association mapping population described by <xref ref-type="bibr" rid="B5">Bauer et&#xa0;al. (2013)</xref> consisting of 10 half-sib DH families. Double haploid lines were all crossed to the common flint line UH007 and F1 hybrids were phenotypically analyzed for dry matter yield (DMY), dry matter content (DMC), plant height (PH), days till tasseling (DtTAS) and days till silking (DtSILK), as described by <xref ref-type="bibr" rid="B61">Lehermeier et&#xa0;al. (2014)</xref>. All DH lines were genotyped with the Illumina MaizeSNP50 SNP array (<xref ref-type="bibr" rid="B14">Clarke et&#xa0;al., 2016</xref>). The wheat dataset, described in <xref ref-type="bibr" rid="B102">Voss-Fels et&#xa0;al. (2019)</xref>, consists of 191 released wheat varieties from 1966 to 2013 that were tested under three agrichemical treatments for a wide range of agronomic traits including yield, biomass yield, falling number, days till heading, plant height, harvest index kernel spike<sup>-1</sup>, nitrogen use efficiency (NUE), powdery mildew resistance, protein content, protein yield sedimentation value spike m<sup>-2</sup>, stripe rust and thousand kernel weight (TKW). All lines were genotyped with the Illumina 15k wheat SNP array described in <xref ref-type="bibr" rid="B91">Soleimani et&#xa0;al. (2020)</xref>. The soybean dataset consisted out of 1000 lines from the USDA Soybean Germplasm Collection (<xref ref-type="bibr" rid="B36">Grant et&#xa0;al., 2010</xref>) with phenotypic records for protein and oil content (PC, OC) (<xref ref-type="bibr" rid="B3">Bandillo et&#xa0;al., 2015</xref>). For all lines, genotypic information from the Illumina Infinium SoySNP50K BeadChip (<xref ref-type="bibr" rid="B92">Song et&#xa0;al., 2013</xref>) was available.</p>
<p>With the exception of the maize dataset, all phenotypic data represented adjusted trait means per genotype. The published field data from the maize population was adjusted following methods used for phenotypic data analyses from the original publication.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Genotypic data</title>
<p>With exception of the canola dataset, physical SNP marker positions were obtained from the respective reference genome assemblies used in the original publications, namely the <italic>Brassica napus</italic> Express 617 genome (<xref ref-type="bibr" rid="B60">Lee et&#xa0;al., 2020</xref>), the maize B73 AGPv2 genome (<xref ref-type="bibr" rid="B88">Schnable et&#xa0;al., 2009</xref>), the wheat Chinese Spring IWGCS reference Sequence v1.0 (<xref ref-type="bibr" rid="B119">Zimin et&#xa0;al., 2017</xref>) and the soybean Glyma1.01 reference (<xref ref-type="bibr" rid="B87">Schmutz et&#xa0;al., 2010</xref>). In general, only markers with a unique physical position on the reference genome, a minor allele frequency &#x2265; 0.05 and a maximum of 10% missing values in each population were used for further analyses. This left a total of 29385, 32363, 8710 and 35821 markers for the canola, maize, wheat and soybean datasets, respectively. This corresponds to a marker density of 31.78, 15.63, 0.57 and 37.48 SNPs mbp<sup>-1</sup> in canola, maize, wheat and soy respectively. After filtering, markers were imputed with the software &#x201c;<italic>BEAGLE&#x201d;</italic> V5.2 (<xref ref-type="bibr" rid="B10">Browning and Browning, 2007</xref>; <xref ref-type="bibr" rid="B11">Browning et&#xa0;al., 2018</xref>).</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Haplotype block construction</title>
<p>We considered seven haplotype block construction methods based on (i) pre-determined LD thresholds, (ii) fixed windows of adjacent markers, (iii) fixed windows of adjacent base pairs, (iv) <italic>&#x201c;HaploBlocker&#x201d;</italic> (<xref ref-type="bibr" rid="B83">Pook et&#xa0;al., 2019</xref>), (v) the confidence intervals of <italic>D&#xb4;</italic> method described by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref>, (vi) the <italic>Four-gamete Rule</italic> method described by <xref ref-type="bibr" rid="B105">Wang et&#xa0;al. (2002)</xref> and (vii) the <italic>Solid Spine of LD</italic> method (<xref ref-type="bibr" rid="B4">Barrett et&#xa0;al., 2005</xref>). The first three methods were implemented in the r package <italic>&#x201c;SelectionTools&#x201d;</italic> (downloadable at <ext-link ext-link-type="uri" xlink:href="http://population-genetics.uni-giessen.de/~software/">http://population-genetics.uni-giessen.de/~software/</ext-link>), while the latter three are implemented in the software &#x201c;<italic>Haploview&#x201d;</italic> v4.1 (<xref ref-type="bibr" rid="B4">Barrett et&#xa0;al., 2005</xref>). The different approaches are described in detail below. These methods were selected for their widespread use in haplotype block formation and their distinct characteristics. Methods such as the pre-determined LD threshold, confidence intervals of <italic>D&#x2019;</italic>, the <italic>Four-gamete Rule</italic>, and the <italic>Solid Spine of LD</italic> are based on linkage disequilibrium (LD) and gamete frequency. They aim to model historical recombination hotspots and generate meaningful blocks within populations. However, these blocks do not necessarily represent functional groups. Therefore, we also included methods based on fixed windows to assess blocks that would not be constructed based on population-based measures alone. Additionally, while most methods consider block borders across the entire population, it is important to note that subpopulations or genotypes may have different recombination patterns. To account for this, we utilized the method <italic>&#x201c;HaploBlocker&#x201d;</italic> described in <xref ref-type="bibr" rid="B83">Pook et&#xa0;al. (2019)</xref> to construct haplotype blocks specific to different groups.</p>
<sec id="s2_3_1">
<label>2.3.1</label>
<title>LD threshold</title>
<p>LD between markers on the same chromosome was calculated as <italic>r<sup>2</sup>
</italic> (<xref ref-type="bibr" rid="B48">Hill and Robertson, 1968</xref>) in <italic>&#x201c;SelectionTools&#x201d;</italic>. Haplotype blocks were built by starting with the two neighboring markers with the highest LD. If the pairwise LD exceeded a certain threshold, those markers were then assigned to a haplotype block. In the next step, if the LD between the next immediately adjacent markers and the markers at the block border again exceeded the threshold, the block was extended. This was done until no more markers fulfilled this criterion and the algorithm started over again with new markers. To account for misplaced markers, a tolerance parameter of 1 was used, meaning that one marker that did not fulfill the LD threshold was accepted if the next flanking marker fulfilled the LD criterion. Thresholds were set sequentially from 0.01 to 1 with a step size of 0.01, resulting in 100 different LD thresholds. Using very high thresholds to form blocks effectively eliminates redundant information, making these scenarios similar to LD pruning, which has been shown to improve prediction accuracy (<xref ref-type="bibr" rid="B113">Ye et&#xa0;al., 2019</xref>). On the other hand, very low thresholds result in the formation of large blocks commonly observed in introgression breeding, where recombination is sometimes very limited (<xref ref-type="bibr" rid="B39">Hao et&#xa0;al., 2020</xref>).</p>
</sec>
<sec id="s2_3_2">
<label>2.3.2</label>
<title>Fixed windows of adjacent markers</title>
<p>Starting at the beginning of each chromosome, haplotype blocks consisting of <italic>m</italic> neighboring markers were constructed until all markers on a chromosome were assigned to blocks. We considered &#x2308;2<italic>
<sup>x</sup>
</italic>&#x2309; markers with <italic>x</italic> being {1, 1.5, 2, 2.5 &#x2026;}, until in the most excessive case all markers of a chromosome represented a haplotype block containing all markers of that chromosome. We chose to create blocks of such large size to address scenarios where entire chromosomes or large segments play an important role in traits, as well as scenarios related to introgression breeding, where recombination is limited (<xref ref-type="bibr" rid="B39">Hao et&#xa0;al., 2020</xref>).</p>
</sec>
<sec id="s2_3_3">
<label>2.3.3</label>
<title>Fixed windows of adjacent base pairs</title>
<p>Starting at the beginning of each chromosome, haplotype blocks of <italic>m</italic> consecutive base pairs were constructed until the whole chromosome was partitioned into blocks. We considered &#x2308;2<italic>
<sup>x</sup>
</italic>&#x2309; base pairs with <italic>x</italic> being {10, 10.5, 11, 11.5 &#x2026;} until in the most excessive case a whole chromosome represented a block. Similar to the approach using fixed windows of adjacent markers, we selected to construct blocks of considerable size to accommodate scenarios where entire chromosomes or large segments influence traits, as well as situations related to introgression breeding characterized by limited recombination (<xref ref-type="bibr" rid="B39">Hao et&#xa0;al., 2020</xref>).</p>
</sec>
<sec id="s2_3_4">
<label>2.3.4</label>
<title>HaploBlocker</title>
<p>Since different subpopulations might result in different block borders, we also built haplotype blocks with the algorithm of <xref ref-type="bibr" rid="B83">Pook et&#xa0;al. (2019)</xref>. This algorithm relies on linkage instead of linkage disequilibrium to construct haplotype blocks. Here blocks are defined as consecutive sequence of genetic markers with a predefined frequency, a sequence of haplotype merging and splitting steps is applied to construct subgroup-specific haplotype blocks. This algorithm allows subgroup specific haplotype block borders. The algorithm was conducted with default settings with the r package <italic>&#x201c;HaploBlocker&#x201d;</italic> (<xref ref-type="bibr" rid="B83">Pook et&#xa0;al., 2019</xref>).</p>
</sec>
<sec id="s2_3_5">
<label>2.3.5</label>
<title>Gabriel algorithm</title>
<p>The algorithm developed by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref> (GAB) for the Human Haplotype Map generates 95% confidence bounds on <italic>D&#xb4;</italic> between all intrachromosomal marker pairs. Marker pairs are considered in &#x201c;strong LD&#x201d; if the one-sided upper 95% <italic>D&#xb4;</italic> confidence bound is higher than 0,98 and the lower bound is higher than 0.7. Markers in &#x201c;strong LD&#x201d; are consequently grouped into blocks. Blocks are extended until the outermost marker pairs don&#xb4;t fulfill this criterion anymore.</p>
</sec>
<sec id="s2_3_6">
<label>2.3.6</label>
<title>Four gamete rule</title>
<p>The <italic>Four Gamete Rule</italic> (GAM) described by <xref ref-type="bibr" rid="B105">Wang et&#xa0;al. (2002)</xref> groups consecutive markers into haplotype blocks if no evidence for a historical recombination event can be found between all marker pairs of a block. A historical recombination is defined if all four haplotypes of the new marker and any other previous marker are found with at least 1% frequency. If this is the case, a block border is created between those markers and the algorithm starts with a new block.</p>
</sec>
<sec id="s2_3_7">
<label>2.3.7</label>
<title>Solid spine of LD</title>
<p>The <italic>Solid Spine of LD</italic> method (SPI), introduced by the developers of &#x201c;<italic>Haploview</italic>&#x201d; (<xref ref-type="bibr" rid="B4">Barrett et&#xa0;al., 2005</xref>), searches for a spine of strong LD by calculation of LD between all intrachromosomal marker pairs. In this method, two markers on the same chromosome form a block border if the pairwise <italic>D&#xb4;</italic> is higher than 0.8. All markers in that window form the block. This allows for intermediate markers to not be in LD.</p>
</sec>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Genomic prediction models</title>
<p>In total, four genomic selection models were used to predict testcross (maize, canola) and inbred line (soybean, wheat) performance, respectively. The models represent two variations of the GBLUP and two models implemented in a Bayesian framework. The frequentist models were GBLUP (<xref ref-type="bibr" rid="B7">Bernardo, 1994</xref>; <xref ref-type="bibr" rid="B75">Meuwissen et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B97">VanRaden, 2008</xref>) and extended GBLUP to account for second-order additive*additive epistasis, following the EGBLUP model of <xref ref-type="bibr" rid="B55">Jiang and Reif (2015)</xref>. The Bayesian model included the Bayesian LASSO model (<xref ref-type="bibr" rid="B80">Park and Casella, 2008</xref>) which offers the capability of marker-specific shrinkage, and the semiparametric RKHS regression model (<xref ref-type="bibr" rid="B27">de los Campos et&#xa0;al., 2009</xref>) which allows modeling of higher order epistasis.</p>
<p>In the GBLUP and EGBLUP the underlying model is assumed to be:</p>
<disp-formula>
<mml:math display="block" id="M1">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>X</mml:mi>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
<mml:mi>a</mml:mi>
<mml:mo>+</mml:mo>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mi>i</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im1">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula> is a vector of observations for a trait under consideration, <inline-formula>
<mml:math display="inline" id="im2">
<mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is a vector of fixed non-genetic effects, <inline-formula>
<mml:math display="inline" id="im3">
<mml:mi>a</mml:mi>
</mml:math>
</inline-formula> is a vector of random additive effects, <inline-formula>
<mml:math display="inline" id="im4">
<mml:mi>i</mml:mi>
</mml:math>
</inline-formula> is a vector of random epistatic effects and <inline-formula>
<mml:math display="inline" id="im5">
<mml:mi>e</mml:mi>
</mml:math>
</inline-formula> is the random residual term. <inline-formula>
<mml:math display="inline" id="im6">
<mml:mrow>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mi>a</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im7">
<mml:mrow>
<mml:msub>
<mml:mi>Z</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are design matrices relating the random effects to the phenotypic records. <inline-formula>
<mml:math display="inline" id="im8">
<mml:mi>X</mml:mi>
</mml:math>
</inline-formula> is the design matrix for fixed effects and, in the case of the canola and soybean datasets, a vector of ones modeling the intercept ( <inline-formula>
<mml:math display="inline" id="im9">
<mml:mrow>
<mml:msub>
<mml:mn>1</mml:mn>
<mml:mi>n</mml:mi>
</mml:msub>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>). In the wheat dataset, two additional fixed effects for N fertilization and fungicide treatment were added, while in the maize dataset an additional 10 columns were added to assign individuals to half-sib families.</p>
<p>It is assumed that</p>
<disp-formula>
<mml:math display="block" id="M2">
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mo>~</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>G</mml:mi>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>i</mml:mi>
<mml:mo>~</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>a</mml:mi>
<mml:mi>n</mml:mi>
<mml:mi>d</mml:mi>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>~</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>I</mml:mi>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>e</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im10">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>a</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im11">
<mml:mrow>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>e</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
</mml:math>
</inline-formula> are additive genetic variance, epistatic genetic variance and residual variance respectively. <inline-formula>
<mml:math display="inline" id="im12">
<mml:mi>G</mml:mi>
</mml:math>
</inline-formula> and <inline-formula>
<mml:math display="inline" id="im13">
<mml:mrow>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> are the additive and epistatic relationship matrices, respectively. <inline-formula>
<mml:math display="inline" id="im14">
<mml:mi>I</mml:mi>
</mml:math>
</inline-formula> is an identity matrix. Depending on inclusion of epistatic effects the epistasis terms were included or omitted.</p>
<p>The additive genomic relationship matrix was calculated following <xref ref-type="bibr" rid="B97">VanRaden (2008)</xref>:</p>
<disp-formula>
<mml:math display="block" id="M3">
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mo>=</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>Z</mml:mi>
<mml:mi>Z</mml:mi>
<mml:mo>&#xb4;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mo>&#x2211;</mml:mo>
<mml:mo>&#x200b;</mml:mo>
</mml:msup>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:math>
</disp-formula>
<p>with the elements of <inline-formula>
<mml:math display="inline" id="im15">
<mml:mi>Z</mml:mi>
</mml:math>
</inline-formula> being (0-2p<sub>i</sub>) for genotype H<sub>i</sub>H<sub>i</sub>, (1-2p<sub>i</sub>) for genotype H<sub>i</sub>H<sub>j</sub> and (2-2p<sub>i</sub>) for genotype H<sub>j</sub>H<sub>j</sub>, where H<sub>j</sub> is the haplotype (treated as a single marker) within a haplotype block, H<sub>i</sub> is any other haplotype within that haplotype block except H<sub>i</sub>, and <inline-formula>
<mml:math display="inline" id="im16">
<mml:mrow>
<mml:msub>
<mml:mi>p</mml:mi>
<mml:mi>i</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is the frequency of the <italic>i</italic>th haplotype in a haplotype block. Haplotype blocks with only two haplotypes were treated like standard biallelic markers. For the canola dataset, prior to construction of the genomic relationship matrix, parental genotypes were crossed <italic>in silico</italic> to derive hybrid genotypes, as described by <xref ref-type="bibr" rid="B108">Werner et&#xa0;al. (2018a)</xref>.</p>
<p>According to <xref ref-type="bibr" rid="B43">Henderson (1985)</xref> and <xref ref-type="bibr" rid="B55">Jiang and Reif (2015)</xref>, the second order (additive*additive) epistatic relationship matrix can be approximated with <inline-formula>
<mml:math display="inline" id="im17">
<mml:mrow>
<mml:msub>
<mml:mi>G</mml:mi>
<mml:mrow>
<mml:mi>a</mml:mi>
<mml:mi>a</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>=</mml:mo>
<mml:mi>G</mml:mi>
<mml:mo>#</mml:mo>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula>, with <inline-formula>
<mml:math display="inline" id="im18">
<mml:mo>#</mml:mo>
</mml:math>
</inline-formula> denoting the pointwise (hadamard) product operation.</p>
<p>GBLUP and EGBLUP were implemented and solved with the R package <italic>sommer</italic> (<xref ref-type="bibr" rid="B16">Covarrubias-Pazaran, 2016</xref>; <xref ref-type="bibr" rid="B17">Covarrubias-Pazaran, 2018</xref>).</p>
<p>The general formula describing the model Bayesian LASSO model of <xref ref-type="bibr" rid="B80">Park and Casella (2008)</xref> is:</p>
<disp-formula>
<mml:math display="block" id="M4">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>X</mml:mi>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>M</mml:mi>
<mml:mi>f</mml:mi>
<mml:mo>+</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im19">
<mml:mi>y</mml:mi>
</mml:math>
</inline-formula> is the vector of observations for a trait under consideration, <inline-formula>
<mml:math display="inline" id="im20">
<mml:mi>&#x3b2;</mml:mi>
</mml:math>
</inline-formula> is a vector of fixed non-genetic effects, a is a vector of additive effects. <inline-formula>
<mml:math display="inline" id="im21">
<mml:mi>X</mml:mi>
</mml:math>
</inline-formula> is the design matrix as described in the GBLUP section. <inline-formula>
<mml:math display="inline" id="im22">
<mml:mi>M</mml:mi>
</mml:math>
</inline-formula> is an incidence matrix relating phenotypic records with the respective marker/haplotype profiles coded 0, 1, 2. The coefficients of the fixed ( <inline-formula>
<mml:math display="inline" id="im23">
<mml:mi>&#x3b2;</mml:mi>
</mml:math>
</inline-formula>) effects are assigned flat priors, while the coefficients of the marker/haplotype effects ( <inline-formula>
<mml:math display="inline" id="im24">
<mml:mi>f</mml:mi>
</mml:math>
</inline-formula>) are assigned double-exponential priors. This allows the shrinkage of some marker/haplotype effects to effectively zero, introducing sparsity into the model. This model was tested because we assumed that some marker variants and particularly some haplotypes would have no effect on some traits. Here, <inline-formula>
<mml:math display="inline" id="im25">
<mml:mi>e</mml:mi>
</mml:math>
</inline-formula> is the random residual term. In the Bayesian LASSO, only additive effects were modeled, because additional effects in this framework would increase the computational burden to an unacceptable degree. This model was conducted in the r software with the package <italic>BGLR</italic> (<xref ref-type="bibr" rid="B81">P&#xe9;rez and de los Campos, 2014</xref>) using the default parameters.</p>
<p>Following <xref ref-type="bibr" rid="B27">de los Campos et&#xa0;al. (2009)</xref> with kernel averaging, the RKHS model has following form:</p>
<disp-formula>
<mml:math display="block" id="M5">
<mml:mrow>
<mml:mi>y</mml:mi>
<mml:mo>=</mml:mo>
<mml:mi>X</mml:mi>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>+</mml:mo>
<mml:msubsup>
<mml:mo>&#x2211;</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>3</mml:mn>
</mml:mrow>
<mml:mi>L</mml:mi>
</mml:msubsup>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:mo>+</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
</mml:math>
</disp-formula>
<p>with</p>
<disp-formula>
<mml:math display="block" id="M6">
<mml:mrow>
<mml:mi>p</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mn>1</mml:mn>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>L</mml:mi>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x221d;</mml:mo>
<mml:mtext>&#xa0;</mml:mtext>
<mml:msubsup>
<mml:mo>&#x220f;</mml:mo>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>=</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mi>L</mml:mi>
</mml:msubsup>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:msub>
<mml:mi>K</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mrow>
<mml:mi>u</mml:mi>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mn>2</mml:mn>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mo stretchy="false">(</mml:mo>
<mml:mi>e</mml:mi>
<mml:mo>&#x2223;</mml:mo>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:mo>&#xa0;</mml:mo>
<mml:mi>I</mml:mi>
<mml:msubsup>
<mml:mi>&#x3c3;</mml:mi>
<mml:mi>e</mml:mi>
<mml:mn>2</mml:mn>
</mml:msubsup>
<mml:mo stretchy="false">)</mml:mo>
<mml:mo>&#xa0;</mml:mo>
</mml:mrow>
</mml:math>
</disp-formula>
<p>where <inline-formula>
<mml:math display="inline" id="im26">
<mml:mrow>
<mml:msub>
<mml:mi>K</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is an <inline-formula>
<mml:math display="inline" id="im27">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>*</mml:mo>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> kernel. It is calculated from the Euclidean distance between genotypes based on their marker/haplotype profile. We selected a Gaussian kernel with the <italic>l</italic>th value of the bandwidth parameter {0.1, 0.5, 2.5}. <inline-formula>
<mml:math display="inline" id="im28">
<mml:mrow>
<mml:mi>X</mml:mi>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
</mml:math>
</inline-formula> is treated in a similar manner to the Bayesian LASSO and <inline-formula>
<mml:math display="inline" id="im29">
<mml:mrow>
<mml:msub>
<mml:mi>u</mml:mi>
<mml:mi>l</mml:mi>
</mml:msub>
</mml:mrow>
</mml:math>
</inline-formula> is assumed to be the random genomic effect. That way the different random effects, i.e. the three kernel matrices from the three bandwidth parameters, are weighted by their variance components. Here, <inline-formula>
<mml:math display="inline" id="im30">
<mml:mi>e</mml:mi>
</mml:math>
</inline-formula> is the random residual term. As for the Bayesian LASSO, the RKHS model was conducted in the r software with the package <italic>BGLR</italic> (<xref ref-type="bibr" rid="B81">P&#xe9;rez and de los Campos, 2014</xref>) using the default parameters.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Genomic relationship</title>
<p>Generally, constructing haplotype blocks applies a transformation to the original marker data. To assess how well the marker data is also captured by haplotype blocks, we used the relationship coefficients obtained from the relationship matrix calculated following <xref ref-type="bibr" rid="B97">VanRaden (2008)</xref> (see above) and calculated the Pearson correlation between relationship coefficients obtained from SNPs and those obtained from haplotype blocks.</p>
</sec>
<sec id="s2_6">
<label>2.6</label>
<title>Evaluation of prediction accuracy</title>
<p>For all the four datasets, model performance was assessed by running 100 cross-validation runs, where each cycle consisted of splitting the population into 80% training population and 20% validation population. Each model was trained on the training population and then this model was used to predict the validation population with masked phenotypic data. Furthermore, in the maize dataset, a family wise cross validation was conducted. This was done to test how predictive haplotype blocks are to predict genetically distant individuals. Here, the dataset was split according to the family assignment of the nested association mapping population and each family served once as validation set. In both cross validation schemes, the Pearson correlation coefficient (r) between observed and predicted phenotypic values of the validation population was used as a measure of prediction accuracy.</p>
</sec>
</sec>
<sec id="s3" sec-type="results">
<label>3</label>
<title>Results</title>
<sec id="s3_1">
<label>3.1</label>
<title>Haplotype block properties</title>
<p>In all the datasets analyzed, haplotypes of varying sizes were examined. The haplotype blocks had average physical sizes ranging from 1.02 kbp to 47453.13 kbp, 379625.06 kbp, 1073741.82 kbp, and 47453.13 kbp, respectively, for canola, maize, wheat and soybean. A summary of the average size distributions can be found in <xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>. Notably, the fixed window approaches allowed for the construction of both the smallest haplotype blocks (1.02 kbp) and the largest haplotype blocks (<xref ref-type="table" rid="T1">
<bold>Table&#xa0;1</bold>
</xref>).</p>
<table-wrap id="T1" position="float">
<label>Table&#xa0;1</label>
<caption>
<p>Average size ranges of haplotype blocks constructed by LD, fixed window of adjacent markers and fixed window of adjacent base pairs in the canola, maize, wheat and soybean dataset.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Dataset</th>
<th valign="middle" align="center">Method</th>
<th valign="middle" align="center">minimal average size (kbp)</th>
<th valign="middle" align="center">maximal average size (kbp)</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="3" align="center">Canola</td>
<td valign="middle" align="center">LD</td>
<td valign="middle" align="center">97.49 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1)</td>
<td valign="middle" align="center">2629.87 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.01)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent marker</td>
<td valign="middle" align="center">25.46 (<italic>nSNP</italic> = 2)</td>
<td valign="middle" align="center">39801.09 (<italic>nSNP</italic> = 2048)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent base pairs</td>
<td valign="middle" align="center">1.02 kbp</td>
<td valign="middle" align="center">47453.13 kbp</td>
</tr>
<tr>
<td valign="middle" rowspan="3" align="center">Maize</td>
<td valign="middle" align="center">LD</td>
<td valign="middle" align="center">8.08 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1)</td>
<td valign="middle" align="center">21556.53 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.01)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent marker</td>
<td valign="middle" align="center">64.31 (<italic>nSNP</italic> = 2)</td>
<td valign="middle" align="center">205312.88 (<italic>nSNP</italic> = 5793)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent base pairs</td>
<td valign="middle" align="center">1.02 kbp</td>
<td valign="middle" align="center">379625.06 kbp</td>
</tr>
<tr>
<td valign="middle" rowspan="3" align="center">Wheat</td>
<td valign="middle" align="center">LD</td>
<td valign="middle" align="center">106.79 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1)</td>
<td valign="middle" align="center">64954.10 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.01)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent marker</td>
<td valign="middle" align="center">1544.58 (<italic>nSNP</italic> = 2)</td>
<td valign="middle" align="center">667692.8 (<italic>nSNP</italic> = 1024)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent base pairs</td>
<td valign="middle" align="center">1.02 kbp</td>
<td valign="middle" align="center">1073741.82 kbp</td>
</tr>
<tr>
<td valign="middle" rowspan="3" align="center">Soybean</td>
<td valign="middle" align="center">LD</td>
<td valign="middle" align="center">138.55 bp (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1)</td>
<td valign="middle" align="center">1587.07 (<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.01)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent marker</td>
<td valign="middle" align="center">430.27 (<italic>nSNP</italic> = 2)</td>
<td valign="middle" align="center">1526.61 (<italic>nSNP</italic> = 2897)</td>
</tr>
<tr>
<td valign="middle" align="center">fixed window of adjacent base pairs</td>
<td valign="middle" align="center">1.02 kbp</td>
<td valign="middle" align="center">47453.13 kbp</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Within all datasets using the methods implemented in <italic>&#x201c;Haploview&#x201d;</italic> (GAB, GAM, SPI), the number of haplotype blocks was consistently lower than the number of total SNPs (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1A, E, I, M</bold>
</xref>). However, a significant portion of those blocks consisted of only a single SNP (unblocked SNPs) (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1A, E, I, M</bold>
</xref>). Moreover, the total number of haplotypes available for genomic prediction (excluding single SNP blocks) increased in the canola and soybean datasets, remained similar to the number of SNPs in wheat, and decreased in maize (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1A, E, I, M</bold>
</xref>). Across all datasets, the number of blocks based on LD increased with higher LD thresholds. Additionally, in the case of maize, the number of haplotypes exhibited a similar pattern. With LD-based haplotype blocks, the number of haplotypes (excluding single SNP blocks) exceeded the total number of SNPs across all LD thresholds in soybean and was lower across all thresholds in maize (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1B, F, J, N</bold>
</xref>). In canola, thresholds above <italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.75 resulted in fewer haplotypes than SNPs, while lower thresholds yielded higher numbers. Conversely, in wheat, only relatively small blocks (<italic>r<sup>2</sup>
</italic> &#x2264; 0.10) increased the number of haplotypes compared to the number of SNPs (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1B</bold>
</xref>). With fixed window blocks, the number of haplotype blocks generally decreased with increasing block size (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1C, D, G, H, K, L, O, P</bold>
</xref>). Here, the number of haplotypes was the highest with relatively small blocks, with increasing block size, the number of haplotypes decreased (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1C, D, G, H, K, L, O, P</bold>
</xref>). Notably, in comparison to SNPs, the number of haplotypes was higher for blocks smaller than 1024, 6, 128, and 1449 SNPs, or 23726.57 kbp, 92.68 kbp, 134217.73 kbp, and 33554.43 kbp in the canola, maize, wheat, and soybean datasets, respectively (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1C, D, G, H, K, L, O, P</bold>
</xref>). In all scenarios, increasing block size resulted in fewer unblocked markers, especially with the fixed window approaches. In all datasets, the <italic>&#x201c;HaploBlocker&#x201d;</italic> method produced the fewest haplotypes, considerably fewer than the number of SNPs (<xref ref-type="fig" rid="f1">
<bold>Figures&#xa0;1A, E, I, M</bold>
</xref>). Furthermore, across all datasets and methods, except for blocks based on <italic>&#x201c;HaploBlocker&#x201d;</italic>, most of the introduced haplotypes can be classified as rare (Frequency &#x2264; 0.05) or very rare (Frequency &#x2264; 0.01) (<xref ref-type="fig" rid="f1">
<bold>Figure&#xa0;1</bold>
</xref>).</p>
<fig id="f1" position="float">
<label>Figure&#xa0;1</label>
<caption>
<p>Numbers of SNPs (dark green horizontal line), haplotype blocks (dark orange), haplotypes (light orange), haplotypes with frequency &#x2264; 0.05 (light green), haplotypes with frequency &#x2264; 0.01 (light yellow) and unblocked SNP markers (green) identified by GAB, GAM, SPI, <italic>&#x201c;HaploBlocker&#x201d;</italic> <bold>(A, E, I, M)</bold>, LD <bold>(B, F, J, N)</bold>, fixed window of adjacent markers <bold>(C, G, K, O)</bold> and fixed window of adjacent base pairs <bold>(D, H, I, P)</bold> in canola <bold>(A&#x2013;D)</bold>, maize <bold>(E&#x2013;H)</bold>, wheat <bold>(I&#x2013;L)</bold> and soybean <bold>(M&#x2013;P)</bold>.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1217589-g001.tif"/>
</fig>
<p>Across the four datasets, the examination of the correlations between relationship coefficients derived from SNPs and haplotypes revealed high redundancy between the two marker types in many method/parameter combinations. The methods implemented in <italic>&#x201c;Haploview&#x201d;</italic> resulted in relationship coefficients that were highly correlated to those obtained from SNPs, closely approaching a correlation coefficient of 1, in canola, wheat, and soybean (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1A, E, I, M</bold>
</xref>). However, in maize, these methods only produced intermediate correlations (GAB = 0.60, GAM = 0.50, SPI = 0.46) (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S1E</bold>
</xref>). In all datasets, relationship coefficients from haplotypes from LD-based haplotype blocks were highly correlated to those obtained from SNPs (<italic>r</italic> &gt; 0.75) with little variation observed across LD thresholds. Only at very low LD thresholds, this correlation was slightly lower, while it was slightly higher for very high thresholds (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1B, F, J, N</bold>
</xref>). Additionally, small fixed window blocks resulted in relationship coefficients similar to those obtained from SNPs, closely approaching a correlation coefficient of 1. However, this similarity eroded drastically with increasing block size (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1C, D, G, H, K, L, O, P</bold>
</xref>). Notably, in Soybean, while the correlation between relationship coefficients from SNPs and haplotypes decreased with increasing block size of the fixed window of adjacent base pairs, it slightly increased again with the largest blocks (<italic>nKB</italic> = 67108.86) (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S1P</bold>
</xref>). In canola and soybean, relationship coefficients obtained from <italic>&#x201c;HaploBlocker&#x201d;</italic> were highly correlated to those obtained from SNPs (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1A, M</bold>
</xref>). In wheat, this correlation was lower (r = 0.75), and in maize, it was close to zero (r = 0.058), indicating that these blocks capture different information (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S1E, I</bold>
</xref>).</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Genomic prediction</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Canola</title>
<p>Within the canola dataset, the prediction accuracy across different models ranged from 0.3 to 0.85, with a strong dependence on the specific trait. Notably, for oil yield, field emergence, glucosinolate content, and lodging, the models considering epistatic effects (EGBLUP and RKHS) consistently outperformed by the other SNP-based models (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S2</bold>
</xref>). However, this effect did not consistently translate to haplotype-based predictions. Prediction accuracy showed little variation across LD threshold as well as between LD base, <italic>&#x201c;Haploview&#x201d;</italic> or <italic>&#x201c;HaploBlocker&#x201d;</italic> methods (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2A, B</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S2</bold>
</xref>). On the other hand, the fixed-window approaches exhibited the most variation, with a substantial decrease in prediction accuracy as the block size increased for every trait, while small blocks based on fixed windows resulted in prediction accuracies similar to those based on SNPs or the remaining methods (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2C, D</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S2</bold>
</xref>).</p>
<fig id="f2" position="float">
<label>Figure&#xa0;2</label>
<caption>
<p>Prediction accuracy (r) of GBLUP (red), Bayesian LASSO (blue), EGBLUP (black) and RKHS (grey) with SNPs, GAB, GAM, SPI, <italic>&#x201c;HaploBlocker&#x201d;</italic> <bold>(A, E, I, M)</bold>, LD <bold>(B, F, J, N)</bold>, fixed window of adjacent markers <bold>(C, G, K, O)</bold> and fixed window of adjacent base pairs <bold>(D, H, L, P)</bold> based haplotype blocks, in canola seed yield <bold>(A&#x2013;D)</bold>, maize DMY <bold>(E&#x2013;H)</bold>, wheat seed yield <bold>(I&#x2013;L)</bold> and soybean oil content <bold>(M&#x2013;P)</bold>. Individual points in the line plots represent the mean over all cross validation runs for each haplotype block parameter and model combination.</p>
</caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="fpls-14-1217589-g002.tif"/>
</fig>
<p>Comparing haplotype blocks to SNP-based prediction, the improvement in prediction accuracy ranged from 0.007 to 0.021 for GBLUP, 0.008 to 0.024 for Bayesian LASSO, 0.008 to 0.023 for EGBLUP, and 0.007 to 0.022 for RKHS. These values were based on the haplotyping method that yielded the highest prediction accuracy for each specific trait and model (<xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S1</bold>
</xref>). Interestingly, the use of haplotypes seemed to have the least impact on oil yield (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S1</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>). Except for flowering time with RKHS, the LD-based methods generally resulted in the most significant improvements. However, no ideal LD threshold or range of thresholds could be identified (<xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>). In the case of flowering time with RKHS, the optimal haplotyping method involved a fixed window of adjacent base pairs measuring 20987.15 kbp.</p>
<table-wrap id="T2" position="float">
<label>Table&#xa0;2</label>
<caption>
<p>Average prediction accuracy of SNP based prediction compared to the best haplotyping method of canola, maize, wheat and soybean for some example traits.</p>
</caption>
<table frame="hsides">
<thead>
<tr>
<th valign="middle" align="center">Dataset</th>
<th valign="middle" align="center">Trait</th>
<th valign="middle" align="center">Model</th>
<th valign="middle" align="center">SNP prediction accuracy</th>
<th valign="middle" align="center">Best haplotyping algorithm</th>
<th valign="middle" align="center">Prediction accuracy by best haplotyping algorithm</th>
<th valign="middle" align="center">Improvement by best haplotyping algorithm</th>
</tr>
</thead>
<tbody>
<tr>
<td valign="middle" rowspan="8" align="center">Canola</td>
<td valign="middle" rowspan="4" align="center">yield</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.464</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.6</td>
<td valign="middle" align="center">0.485</td>
<td valign="middle" align="center">0.021</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.462</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.59</td>
<td valign="middle" align="center">0.486</td>
<td valign="middle" align="center">0.024</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.471</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.6</td>
<td valign="middle" align="center">0.492</td>
<td valign="middle" align="center">0.021</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.474</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.71</td>
<td valign="middle" align="center">0.496</td>
<td valign="middle" align="center">0.022</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">flowering time</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.709</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.12</td>
<td valign="middle" align="center">0.721</td>
<td valign="middle" align="center">0.012</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.697</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.15</td>
<td valign="middle" align="center">0.721</td>
<td valign="middle" align="center">0.024</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.711</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.14</td>
<td valign="middle" align="center">0.723</td>
<td valign="middle" align="center">0.012</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.704</td>
<td valign="middle" align="center">
<italic>nKB</italic> = 2097.15</td>
<td valign="middle" align="center">0.719</td>
<td valign="middle" align="center">0.015</td>
</tr>
<tr>
<td valign="middle" rowspan="8" align="center">Maize</td>
<td valign="middle" rowspan="4" align="center">DMY</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.624</td>
<td valign="middle" align="center">GAB</td>
<td valign="middle" align="center">0.635</td>
<td valign="middle" align="center">0.011</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.616</td>
<td valign="middle" align="center">GAB</td>
<td valign="middle" align="center">0.621</td>
<td valign="middle" align="center">0.006</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.620</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 8</td>
<td valign="middle" align="center">0.622</td>
<td valign="middle" align="center">0.002</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.608</td>
<td valign="middle" align="center">GAB</td>
<td valign="middle" align="center">0.631</td>
<td valign="middle" align="center">0.023</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">DtTAS</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.847</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.847</td>
<td valign="middle" align="center">0.000</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.846</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1</td>
<td valign="middle" align="center">0.842</td>
<td valign="middle" align="center">-0.004</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.845</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.846</td>
<td valign="middle" align="center">0.000</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.846</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1</td>
<td valign="middle" align="center">0.842</td>
<td valign="middle" align="center">-0.003</td>
</tr>
<tr>
<td valign="middle" rowspan="8" align="center">Wheat</td>
<td valign="middle" rowspan="4" align="center">yield</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.805</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.23</td>
<td valign="middle" align="center">0.813</td>
<td valign="middle" align="center">0.008</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.697</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 46</td>
<td valign="middle" align="center">0.818</td>
<td valign="middle" align="center">0.122</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.811</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.1</td>
<td valign="middle" align="center">0.815</td>
<td valign="middle" align="center">0.005</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.765</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.1</td>
<td valign="middle" align="center">0.814</td>
<td valign="middle" align="center">0.049</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">sedimentation value</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.493</td>
<td valign="middle" align="center">
<italic>nKB</italic> = 1073741.82</td>
<td valign="middle" align="center">0.619</td>
<td valign="middle" align="center">0.126</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.488</td>
<td valign="middle" align="center">
<italic>nKB</italic> = 1073741.82</td>
<td valign="middle" align="center">0.636</td>
<td valign="middle" align="center">0.148</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.620</td>
<td valign="middle" align="center">
<italic>nKB</italic> = 1073741.82</td>
<td valign="middle" align="center">0.627</td>
<td valign="middle" align="center">0.006</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.610</td>
<td valign="middle" align="center">
<italic>nKB</italic> = 1073741.82</td>
<td valign="middle" align="center">0.631</td>
<td valign="middle" align="center">0.021</td>
</tr>
<tr>
<td valign="middle" rowspan="8" align="center">Soybean</td>
<td valign="middle" rowspan="4" align="center">oil content</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.674</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.24</td>
<td valign="middle" align="center">0.682</td>
<td valign="middle" align="center">0.008</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.675</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.24</td>
<td valign="middle" align="center">0.683</td>
<td valign="middle" align="center">0.008</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.674</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.24</td>
<td valign="middle" align="center">0.682</td>
<td valign="middle" align="center">0.008</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.677</td>
<td valign="middle" align="center">
<italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.26</td>
<td valign="middle" align="center">0.691</td>
<td valign="middle" align="center">0.014</td>
</tr>
<tr>
<td valign="middle" rowspan="4" align="center">protein content</td>
<td valign="middle" align="center">GBLUP</td>
<td valign="middle" align="center">0.601</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.606</td>
<td valign="middle" align="center">0.006</td>
</tr>
<tr>
<td valign="middle" align="center">Bayesian LASSO</td>
<td valign="middle" align="center">0.602</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.608</td>
<td valign="middle" align="center">0.006</td>
</tr>
<tr>
<td valign="middle" align="center">EGBLUP</td>
<td valign="middle" align="center">0.609</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.611</td>
<td valign="middle" align="center">0.003</td>
</tr>
<tr>
<td valign="middle" align="center">RKHS</td>
<td valign="middle" align="center">0.609</td>
<td valign="middle" align="center">
<italic>nSNP</italic> = 4</td>
<td valign="middle" align="center">0.613</td>
<td valign="middle" align="center">0.003</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Maize</title>
<p>Prediction accuracy obtained from the random cross validation ranged from 0.4 to 0.9 and was trait-dependent. Here, little difference between models was observed with SNP-based prediction (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2E</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S3</bold>
</xref>). With haplotypes, however, there were considerable differences between Models implemented in a Bayesian framework and frequentist models (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S3</bold>
</xref>). With haplotypes based on LD, prediction accuracy decreased with higher LD thresholds for GBLUP and EGBLUP and increased for Bayesian LASSO and RKHS, respectively (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2F</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S4</bold>
</xref>). Here for DMY, DMC and PH, respectively, all models approached a similar prediction accuracy around <italic>r<sup>2</sup>
</italic>~0.75 (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S3</bold>
</xref>). And for DtTAS and DtSILK all models approached the same prediction accuracy around <italic>r<sup>2</sup>
</italic>~0.55 (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S3</bold>
</xref>). The same behavior could not be observed with the fixed window haplotypes, where prediction accuracy obtained from GBLUP and EGBLUP decreased drastically with increasing block size. Here, for models implemented in a Bayesian framework the prediction accuracy remained low independent of the block size (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2G, H</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S3</bold>
</xref>). Except for DMY, where the GAB method slightly improved prediction accuracy, haplotypes based on the algorithms implemented in &#x201c;<italic>Haploview&#x201d;</italic> decreased prediction accuracy in every scenario (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2E</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S4</bold>
</xref>). In general, there was no discernable improvement of prediction accuracy by haplotypes compared to SNP-based predictions. In all traits but DMY, the haplotyping method with the highest prediction accuracy even decreased prediction accuracy with Bayesian LASSO and RKHS (<xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S1</bold>
</xref>), whereas for GBLUP and EGBLUP prediction accuracy did not or only slightly increased prediction accuracy compared to SNP based prediction. DMY profited most from haplotypes, whereby GBLUP, Bayesian LASSO and RKHS worked best with the GAB method while for EGBLUP a fixed window of 8 SNPs was ideal (<xref ref-type="fig" rid="f2">
<bold>Figure&#xa0;2</bold>
</xref>; <xref ref-type="table" rid="T2">
<bold>Table&#xa0;2</bold>
</xref>). Besides DMY, for Bayesian LASSO and RKHS haplotypes worked best with an LD threshold of <italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>1, however, prediction accuracy was still worse than SNP based prediction (<xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S4</bold>
</xref>). For the same traits, Frequentists model worked best with varying fixed window size haplotypes, with a maximal improvement of 0.002 (<xref ref-type="supplementary-material" rid="SM1">
<bold>Table S1</bold>
</xref>). The <italic>&#x201c;HaploBlocker&#x201d;</italic> method together with very large fixed window blocks yielded the lowest prediction accuracies across all traits.</p>
<p>The family-wise cross validation generally yielded considerably lower prediction accuracies than its random counterpart (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S4</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM1">
<bold>Table S2</bold>
</xref>). The ranking in prediction accuracies obtained from haplotype blocks followed the pattern of the random counterpart, albeit being lower (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figure S4</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM3">
<bold>Table S2</bold>
</xref>). Mentionable, prediction accuracy approached zero for the <italic>HaploBlocker&#x201d;</italic> method together with very large fixed window blocks.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Wheat</title>
<p>Prediction accuracy in the wheat dataset exhibited much greater variability between traits compared to the other three datasets, ranging from -0.4 to 0.9, depending on the specific trait. Interestingly, even with SNP-based predictions, considerable differences in prediction accuracy were observed across (i) models that consider epistasis and those that do not, (ii) frequentist and models implemented in a Bayesian framework, and (iii) combinations of (i) and (ii) (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S5</bold>
</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">
<bold>S7</bold>
</xref>). However, when haplotype blocks were utilized, all models achieved at least the average prediction accuracy of the best SNP-based model for 13 out of 15 traits (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S5</bold>
</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">
<bold>S7</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>). This was achieved by using haplotype blocks constructed with varying methods, including even the largest possible haplotype blocks based on fixed windows (e.g., using whole chromosomes as blocks) (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S5</bold>
</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">
<bold>S7</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>).</p>
<p>Furthermore, for traits such as yield, biomass yield, NUE, protein yield, sedimentation value, stripe rust, and falling number, the previously worst-performing SNP-based model became the best-performing model when using haplotype blocks (<xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>). Additionally, for traits with very low or even negative prediction accuracy based on SNPs (e.g., plant height, TKW, days till heading, falling number, powdery mildew, and stripe rust), strong improvements were achieved through the use of haplotypes (<xref ref-type="supplementary-material" rid="SM1">
<bold>Figures S5</bold>
</xref>&#x2013;<xref ref-type="supplementary-material" rid="SM1">
<bold>S7</bold>
</xref>; <xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>). Models implemented in a Bayesian framework seemed to benefit the most from the utilization of haplotypes, with changes in prediction accuracy ranging from -0.039 to 0.170 for GBLUP, from 0.006 to 0.277 for Bayesian LASSO, from -0.003 to 0.085 for EGBLUP, and from 0.025 to 0.291 for RKHS (<xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S1</bold>
</xref>). The most notable improvements were typically seen when prediction accuracy varied considerably between models using SNP data. Only for cases, such as falling number with RKHS, kernel spike<sup>-1</sup> with EGBLUP, spike m<sup>-2</sup> with RKHS and EGBLUP, and stripe rust resistance with GBLUP, did the prediction accuracy decrease compared to SNP-based prediction when using haplotype blocks (<xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>).</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Soybean</title>
<p>The prediction accuracy in the soybean dataset ranged from 0.5 to 0.8 and exhibited a striking similarity between oil content and protein content. No noticeable differences were observed between models based on SNPs (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2M</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S8A, E</bold>
</xref>). Moreover, there was minimal variation in prediction accuracy across different LD thresholds, with only a slight decrease in accuracy observed between <italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.01 and 0.05 (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2N</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S8B, F</bold>
</xref>).</p>
<p>When using fixed windows of adjacent marker blocks, the prediction accuracy experienced a decline with increasing block size for all models (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2O</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S8C, G</bold>
</xref>). Similar behavior was observed for fixed windows of adjacent base pairs blocks, except for a marginal increase in prediction accuracy with blocks of size 47453.13 kbp (<xref ref-type="fig" rid="f2">
<bold>Figures&#xa0;2P</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM1">
<bold>S8D, H</bold>
</xref>). However, it is worth noting that the prediction accuracy remained lower than the SNP-based prediction in that case. Overall, the improvements achieved with haplotypes were relatively minor (<xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM2">
<bold>S1</bold>
</xref>). For oil and protein content, the best haplotype block method and parameter improved the prediction accuracy by only 0.006 and 0.008 with GBLUP, 0.006 and 0.009 with Bayesian LASSO, 0.003 and 0.008 with EGBLUP, and 0.003 and 0.014 with RKHS, respectively, compared to the SNP-based prediction (<xref ref-type="table" rid="T2">
<bold>Tables&#xa0;2</bold>
</xref>, <xref ref-type="supplementary-material" rid="SM3">
<bold>S1</bold>
</xref>).</p>
<p>Interestingly, within the traits, it was observed that the models worked best with the same haplotype block method: an LD threshold of <italic>r<sup>2</sup>&#xa0;=&#xa0;</italic>0.24-0.26 for oil content and a fixed window size of <italic>nSNP</italic> = 4 for protein content.</p>
</sec>
</sec>
</sec>
<sec id="s4" sec-type="discussion">
<label>4</label>
<title>Discussion</title>
<p>Using datasets from four diverse crops and haplotype blocks constructed using a broad range of construction parameters, we show how haplotype blocks change in size and influence the effective number of predictors for genomic prediction. While haplotype blocks sometimes drastically change the number of predictors, genomic prediction accuracy was only marginally affected with no consistent improvement for any method and trait.</p>
<p>Haplotype blocks were built based on LD (<italic>r<sup>2</sup>
</italic>), fixed window sizes of adjacent marker or base pairs as well as the three algorithms implemented in the software &#x201c;<italic>Haploview&#x201d;</italic> and the method <italic>&#x201c;Haploblocker&#x201d;</italic>. The <italic>r<sup>2</sup>
</italic> measurement of LD between markers (<xref ref-type="bibr" rid="B48">Hill and Robertson, 1968</xref>; <xref ref-type="bibr" rid="B47">Hill, 1981</xref>) is highly correlated to <italic>D&#xb4;</italic> (<xref ref-type="bibr" rid="B96">VanLiere and Rosenberg, 2008</xref>), which is more commonly used in tagSNP methods where it showed superior performance to other measures (<xref ref-type="bibr" rid="B12">Carlson et&#xa0;al., 2004</xref>; <xref ref-type="bibr" rid="B26">de Bakker et&#xa0;al., 2005</xref>). According to <xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al. (2014)</xref>, <italic>r<sup>2</sup>
</italic> and <italic>D&#xb4;</italic> show no difference in terms of prediction accuracy in genomic prediction. The high resolution of haplotype blocking methods and construction parameters allowed an examination of a wide range of haplotype block sizes that are normally not considered in genomic prediction. Most studies in this regard only include single or few construction methods or parameters (<xref ref-type="bibr" rid="B68">Lorenz et&#xa0;al., 2010</xref>; <xref ref-type="bibr" rid="B2">Ballesta et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B71">Maldonado et&#xa0;al., 2019</xref>), although our results show that the method of haplotype construction can potentially impact prediction quality. We included haplotype blocks of relatively large sizes, such as a LD threshold of 0.01 and whole chromosome blocks, which may initially seem unrealistic. However, we included these large blocks to account for scenarios in which traits are controlled by large chromosome segments (<xref ref-type="bibr" rid="B102">Voss-Fels et&#xa0;al., 2019</xref>), possibly resulting from introgression breeding with suppressed recombination (<xref ref-type="bibr" rid="B39">Hao et&#xa0;al., 2020</xref>).</p>
<p>Here, in three datasets the number of haplotypes could be increased substantially compared to the number of SNPs. The number of haplotypes we observed in the four examined datasets was lower than observed in cattle (<xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B23">Cuyabano et&#xa0;al., 2015</xref>; <xref ref-type="bibr" rid="B65">Li et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B64">Li et&#xa0;al., 2022</xref>) and human (<xref ref-type="bibr" rid="B66">Liang et&#xa0;al., 2020</xref>) but similar to previous reports in plants including <italic>Eucalyptus globulus</italic> (<xref ref-type="bibr" rid="B2">Ballesta et&#xa0;al., 2019</xref>), maize (<xref ref-type="bibr" rid="B74">Matias et&#xa0;al., 2017</xref>) and rice (<xref ref-type="bibr" rid="B74">Matias et&#xa0;al., 2017</xref>). These variations may arise from differences in population diversity, marker density, and sequencing technology. The haplotype number detected in maize by <xref ref-type="bibr" rid="B74">Matias et&#xa0;al. (2017)</xref> was comparable to that observed in our analysis using around ten times fewer SNP markers, indicating that haplotype number is not (solely) dependent on marker density. However, as expected there is a relationship between the population size and haplotype number, with more (diverse) genotypes causing more haplotypes. The number of haplotypes we detected corresponded to the population size used for each crop, with wheat having the fewest haplotypes and soybean the most, independent of the method. Nevertheless, an effect of genetic diversity within a species or population cannot be discounted without comparative within-species analyses of alternative populations. Some authors argue that use of haplotype blocks can help to reduce dimensionality (<xref ref-type="bibr" rid="B57">Kim et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B83">Pook et&#xa0;al., 2019</xref>). However, depending on the methods and parameters for haplotype construction the number of haplotypes was sometimes higher in the examined datasets than the number of SNPs. This may reflect lower marker numbers and different methods compared to <xref ref-type="bibr" rid="B57">Kim et&#xa0;al. (2019)</xref>. Dimensionality can certainly be decreased if rare haplotypes would be excluded (<xref ref-type="bibr" rid="B44">Hess et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B64">Li et&#xa0;al., 2022</xref>). The method <italic>&#x201c;HaploBlocker&#x201d;</italic> described by <xref ref-type="bibr" rid="B83">Pook et&#xa0;al. (2019)</xref> decreased the dimensionality in every examined dataset. In all cases, the major drawback of the large number of additional variants is the very low frequency at which the haplotypes occur. However, low frequency variants are often assumed to be in higher LD with recent causal mutations (<xref ref-type="bibr" rid="B9">Bloom et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B104">Wainschtein et&#xa0;al., 2022</xref>), implying that their detection and use for predictions could be beneficial. However, caution is needed when considering all haplotypes, especially rare ones. In genomic predictions. effect estimation of rare variants require large populations to be estimated accurately (<xref ref-type="bibr" rid="B75">Meuwissen et&#xa0;al., 2001</xref>; <xref ref-type="bibr" rid="B35">Goddard and Hayes, 2007</xref>). In large populations, rare variants can be observed at higher frequencies which enables a more accurate estimation of their trait effects. In SNP based prediction markers are commonly excluded if they have a minor allele frequency &#x2264; 0.05 (<xref ref-type="bibr" rid="B93">Technow et&#xa0;al., 2012</xref>; <xref ref-type="bibr" rid="B19">Crossa et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B51">Jan et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B108">Werner et&#xa0;al., 2018a</xref>; <xref ref-type="bibr" rid="B116">Zhang et&#xa0;al., 2018</xref>). With large populations, filtering could be shifted from frequencies to allele counts, potentially leading to more reliable effect estimates of rare haplotypes. However, increasing the population could again increases the number of rare new haplotypes. In all four datasets, the number of unblocked SNPs decreased with increasing block size. With LD based haplotype blocks, increasing the LD threshold resulted in an increase of unblocked SNPs.</p>
<p>Genomic prediction was conducted using four models: GBLUP, EGBLUP, Bayesian LASSO, and RKHS regression, with the latter two implemented within a Bayesian framework. GBLUP, being the golds standard of genomic prediction, is a widely employed prediction models in breeding, hence we included it in the analysis. However, GBLUP assumes that all markers or haplotypes contribute to the trait (through relationship), prompting the inclusion of Bayesian LASSO, which allows for marker or haplotype-specific shrinkage of effects towards zero. This is beneficial in scenarios where not all markers or haplotypes have an impact on the trait. Given the assumption that haplotypes capture local epistatic effects (<xref ref-type="bibr" rid="B56">Jiang et&#xa0;al., 2018</xref>), EGBLUP and RKHS regression were employed to assess whether considering global epistasis between haplotype blocks could yield a substantial improvement in genomic prediction. Although haplotype blocks are typically fewer in number compared to SNPs, the number of haplotypes used for prediction was often comparable to or even greater than the number of SNPs. Therefore, we selected prediction models capable of handling the challenges posed by the large p small n scenario, opting not to explore machine learning models. Furthermore, the application of machine learning methods would have required extensive hyperparameter optimization, which would have significantly exceeded the computational time required for the four prediction models employed in this study. Lastly, the objective of this study was to compare various haplotype blocking methods and parameters, rather than comparing different prediction models.</p>
<p>Generally, genomic prediction accuracies based on SNPs were similar to those reported in the literature across all datasets. In the canola dataset, accuracies closely matched <xref ref-type="bibr" rid="B51">Jan et&#xa0;al. (2016)</xref>, with a small improvement likely due to the higher number of markers remaining after filtering. Trait prediction accuracies in canola/rapeseed were mostly consistent with previous reports, with minor variations observed for field emergence, and glucosinolate content (<xref ref-type="bibr" rid="B112">W&#xfc;rschum et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B51">Jan et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B108">Werner et&#xa0;al., 2018a</xref>; <xref ref-type="bibr" rid="B109">Werner et&#xa0;al., 2018b</xref>). Also in the maize dataset, SNP-based genomic prediction accuracy roughly matched the original publication (<xref ref-type="bibr" rid="B61">Lehermeier et&#xa0;al., 2014</xref>), with expected differences due to varying cross-validation schemes. Maize hybrids exhibited high prediction accuracies as previously reported (<xref ref-type="bibr" rid="B93">Technow et&#xa0;al., 2012</xref>; <xref ref-type="bibr" rid="B21">Crossa et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B76">Millet et&#xa0;al., 2019</xref>). In wheat, prediction accuracies based on SNPs for seed yield and yield components were on a very high level (<xref ref-type="supplementary-material" rid="SM2">
<bold>Table S1</bold>
</xref>) compared to many previously published reports (<xref ref-type="bibr" rid="B58">Lado et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B118">Zhao et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B21">Crossa et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B24">Daetwyler et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B20">Crossa et&#xa0;al., 2016</xref>; <xref ref-type="bibr" rid="B32">Edwards et&#xa0;al., 2019</xref>). Furthermore, prediction accuracies based on SNPs for stripe rust resistance, despite population differences, showed a similar level than observed by <xref ref-type="bibr" rid="B24">Daetwyler et&#xa0;al. (2014)</xref>. Whereas protein content had a higher prediction accuracy compared to <xref ref-type="bibr" rid="B20">Crossa et&#xa0;al. (2016)</xref>, sedimentation value was predicted equally well. In soybean, prediction accuracies based solely on SNPs were comparable to levels reported by <xref ref-type="bibr" rid="B53">Jarquin et&#xa0;al. (2016)</xref> for oil content and protein content, despite considerable differences in the cross-validation and modeling schemes. The lack of differences in prediction accuracies may be explained by the narrow genetic diversity in soybean breeding material due to genetic bottlenecks (<xref ref-type="bibr" rid="B50">Hyten et&#xa0;al., 2006</xref>).</p>
<p>Genomic prediction with LD-based haplotype blocks in canola resulted in the highest accuracy improvements for most model/trait combinations. Variation in prediction accuracy across LD thresholds was minimal. The optimal threshold varied significantly by trait and model, ranging from very low (0.01) to high (0.89). In wheat, LD-based haplotype blocks were superior to the other haplotyping methods for 20 out of 60 model/trait combinations, but accuracy didn&#x2019;t always improve compared to SNP-based prediction. Similar low variation across LD thresholds was observed in soybean. For soybean&#x2019;s oil content, the ideal LD threshold for accuracy estimates across all models was 0.24-0.26. In maize, only the Bayesian LASSO and RKHS models achieved the highest improvements with LD based haplotype blocks with a threshold of 1, effectively removing redundant information. In this scenario, only markers in complete LD were grouped into a block, effectively removing redundant information. This process, is similar to LD pruning, which has been demonstrated to enhance prediction accuracy (<xref ref-type="bibr" rid="B113">Ye et&#xa0;al., 2019</xref>). Intriguing patterns were observed with LD-based haplotypes in maize, the two models implemented in a Bayesian framework (Bayesian LASSO and RKHS) behaved in an opposite direction to the other (frequentist) models, potentially due to different estimation procedures. In contrast to <xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al. (2014)</xref>, we generally did not find an ideal LD threshold or even an ideal threshold specific to each dataset and mostly not even an ideal threshold within one trait. The prediction accuracy variation along LD thresholds reported in cattle (<xref ref-type="bibr" rid="B22">Cuyabano et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B65">Li et&#xa0;al., 2021</xref>; <xref ref-type="bibr" rid="B64">Li et&#xa0;al., 2022</xref>) was similar to the variation observed in our analyses. This suggests that any LD threshold is reasonable for genomic prediction due to low variation of prediction accuracy. We propose that even with extreme LD thresholds, reasonably accurate haplotype blocks are constructed, which explains the low variation observed across LD thresholds in all datasets. Additionally, in all datasets, the correlation between relationship coefficients obtained from markers and haplotypes was consistently high, with little variation across LD thresholds. This suggests that relationship representation remains consistent when using LD-based haplotype blocks.</p>
<p>The use of small fixed window blocks led to prediction accuracies comparable to those achieved with individual SNPs. Additionally, in maize, our findings aligned with those of <xref ref-type="bibr" rid="B56">Jiang et&#xa0;al. (2018)</xref> in Flint material, showing similar prediction accuracy patterns for frequentist models using small fixed window size haplotype blocks (2-5 markers). Interestingly, in maize, prediction accuracy eroded with the two frequentist models and increasing block size based on fixed windows, whereas for two Bayesian models the prediction accuracy was low across all parameters. Except for the wheat dataset, using excessively large fixed windows to build haplotype blocks considerably reduced prediction accuracy, as observed in previous studies with cattle (<xref ref-type="bibr" rid="B44">Hess et&#xa0;al., 2017</xref>). Unrealistically large blocks likely obscure the effects of true QTL within them. Furthermore, these larger blocks are generally more prone to errors in genotyping, and imputation, which accumulate in large blocks and limit prediction accuracy of genomic prediction models utilizing these blocks. These errors can also introduce false rare haplotypes, exacerbating issues related to rare variants. Additionally, as block size increases, haplotypes become more specific to genotypes or subpopulations, resulting in the absence of certain haplotypes in the training set but presence in the validation set. This lack of overlap leads to inaccurate estimation of the effects for those haplotypes, thus decreasing prediction accuracy due to the limited shared haplotypes between the training and validation sets. In the case of wheat, however, using very large blocks, such as whole chromosomes, resulted in considerable improvements in prediction accuracy. Mentionable improvements were observed for traits such as wheat stripe rust resistance, powdery mildew resistance, and kernel spike<sup>-1</sup>. This improvement can likely be attributed to introgression breeding in wheat, where large chromosome segments are introgressed and preserved due to restricted recombination (<xref ref-type="bibr" rid="B39">Hao et&#xa0;al., 2020</xref>). Furthermore, the wheat D-subgenome exhibits large LD haplotype blocks that are important for yield and biomass-related traits (<xref ref-type="bibr" rid="B102">Voss-Fels et&#xa0;al., 2019</xref>). However, it should be noted that these improvements were observed in cases where the model performance was initially at a very low level with SNPs. The correlation between relationship coefficients obtained from markers and haplotypes was high for small fixed window blocks but decreased as block size increased. This suggests that crucial relationship information is lost or encoded within large haplotype blocks, which cannot be accessed for accurate prediction. As a result, the prediction accuracy in canola, maize, and soybean is reduced. However, it is important to highlight that large blocks can potentially introduce additional trait information, as demonstrated by their impact in some of the wheat traits.</p>
<p>The widely used algorithms implemented in <italic>&#x201c;Haploview&#x201d;</italic> did not exhibit superiority in terms of prediction accuracy compared to other methods. Although the method proposed by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref> showed a slight improvement, particularly in maize DMY, these gains remained modest when compared to SNP-based prediction. In contrast to the findings of <xref ref-type="bibr" rid="B74">Matias et&#xa0;al. (2017)</xref>, our analysis generally revealed a decrease in prediction accuracy rather than a benefit from haplotypes based on <italic>&#x201c;Haploview&#x201d;</italic> in the maize dataset. This discrepancy could be attributed to differences in the plant materials studied. While <xref ref-type="bibr" rid="B74">Matias et&#xa0;al. (2017)</xref> examined a diverse collection of tropical maize lines, our analyses focused on European dent material characterized by a relatively strong population structure (<xref ref-type="bibr" rid="B61">Lehermeier et&#xa0;al., 2014</xref>). Moreover, the population studied by <xref ref-type="bibr" rid="B74">Matias et&#xa0;al. (2017)</xref> was nearly twice the size of our investigation, potentially leading to increased recombination events between loci and reducing the potential size of haplotype blocks. Another contributing factor may be the limited representation of relationship captured by those haplotypes, as evidenced by the intermediate correlation between relationship coefficients obtained from markers and haplotypes. In contrast, canola, wheat, and soybean exhibited a high correlation in this regard. Unlike the findings of <xref ref-type="bibr" rid="B69">Ma et&#xa0;al. (2016)</xref> suggest, our study did not observe improved prediction accuracies in soybean using the method proposed by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref>. This discrepancy could be attributed to several factors, including differences in the traits under examination, as well as substantial variations in population size and marker density. It is worth noting that the method proposed by <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref> shares similarities with the LD-based method described earlier, implying that haplotype blocks formed using this method may already be represented using a specific LD threshold.</p>
<p>The <italic>&#x201c;HaploBlocker&#x201d;</italic> method (<xref ref-type="bibr" rid="B82">Pook et&#xa0;al., 2020</xref>) has the advantage of constructing subgroup-specific haplotype blocks and was implemented to address this aspect. However, this approach did not improve prediction accuracy and even led to a decrease of prediction accuracy in some cases. In canola and soybean, haplotype blocks from <italic>&#x201c;HaploBlocker&#x201d;</italic> effectively captured the genomic relationship represented by SNPs. In wheat, the representation was reasonable, but in maize, it was notably inadequate. Similar to the large fixed windows, haplotypes generated by this method are specific to genotypes or subpopulations. Consequently, haplotypes present in the validation set may not be observed in the training set, resulting in the inability to estimate their effects accurately and leading to decreased prediction accuracy due to the limited number of shared haplotypes between the training and validation sets. Particularly in the maize population, which exhibited strong population structure, the <italic>&#x201c;HaploBlocker&#x201d;</italic> method resulted in comparatively low prediction accuracies. This was pronounced with the family-wise cross validation, where the accuracies were diminished to nearly zero. In this scenario, even when using SNPs, the number of shared alleles or haplotypes between the training and validation sets will be minimized. This effect will be particularly prominent when employing a method that constructs subgroup-specific blocks.</p>
<p>In general, with the exception of wheat, prediction accuracies based on haplotype blocks using GBLUP and EGBLUP followed the correlation observed between relationship coefficients obtained from SNPs and haplotype blocks. This suggests that a portion of the prediction accuracy achieved with haplotypes is derived from reinterpreting the SNP information. However, in the case of wheat, this pattern did not hold true, even when using large fixed window blocks. Furthermore, considerable prediction accuracy differences were observed across models for wheat traits, but these differences were consistently compensated for by utilizing haplotype blocks with varying methods and parameters. This indicates that additional information beyond genetic relatedness contributes to the prediction accuracy when using haplotype blocks. One possible explanation is that haplotype blocks are generally considered to exhibit higher LD with QTL compared to individual markers (<xref ref-type="bibr" rid="B56">Jiang et&#xa0;al., 2018</xref>).</p>
<p>Multiple factors contribute to the accuracy of genomic prediction. One crucial factor is the relationship among genotypes, which is overlooked in random cross-validation approaches. In such cases, closely related genotypes may be included in both the training and validation sets, leading to higher prediction accuracies for related individuals (<xref ref-type="bibr" rid="B73">Massman et&#xa0;al., 2013</xref>; <xref ref-type="bibr" rid="B46">Hickey et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B107">Werner et&#xa0;al., 2020</xref>). Consequently, the prediction accuracies obtained from random cross-validation are population-specific and cannot be readily adopted to all breeding populations (<xref ref-type="bibr" rid="B107">Werner et&#xa0;al., 2020</xref>). To address this issue, we conducted a family-wise cross-validation in the maize dataset to assess the predictive performance of haplotype blocks for less related individuals. As expected from <xref ref-type="bibr" rid="B107">Werner et&#xa0;al. (2020)</xref>, we observed a decrease in prediction accuracy compared to random cross-validation. However, the relative ranking of haplotype block methods and parameters remained consistent with that of the random cross-validation, indicating no added benefit from haplotypes in predicting the breeding values of genetically distinct materials.</p>
<p>Moreover, GBLUP models trained with small haplotype blocks exhibited very similar prediction accuracies to models trained with SNPs. This is expected since haplotype effects can be partially defined as the sum of individual marker effects within their respective block. Another advantage of haplotype effects is their ability to capture local epistasis, as demonstrated by <xref ref-type="bibr" rid="B56">Jiang et&#xa0;al. (2018)</xref>. However, it is worth noting that purely additive models, especially in prediction methods like GBLUP where marker effects are estimated simultaneously, already implicitly capture local epistasis among markers in complete LD.</p>
<p>The use of haplotypes has been proposed as a means to address the challenges associated with apparent or phantom epistasis (<xref ref-type="bibr" rid="B111">Wood et&#xa0;al., 2014</xref>). Apparent or phantom epistasis can occur when two markers are in incomplete LD with QTL, resulting in statistically significant marker interactions in association studies and enhanced prediction accuracies in genomic prediction with models considering epistasis (<xref ref-type="bibr" rid="B111">Wood et&#xa0;al., 2014</xref>; <xref ref-type="bibr" rid="B28">de los Campos et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B89">Schrauf et&#xa0;al., 2020</xref>). This effect may be particularly pronounced in the wheat dataset, which had a significantly lower marker density compared to the other three datasets. Consequently, the use of haplotype blocks sometimes led to considerable improvements in prediction accuracy.</p>
<p>There is a multitude of factors affecting the accurate assembly of haplotype blocks and their respective haplotypes. Especially in complex plant genomes like the allopolyploids canola and wheat, SNP array markers can potentially be non-specific in terms of physical position, representing different homoeologous loci in different individuals (<xref ref-type="bibr" rid="B72">Mason et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B70">Makhoul et&#xa0;al., 2020</xref>). Furthermore, all methods to build haplotype blocks rely on known marker positions along the genome. These positions are obtained from a reference genome and are not necessarily the same in every population or even genotype. Especially if the reference genome is only distantly related. In such cases, a lack of precision in assembled haplotype blocks and their corresponding haplotypes may limit their potential in genomic prediction. Furthermore, haplotype block borders are not necessarily the same across populations and generations. Even though, <xref ref-type="bibr" rid="B33">Gabriel et&#xa0;al. (2002)</xref> showed high harmony of block structure across different human populations, however in plant breeding, with selection favoring positive alleles or haplotypes, this could ultimately change. Especially LD based haplotype blocks may only be useful for very few generations, since initially defined blocks will rapidly be disrupted by recombination or extended due to selection in later generations as the breeding program progresses. Indeed, an important goal of breeding is to accumulate favorable alleles through selection and recombination. This underlines the need for constant updating of both, the haplotype block assignment and the prediction model. Furthermore, besides the two fixed window approaches, all of the methods tested are only capable of identifying a proxy to true chromosomal recombination breakpoints. Even though crossovers tends to aggregate in recombination hotspots (<xref ref-type="bibr" rid="B63">Li and Stephens, 2003</xref>; <xref ref-type="bibr" rid="B77">Myers et&#xa0;al., 2006</xref>), haplotype blocking methods with limited marker density and population size may not necessarily be able to detect these hotspots. Therefore, there is a need to develop enhanced haplotype blocking pipelines that can effectively capture natural recombination patterns and address challenges associated with polyploidy, structural variations, and chromosomal rearrangements commonly observed in crop plants (<xref ref-type="bibr" rid="B72">Mason et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B86">Schiessl et&#xa0;al., 2019</xref>). Consequently, ongoing efforts focus on the development of innovative methods to capture local epistatic effects (<xref ref-type="bibr" rid="B82">Pook et&#xa0;al., 2020</xref>).</p>
<p>Unfortunately, we could not identify a single optimal haplotype blocking method that suits all datasets. Therefore, it is important to consider haplotype block construction methods and parameters as hyperparameters that require careful optimization, rather than fixed biological parameters. A breeding program that adopts haplotype block-based genomic selection should explore multiple haplotype blocking methods with different parameter settings. In general, the selected method should effectively capture relationships among individuals. Additionally, it is worth examining blocks of large size, as, in the case of the wheat dataset, larger blocks proved beneficial in improving prediction accuracy. The wheat dataset, which had the lowest marker density, generally showed the greatest improvements. This suggests that haplotype block-based genomic selection could be particularly valuable for breeding programs lacking access to high-density SNP arrays. However, further investigation is required in other datasets with varying SNP densities to validate these findings.</p>
<p>Although we observed only marginal beneficial effects of haplotype blocks in the canola, maize and soybean datasets on genomic prediction, they can still have a beneficial effect when used in other contexts. For example, haplotype blocks can help to identify regions of interest for the identification of candidate genes near significant marker-trait associations, or to compare different genotype groups at such loci (<xref ref-type="bibr" rid="B13">Clark, 2004</xref>; <xref ref-type="bibr" rid="B62">Li et&#xa0;al., 2017</xref>; <xref ref-type="bibr" rid="B101">Vollrath et&#xa0;al., 2021</xref>). Moreover, even if the majority of SNP markers exhibit intermediate minor allele frequency in a population, specific combinations of alleles represented as haplotypes may not be common in a population. Therefore, haplotypes can assist in identifying rare variants that have a potential impact on phenotypic traits. (<xref ref-type="bibr" rid="B9">Bloom et&#xa0;al., 2019</xref>; <xref ref-type="bibr" rid="B104">Wainschtein et&#xa0;al., 2022</xref>; <xref ref-type="bibr" rid="B106">Wang et&#xa0;al., 2023</xref>). Furthermore, especially in highly quantitative traits like yield where markers tend to have very small effects on traits, haplotype blocks can identify positive or negative chromosomal segments. This information can be implemented for cross designs to recombine haplotypes with positive effects (<xref ref-type="bibr" rid="B8">Bernardo and Thompson, 2016</xref>; <xref ref-type="bibr" rid="B108">Werner et&#xa0;al., 2018a</xref>). This can be considerably easier than selecting for single positive SNPs, as their positive effect can be obscured by deleterious SNPs in proximity that are only rarely separated by recombination in subsequent generations.</p>
</sec>
<sec id="s5" sec-type="conclusions">
<label>5</label>
<title>Conclusion</title>
<p>As anticipated based on numerous previous reports, our study confirms that haplotype blocks have the potential to enhance genomic selection, although the magnitude of improvement is sometimes only marginal. Haplotype blocks can particularly compensate for model differences when there is considerable variation in model performance across different prediction models. The extent of improvement with haplotypes compared to SNP-based predictions seem to be highly dependent on factors such as population, population structure, trait, and model. for a multitude of different traits from different crop species with different genome properties and breeding schemes, we were unable to identify optimal methods or parameters for constructing haplotype blocks in terms of prediction accuracy. Approaches based on LD resulted in improved prediction accuracies across various traits and demonstrated robustness in LD-threshold selection. However, the greatest improvements were observed with haplotype blocks consisting of entire chromosomes. Therefore, we recommend treating haplotype block definition as a tunable hyperparameter when employing genomic selection, taking into account extremely large haplotype blocks.</p>
</sec>
<sec id="s6" sec-type="data-availability">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: Please refer to the original publications of the four datasets.</p>
</sec>
<sec id="s7" sec-type="author-contributions">
<title>Author contributions</title>
<p>SW and RS designed the study. SW conceived the analysis, MF developed the software for Linkage Disequilibrium (LD) based haplotyping. KV-F and MF supervised the statistical analysis. SW wrote the manuscript. RS and KV-F revised the manuscript. All authors contributed to the article and approved the submitted version.</p>
</sec>
</body>
<back>
<sec id="s8" sec-type="funding-information">
<title>Funding</title>
<p>The work was funded by grant FKZ 031B0890A from the German Federal Ministry of Education and Research (BMBF) to MF and RS. Informatics infrastructure was provided by the BMBF-funded de.NBI Cloud within the German Network for Bioinformatics (de.NBI).</p>
</sec>
<ack>
<title>Acknowledgments</title>
<p>The authors thank Benjamin Wittkop, Christian Obermeier, Carola Zenke-Philippi and Lennard Ehrig for discussions on potential applications of haplotype blocks in plant breeding.</p>
</ack>
<sec id="s9" sec-type="COI-statement">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec id="s10" sec-type="disclaimer">
<title>Publisher&#x2019;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec id="s11" sec-type="supplementary-material">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fpls.2023.1217589/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fpls.2023.1217589/full#supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet_1.docx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.wordprocessingml.document"/>
<supplementary-material xlink:href="Table_1.xlsx" id="SM2" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
<supplementary-material xlink:href="Table_2.xlsx" id="SM3" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Atanda</surname> <given-names>S. A.</given-names>
</name>
<name>
<surname>Govindan</surname> <given-names>V.</given-names>
</name>
<name>
<surname>Singh</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Robbins</surname> <given-names>K. R.</given-names>
</name>
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Bentley</surname> <given-names>A. R.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>Sparse testing using genomic prediction improves selection for breeding targets in elite spring wheat</article-title>. <source>Theor. Appl. Genet.</source> <volume>135</volume>, <fpage>1939</fpage>&#x2013;<lpage>1950</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-022-04085-0</pub-id>
</citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ballesta</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Maldonado</surname> <given-names>C.</given-names>
</name>
<name>
<surname>P&#xe9;rez-Rodr&#xed;guez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Mora</surname> <given-names>F.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>SNP and haplotype-based genomic selection of quantitative traits in eucalyptus globulus</article-title>. <source>Plants</source> <volume>8</volume>, <elocation-id>331</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/plants8090331</pub-id>
</citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bandillo</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Jarquin</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Cregan</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Specht</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2015</year>). <article-title>A population structure and genome-wide association analysis on the USDA soybean germplasm collection</article-title>. <source>Plant Genome</source> <volume>8</volume>, <elocation-id>plantgenome2015.04.0024</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3835/plantgenome2015.04.0024</pub-id>
</citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Barrett</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Fry</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Maller</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Daly</surname> <given-names>M. J.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Haploview: analysis and visualization of LD and haplotype maps</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>263</fpage>&#x2013;<lpage>265</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/bth457</pub-id>
</citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bauer</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Falque</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Walter</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Bauland</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Camisan</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Campo</surname> <given-names>L.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Intraspecific variation of recombination rate in maize</article-title>. <source>Genome Biol.</source> <volume>9</volume> (<issue>14</issue>). doi:&#xa0;<pub-id pub-id-type="doi">10.1186/gb-2013-14-9-r103</pub-id>
</citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bayer</surname> <given-names>P. E.</given-names>
</name>
<name>
<surname>Petereit</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Danilevicz</surname> <given-names>M. F.</given-names>
</name>
<name>
<surname>Anderson</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Batley</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Edwards</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>The application of pangenomics and machine learning in genomic selection in plants</article-title>. <source>Plant Genome</source> <volume>14</volume>, <elocation-id>e20112</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/tpg2.20112</pub-id>
</citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernardo</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>1994</year>). <article-title>Prediction of maize single-cross performance using RFLPs and information from related hybrids</article-title>. <source>Crop Sci.</source> <volume>34</volume>, <elocation-id>cropsci1994.0011183X003400010003x</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.2135/cropsci1994.0011183X003400010003x</pub-id>
</citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bernardo</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>A. M.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Germplasm architecture revealed through chromosomal effects for quantitative traits in maize</article-title>. <source>Plant Genome</source> <volume>9</volume>, <elocation-id>plantgenome2016.03.0028</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3835/plantgenome2016.03.0028</pub-id>
</citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bloom</surname> <given-names>J. S.</given-names>
</name>
<name>
<surname>Boocock</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Treusch</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Sadhu</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Day</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Oates-Barker</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Rare variants contribute disproportionately to quantitative trait variation in yeast</article-title>. <source>eLife</source> <volume>8</volume>, <elocation-id>e49212</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.7554/eLife.49212</pub-id>
</citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Browning</surname> <given-names>S. R.</given-names>
</name>
<name>
<surname>Browning</surname> <given-names>B. L.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Rapid and accurate haplotype phasing and missing-data inference for whole-genome association studies by use of localized haplotype clustering</article-title>. <source>Am. J. Hum. Genet.</source> <volume>81</volume>, <fpage>1084</fpage>&#x2013;<lpage>1097</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1086/521987</pub-id>
</citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Browning</surname> <given-names>B. L.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Browning</surname> <given-names>S. R.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>A one-penny imputed genome from next-generation reference panels</article-title>. <source>Am. J. Hum. Genet.</source> <volume>103</volume>, <fpage>338</fpage>&#x2013;<lpage>348</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.ajhg.2018.07.015</pub-id>
</citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Carlson</surname> <given-names>C. S.</given-names>
</name>
<name>
<surname>Eberle</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Rieder</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Yi</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Kruglyak</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Nickerson</surname> <given-names>D. A.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Selecting a maximally informative set of single-nucleotide polymorphisms for association analyses using linkage disequilibrium</article-title>. <source>Am. J. Hum. Genet.</source> <volume>74</volume>, <fpage>106</fpage>&#x2013;<lpage>120</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1086/381000</pub-id>
</citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Clark</surname> <given-names>A. G.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>The role of haplotypes in candidate gene studies</article-title>. <source>Genet. Epidemiol.</source> <volume>27</volume>, <fpage>321</fpage>&#x2013;<lpage>333</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/gepi.20025</pub-id>
</citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Clarke</surname> <given-names>W. E.</given-names>
</name>
<name>
<surname>Higgins</surname> <given-names>E. E.</given-names>
</name>
<name>
<surname>Plieske</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Wieseke</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Sidebottom</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Khedikar</surname> <given-names>Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>A high-density SNP genotyping array for Brassica napus and its ancestral diploid species based on optimised selection of single-locus markers in the allotetraploid genome</article-title>. <source>Theor. Appl. Genet.</source> <volume>129</volume>, <fpage>1887</fpage>&#x2013;<lpage>1899</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-016-2746-7</pub-id>
</citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Combs</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Bernardo</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Accuracy of genomewide selection for different traits with constant population size, heritability, and number of markers</article-title>. <source>Plant Genome</source> <volume>6</volume>, <elocation-id>plantgenome2012.11.0030</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3835/plantgenome2012.11.0030</pub-id>
</citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Covarrubias-Pazaran</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Genome-assisted prediction of quantitative traits using the R package sommer</article-title>. <source>PLoS One</source> <volume>11</volume>, <elocation-id>e0156744</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0156744</pub-id>
</citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Covarrubias-Pazaran</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Software update: Moving the R package <italic>sommer</italic> to multivariate mixed models for genome-assisted prediction</article-title>. <source>Genetics</source>. doi:&#xa0;<pub-id pub-id-type="doi">10.1101/354639</pub-id>
</citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Crespo-Herrera</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Howard</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Piepho</surname> <given-names>H.-P.</given-names>
</name>
<name>
<surname>P&#xe9;rez-Rodr&#xed;guez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Montesinos-Lopez</surname> <given-names>O.</given-names>
</name>
<name>
<surname>Burgue&#xf1;o</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Genome-enabled prediction for sparse testing in multi-environmental wheat trials</article-title>. <source>Plant Genome</source> <volume>14</volume>, <elocation-id>e20151</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1002/tpg2.20151</pub-id>
</citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Beyene</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Kassa</surname> <given-names>S.</given-names>
</name>
<name>
<surname>P&#xe9;rez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Hickey</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Genomic prediction in maize breeding populations with genotyping-by-sequencing</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>3</volume>, <fpage>1903</fpage>&#x2013;<lpage>1926</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.113.008227</pub-id>
</citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Jarqu&#xed;n</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Franco</surname> <given-names>J.</given-names>
</name>
<name>
<surname>P&#xe9;rez-Rodr&#xed;guez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Burgue&#xf1;o</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Saint-Pierre</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>Genomic prediction of gene bank wheat landraces</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>6</volume>, <fpage>1819</fpage>&#x2013;<lpage>1834</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.116.029637</pub-id>
</citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>P&#xe9;rez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Hickey</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Burgue&#xf1;o</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ornella</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Cer&#xf3;n-Rojas</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Genomic prediction in CIMMYT maize and wheat breeding programs</article-title>. <source>Heredity</source> <volume>112</volume>, <fpage>48</fpage>&#x2013;<lpage>60</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/hdy.2013.16</pub-id>
</citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cuyabano</surname> <given-names>B. C.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lund</surname> <given-names>M. S.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Genomic prediction of genetic merit using LD-based haplotypes in the Nordic Holstein population</article-title>. <source>BMC Genomics</source> <volume>15</volume>, <elocation-id>1171</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1471-2164-15-1171</pub-id>
</citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cuyabano</surname> <given-names>B. C.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Lund</surname> <given-names>M. S.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Selection of haplotype variables from a high-density marker map for genomic prediction</article-title>. <source>Genet. Selection Evol.</source> <volume>47</volume>, <fpage>61</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-015-0143-3</pub-id>
</citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daetwyler</surname> <given-names>H. D.</given-names>
</name>
<name>
<surname>Bansal</surname> <given-names>U. K.</given-names>
</name>
<name>
<surname>Bariana</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Hayden</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Hayes</surname> <given-names>B. J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Genomic prediction for rust resistance in diverse wheat landraces</article-title>. <source>Theor. Appl. Genet.</source> <volume>127</volume>, <fpage>1795</fpage>&#x2013;<lpage>1803</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-014-2341-8</pub-id>
</citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Daly</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Rioux</surname> <given-names>J. D.</given-names>
</name>
<name>
<surname>Schaffner</surname> <given-names>S. F.</given-names>
</name>
<name>
<surname>Hudson</surname> <given-names>T. J.</given-names>
</name>
<name>
<surname>Lander</surname> <given-names>E. S.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>High-resolution haplotype structure in the human genome</article-title>. <source>Nat. Genet.</source> <volume>29</volume>, <fpage>229</fpage>&#x2013;<lpage>232</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng1001-229</pub-id>
</citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de Bakker</surname> <given-names>P. I. W.</given-names>
</name>
<name>
<surname>Yelensky</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Pe&#x2019;er</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Gabriel</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Daly</surname> <given-names>M. J.</given-names>
</name>
<name>
<surname>Altshuler</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Efficiency and power in genetic association studies</article-title>. <source>Nat. Genet.</source> <volume>37</volume>, <fpage>1217</fpage>&#x2013;<lpage>1223</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng1669</pub-id>
</citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de los Campos</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Gianola</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Rosa</surname> <given-names>G. J. M.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Reproducing kernel Hilbert spaces regression: A general framework for genetic evaluation1</article-title>. <source>J. Anim. Sci.</source> <volume>87</volume>, <fpage>1883</fpage>&#x2013;<lpage>1887</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2527/jas.2008-1259</pub-id>
</citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>de los Campos</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Sorensen</surname> <given-names>D. A.</given-names>
</name>
<name>
<surname>Toro</surname> <given-names>M. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Imperfect linkage disequilibrium generates phantom epistasis (&amp; Perils of big data)</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>9</volume>, <fpage>1429</fpage>&#x2013;<lpage>1436</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.119.400101</pub-id>
</citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Devlin</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Risch</surname> <given-names>N.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>A comparison of linkage disequilibrium measures for fine-scale mapping</article-title>. <source>Genomics</source> <volume>29</volume>, <fpage>311</fpage>&#x2013;<lpage>322</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1006/geno.1995.9003</pub-id>
</citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Druet</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Macleod</surname> <given-names>I. M.</given-names>
</name>
<name>
<surname>Hayes</surname> <given-names>B. J.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Toward genomic prediction from whole-genome sequence data: impact of sequencing design on genotype imputation and accuracy of predictions</article-title>. <source>Heredity</source> <volume>112</volume>, <fpage>39</fpage>&#x2013;<lpage>47</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/hdy.2013.13</pub-id>
</citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Edwards</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Batley</surname> <given-names>J.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Plant genome sequencing: applications for crop improvement</article-title>. <source>Plant Biotechnol. J.</source> <volume>8</volume>, <fpage>2</fpage>&#x2013;<lpage>9</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/j.1467-7652.2009.00459.x</pub-id>
</citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Edwards</surname> <given-names>S. M.</given-names>
</name>
<name>
<surname>Buntjer</surname> <given-names>J. B.</given-names>
</name>
<name>
<surname>Jackson</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Bentley</surname> <given-names>A. R.</given-names>
</name>
<name>
<surname>Lage</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Byrne</surname> <given-names>E.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>The effects of training population design on genomic prediction accuracy in wheat</article-title>. <source>Theor. Appl. Genet.</source> <volume>132</volume>, <fpage>1943</fpage>&#x2013;<lpage>1952</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-019-03327-y</pub-id>
</citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gabriel</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Schaffner</surname> <given-names>S. F.</given-names>
</name>
<name>
<surname>Nguyen</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Moore</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Roy</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Blumenstiel</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2002</year>). <article-title>The structure of haplotype blocks in the human genome</article-title>. <source>Science</source> <volume>296</volume>, <fpage>2225</fpage>&#x2013;<lpage>2229</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1126/science.1069424</pub-id>
</citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gianola</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Priors in whole-genome regression: the bayesian alphabet returns</article-title>. <source>Genetics</source> <volume>194</volume>, <fpage>573</fpage>&#x2013;<lpage>596</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.113.151753</pub-id>
</citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goddard</surname> <given-names>M. E.</given-names>
</name>
<name>
<surname>Hayes</surname> <given-names>B. j.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>Genomic selection</article-title>. <source>J. Anim. Breed. Genet.</source> <volume>124</volume>, <fpage>323</fpage>&#x2013;<lpage>330</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/j.1439-0388.2007.00702.x</pub-id>
</citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Grant</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>R. T.</given-names>
</name>
<name>
<surname>Cannon</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Shoemaker</surname> <given-names>R. C.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>SoyBase, the USDA-ARS soybean genetics and genomics database</article-title>. <source>Nucleic Acids Res.</source> <volume>38</volume>, <fpage>D843</fpage>&#x2013;<lpage>D846</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/nar/gkp798</pub-id>
</citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Habier</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Fernando</surname> <given-names>R. L.</given-names>
</name>
<name>
<surname>Garrick</surname> <given-names>D. J.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Genomic BLUP decoded: A look into the black box of genomic prediction</article-title>. <source>Genetics</source> <volume>194</volume>, <fpage>597</fpage>&#x2013;<lpage>607</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.113.152207</pub-id>
</citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Habier</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Tetens</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Seefried</surname> <given-names>F.-R.</given-names>
</name>
<name>
<surname>Lichtner</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Thaller</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>The impact of genetic relationship information on genomic breeding values in German Holstein cattle</article-title>. <source>Genet. Sel Evol.</source> <volume>42</volume>, <elocation-id>5</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1297-9686-42-5</pub-id>
</citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hao</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Ning</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Huang</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wu</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>The resurgence of introgression breeding, as exemplified in wheat improvement</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2020.00252</pub-id>
</citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hayes</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Bowman</surname> <given-names>P. J.</given-names>
</name>
<name>
<surname>Chamberlain</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Goddard</surname> <given-names>M. E.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Invited review: Genomic selection in dairy cattle: Progress and challenges</article-title>. <source>J. Dairy Sci.</source> <volume>92</volume>, <fpage>433</fpage>&#x2013;<lpage>443</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3168/jds.2008-1646</pub-id>
</citation>
</ref>
<ref id="B41">
<citation citation-type="web">
<person-group person-group-type="author">
<name>
<surname>Hayes</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Macleod</surname> <given-names>I. M.</given-names>
</name>
<name>
<surname>Daetwyler</surname> <given-names>H. D.</given-names>
</name>
<name>
<surname>Bowman</surname> <given-names>P. J.</given-names>
</name>
<name>
<surname>Chamberlian</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Vander Jagt</surname> <given-names>C. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>)<article-title>Genomic prediction from whole genome sequence in livestock: the 1000 Bull Genomes Project. in <italic>10</italic>
</article-title>. In: <source>World Congress of Genetics Applied to Livestock Production</source> (<publisher-loc>Vancouver, Canada</publisher-loc>) (Accessed <access-date>November 22, 2022</access-date>).</citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Henderson</surname> <given-names>C. R.</given-names>
</name>
</person-group> (<year>1975</year>). <article-title>Best linear unbiased estimation and prediction under a selection model</article-title>. <source>Biometrics</source> <volume>31</volume>, <fpage>423</fpage>&#x2013;<lpage>447</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2307/2529430</pub-id>
</citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Henderson</surname> <given-names>C. R.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Best linear unbiased prediction of nonadditive genetic merits in noninbred populations</article-title>. <source>J. Anim. Sci.</source> <volume>60</volume>, <fpage>111</fpage>&#x2013;<lpage>117</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2527/jas1985.601111x</pub-id>
</citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hess</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Druet</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Hess</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Garrick</surname> <given-names>D.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Fixed-length haplotypes can improve genomic prediction accuracy in an admixed dairy cattle population</article-title>. <source>Genet. Selection Evol.</source> <volume>49</volume>, <fpage>54</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-017-0329-y</pub-id>
</citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hickey</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Chiurugwi</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Mackay</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Powell</surname> <given-names>W.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Genomic prediction unifies animal and plant breeding programs to form platforms for biological discovery</article-title>. <source>Nat. Genet.</source> <volume>49</volume>, <fpage>1297</fpage>&#x2013;<lpage>1303</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng.3920</pub-id>
</citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hickey</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Dreisigacker</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Hearne</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Babu</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Prasanna</surname> <given-names>B. M.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Evaluation of genomic selection training population designs and genotyping strategies in plant breeding programs using simulation</article-title>. <source>Crop Sci.</source> <volume>54</volume>, <fpage>1476</fpage>&#x2013;<lpage>1488</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2135/cropsci2013.03.0195</pub-id>
</citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hill</surname> <given-names>W. G.</given-names>
</name>
</person-group> (<year>1981</year>). <article-title>Estimation of effective population size from data on linkage disequilibrium1</article-title>. <source>Genet. Res.</source> <volume>38</volume>, <fpage>209</fpage>&#x2013;<lpage>216</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1017/S0016672300020553</pub-id>
</citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hill</surname> <given-names>W. G.</given-names>
</name>
<name>
<surname>Robertson</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>1968</year>). <article-title>Linkage disequilibrium in finite populations</article-title>. <source>Theoret. Appl. Genet.</source> <volume>38</volume>, <fpage>226</fpage>&#x2013;<lpage>231</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/BF01245622</pub-id>
</citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hofheinz</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Frisch</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Heteroscedastic ridge regression approaches for genome-wide prediction with a focus on computational efficiency and accurate effect estimation</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>4</volume>, <fpage>539</fpage>&#x2013;<lpage>546</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.113.010025</pub-id>
</citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hyten</surname> <given-names>D. L.</given-names>
</name>
<name>
<surname>Song</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Choi</surname> <given-names>I.-Y.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>R. L.</given-names>
</name>
<name>
<surname>Costa</surname> <given-names>J. M.</given-names>
</name>
<etal/>
</person-group>. (<year>2006</year>). <article-title>Impacts of genetic bottlenecks on soybean genome diversity</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>103</volume>, <fpage>16666</fpage>&#x2013;<lpage>16671</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1073/pnas.0604379103</pub-id>
</citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jan</surname> <given-names>H. U.</given-names>
</name>
<name>
<surname>Abbadi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>L&#xfc;cke</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Nichols</surname> <given-names>R. A.</given-names>
</name>
<name>
<surname>Snowdon</surname> <given-names>R. J.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Genomic prediction of testcross performance in canola (Brassica napus)</article-title>. <source>PLoS One</source> <volume>11</volume>, <elocation-id>e0147769</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0147769</pub-id>
</citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jarquin</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Howard</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Crossa</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Beyene</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Gowda</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Martini</surname> <given-names>J. W. R.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Genomic prediction enhanced sparse testing for multi-environment trials</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>10</volume>, <fpage>2725</fpage>&#x2013;<lpage>2739</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.120.401349</pub-id>
</citation>
</ref>
<ref id="B53">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jarquin</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Specht</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Lorenz</surname> <given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Prospects of genomic prediction in the USDA soybean germplasm collection: historical data creates robust models for enhancing selection of accessions</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>6</volume>, <fpage>2329</fpage>&#x2013;<lpage>2341</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.116.031443</pub-id>
</citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jeffreys</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Kauppi</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Neumann</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Intensely punctate meiotic recombination in the class II region of the major histocompatibility complex</article-title>. <source>Nat. Genet.</source> <volume>29</volume>, <fpage>217</fpage>&#x2013;<lpage>222</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng1001-217</pub-id>
</citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Reif</surname> <given-names>J. C.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Modeling epistasis in genomic selection</article-title>. <source>Genetics</source> <volume>201</volume>, <fpage>759</fpage>&#x2013;<lpage>768</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.115.177907</pub-id>
</citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Schmidt</surname> <given-names>R. H.</given-names>
</name>
<name>
<surname>Reif</surname> <given-names>J. C.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Haplotype-based genome-wide prediction models exploit local epistatic interactions among markers</article-title>. <source>G3 (Bethesda)</source> <volume>8</volume>, <fpage>1687</fpage>&#x2013;<lpage>1699</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.117.300548</pub-id>
</citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname> <given-names>S. A.</given-names>
</name>
<name>
<surname>Brossard</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Roshandel</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Paterson</surname> <given-names>A. D.</given-names>
</name>
<name>
<surname>Bull</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Yoo</surname> <given-names>Y. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>gpart: human genome partitioning and visualization of high-density SNP data by identifying haplotype blocks</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>4419</fpage>&#x2013;<lpage>4421</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/bioinformatics/btz308</pub-id>
</citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lado</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Matus</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Rodr&#xed;guez</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Inostroza</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Poland</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Belzile</surname> <given-names>F.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Increased genomic prediction accuracy in wheat breeding through spatial adjustment of field trial data</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>3</volume>, <fpage>2105</fpage>&#x2013;<lpage>2114</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.113.007807</pub-id>
</citation>
</ref>
<ref id="B59">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lande</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Thompson</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>1990</year>). <article-title>Efficiency of marker-assisted selection in the improvement of quantitative traits</article-title>. <source>Genetics</source> <volume>124</volume>, <fpage>743</fpage>&#x2013;<lpage>756</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/genetics/124.3.743</pub-id>
</citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lee</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Chawla</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Obermeier</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Dreyer</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Abbadi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Snowdon</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Chromosome-scale assembly of winter oilseed rape brassica napus</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2020.00496</pub-id>
</citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lehermeier</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Kr&#xe4;mer</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Bauer</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Bauland</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Camisan</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Campo</surname> <given-names>L.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Usefulness of multiparental populations of maize (Zea mays L.) for genome-based prediction</article-title>. <source>Genetics</source> <volume>198</volume>, <fpage>3</fpage>&#x2013;<lpage>16</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.114.161943</pub-id>
</citation>
</ref>
<ref id="B62">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Han</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Genome-wide association study discovered candidate genes of Verticillium wilt resistance in upland cotton (Gossypium hirsutum L.)</article-title>. <source>Plant Biotechnol. J.</source> <volume>15</volume>, <fpage>1520</fpage>&#x2013;<lpage>1532</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/pbi.12734</pub-id>
</citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Stephens</surname> <given-names>M.</given-names>
</name>
</person-group> (<year>2003</year>). <article-title>Modeling linkage disequilibrium and identifying recombination hotspots using single-nucleotide polymorphism data</article-title>. <source>Genetics</source> <volume>165</volume>, <fpage>2213</fpage>&#x2013;<lpage>2233</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/genetics/165.4.2213</pub-id>
</citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>H.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Genomic prediction of carcass traits using different haplotype block partitioning methods in beef cattle</article-title>. <source>Evolutionary Appl.</source> <volume>15</volume>, <fpage>2028</fpage>&#x2013;<lpage>2042</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/eva.13491</pub-id>
</citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Zhu</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Zhou</surname> <given-names>P.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Genomic prediction using LD-based haplotypes inferred from high-density chip and imputed sequence variants in Chinese simmental beef cattle</article-title>. <source>Front. Genet.</source> <volume>12</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2021.665382</pub-id>
</citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Tan</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Prakapenka</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Da</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Haplotype analysis of genomic prediction using structural and functional genomic information for seven human phenotypes</article-title>. <source>Front. Genet.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2020.588907</pub-id>
</citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Schmidt</surname> <given-names>R. H.</given-names>
</name>
<name>
<surname>Reif</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Selecting closely-linked SNPs based on local epistatic effects for haplotype construction improves power of association mapping</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>9</volume>, <fpage>4115</fpage>&#x2013;<lpage>4126</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.119.400451</pub-id>
</citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lorenz</surname> <given-names>A. J.</given-names>
</name>
<name>
<surname>Hamblin</surname> <given-names>M. T.</given-names>
</name>
<name>
<surname>Jannink</surname> <given-names>J.-L.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Performance of single nucleotide polymorphisms versus haplotypes for genome-wide association analysis in barley</article-title>. <source>PLoS One</source> <volume>5</volume>, <elocation-id>e14079</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0014079</pub-id>
</citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ma</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Reif</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Jiang</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Z.</given-names>
</name>
<etal/>
</person-group>. (<year>2016</year>). <article-title>Potential of marker selection to increase prediction accuracy of genomic selection in soybean (Glycine max L.)</article-title>. <source>Mol. Breed.</source> <volume>36</volume>, <fpage>113</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s11032-016-0504-9</pub-id>
</citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Makhoul</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Rambla</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Voss-Fels</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Hickey</surname> <given-names>L. T.</given-names>
</name>
<name>
<surname>Snowdon</surname> <given-names>R. J.</given-names>
</name>
<name>
<surname>Obermeier</surname> <given-names>C.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Overcoming polyploidy pitfalls: a user guide for effective SNP conversion into KASP markers in wheat</article-title>. <source>Theor. Appl. Genet.</source> <volume>133</volume>, <fpage>2413</fpage>&#x2013;<lpage>2430</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-020-03608-x</pub-id>
</citation>
</ref>
<ref id="B71">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maldonado</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Mora</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Bertagna</surname> <given-names>F. A. B.</given-names>
</name>
<name>
<surname>Kuki</surname> <given-names>M. C.</given-names>
</name>
<name>
<surname>Scapim</surname> <given-names>C. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>SNP- and haplotype-based GWAS of flowering-related traits in maize with network-assisted gene prioritization</article-title>. <source>Agronomy</source> <volume>9</volume>, <elocation-id>725</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.3390/agronomy9110725</pub-id>
</citation>
</ref>
<ref id="B72">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mason</surname> <given-names>A. S.</given-names>
</name>
<name>
<surname>Higgins</surname> <given-names>E. E.</given-names>
</name>
<name>
<surname>Snowdon</surname> <given-names>R. J.</given-names>
</name>
<name>
<surname>Batley</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Stein</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Werner</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>A user guide to the Brassica 60K Illumina Infinium<sup>TM</sup> SNP genotyping array</article-title>. <source>Theor. Appl. Genet.</source> <volume>130</volume>, <fpage>621</fpage>&#x2013;<lpage>633</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-016-2849-1</pub-id>
</citation>
</ref>
<ref id="B73">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Massman</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Gordillo</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Lorenzana</surname> <given-names>R. E.</given-names>
</name>
<name>
<surname>Bernardo</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Genomewide predictions from maize single-cross data</article-title>. <source>Theor. Appl. Genet.</source> <volume>126</volume>, <fpage>13</fpage>&#x2013;<lpage>22</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-012-1955-y</pub-id>
</citation>
</ref>
<ref id="B74">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Matias</surname> <given-names>F. I.</given-names>
</name>
<name>
<surname>Galli</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Correia Granato</surname> <given-names>I. S.</given-names>
</name>
<name>
<surname>Fritsche-Neto</surname> <given-names>R.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Genomic prediction of autogamous and allogamous plants by SNPs and haplotypes</article-title>. <source>Crop Sci.</source> <volume>57</volume>, <fpage>2951</fpage>&#x2013;<lpage>2958</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2135/cropsci2017.01.0022</pub-id>
</citation>
</ref>
<ref id="B75">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Meuwissen</surname> <given-names>T. H. E.</given-names>
</name>
<name>
<surname>Hayes</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Goddard</surname> <given-names>M. E.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>Prediction of total genetic value using genome-wide dense marker maps</article-title>. <source>Genetics</source> <volume>157</volume>, <fpage>1819</fpage>&#x2013;<lpage>1829</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/genetics/157.4.1819</pub-id>
</citation>
</ref>
<ref id="B76">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Millet</surname> <given-names>E. J.</given-names>
</name>
<name>
<surname>Kruijer</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Coupel-Ledru</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Alvarez Prado</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Cabrera-Bosquet</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Lacube</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Genomic prediction of maize yield across European environmental conditions</article-title>. <source>Nat. Genet.</source> <volume>51</volume>, <fpage>952</fpage>&#x2013;<lpage>956</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41588-019-0414-y</pub-id>
</citation>
</ref>
<ref id="B77">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Myers</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Spencer</surname> <given-names>C. C. A.</given-names>
</name>
<name>
<surname>Auton</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Bottolo</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Freeman</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Donnelly</surname> <given-names>P.</given-names>
</name>
<etal/>
</person-group>. (<year>2006</year>). <article-title>The distribution and causes of meiotic recombination in the human genome</article-title>. <source>Biochem. Soc. Trans.</source> <volume>34</volume>, <fpage>526</fpage>&#x2013;<lpage>530</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1042/BST0340526</pub-id>
</citation>
</ref>
<ref id="B78">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ni</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Cavero</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Fangmann</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Erbe</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Simianer</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Whole-genome sequence-based genomic prediction in laying chickens with different genomic relationship matrices to account for genetic architecture</article-title>. <source>Genet. Selection Evol.</source> <volume>49</volume>, <elocation-id>8</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-016-0277-y</pub-id>
</citation>
</ref>
<ref id="B79">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Norman</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Taylor</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Edwards</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Kuchel</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Optimising genomic selection in wheat: effect of marker density, population size and population structure on prediction accuracy</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>8</volume>, <fpage>2889</fpage>&#x2013;<lpage>2899</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.118.200311</pub-id>
</citation>
</ref>
<ref id="B80">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Park</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Casella</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>The bayesian lasso</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>103</volume>, <fpage>681</fpage>&#x2013;<lpage>686</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1198/016214508000000337</pub-id>
</citation>
</ref>
<ref id="B81">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>P&#xe9;rez</surname> <given-names>P.</given-names>
</name>
<name>
<surname>de los Campos</surname> <given-names>G.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Genome-wide regression and prediction with the BGLR statistical package</article-title>. <source>Genetics</source> <volume>198</volume>, <fpage>483</fpage>&#x2013;<lpage>495</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.114.164442</pub-id>
</citation>
</ref>
<ref id="B82">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pook</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Freudenthal</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Korte</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Simianer</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Using local convolutional neural networks for genomic prediction</article-title>. <source>Front. Genet.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2020.561497</pub-id>
</citation>
</ref>
<ref id="B83">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pook</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Schlather</surname> <given-names>M.</given-names>
</name>
<name>
<surname>de los Campos</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Mayer</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Schoen</surname> <given-names>C. C.</given-names>
</name>
<name>
<surname>Simianer</surname> <given-names>H.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>HaploBlocker: creation of subgroup-specific haplotype blocks and libraries</article-title>. <source>Genetics</source> <volume>212</volume>, <fpage>1045</fpage>&#x2013;<lpage>1061</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/genetics.119.302283</pub-id>
</citation>
</ref>
<ref id="B84">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Raymond</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Bouwman</surname> <given-names>A. C.</given-names>
</name>
<name>
<surname>Schrooten</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Houwing-Duistermaat</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Veerkamp</surname> <given-names>R. F.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Utility of whole-genome sequence data for across-breed genomic prediction</article-title>. <source>Genet. Sel Evol.</source> <volume>50</volume>, <elocation-id>27</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-018-0396-8</pub-id>
</citation>
</ref>
<ref id="B85">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Reich</surname> <given-names>D. E.</given-names>
</name>
<name>
<surname>Cargill</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Bolk</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Ireland</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Sabeti</surname> <given-names>P. C.</given-names>
</name>
<name>
<surname>Richter</surname> <given-names>D. J.</given-names>
</name>
<etal/>
</person-group>. (<year>2001</year>). <article-title>Linkage disequilibrium in the human genome</article-title>. <source>Nature</source> <volume>411</volume>, <fpage>199</fpage>&#x2013;<lpage>204</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/35075590</pub-id>
</citation>
</ref>
<ref id="B86">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schiessl</surname> <given-names>S.-V.</given-names>
</name>
<name>
<surname>Katche</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Ihien</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Chawla</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Mason</surname> <given-names>A. S.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>The role of genomic structural variation in the genetic improvement of polyploid crops</article-title>. <source>Crop J.</source> <volume>7</volume>, <fpage>127</fpage>&#x2013;<lpage>140</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.cj.2018.07.006</pub-id>
</citation>
</ref>
<ref id="B87">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schmutz</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Cannon</surname> <given-names>S. B.</given-names>
</name>
<name>
<surname>Schlueter</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ma</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Mitros</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>W.</given-names>
</name>
<etal/>
</person-group>. (<year>2010</year>). <article-title>Genome sequence of the palaeopolyploid soybean</article-title>. <source>Nature</source> <volume>463</volume>, <fpage>178</fpage>&#x2013;<lpage>183</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/nature08670</pub-id>
</citation>
</ref>
<ref id="B88">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schnable</surname> <given-names>P. S.</given-names>
</name>
<name>
<surname>Ware</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Fulton</surname> <given-names>R. S.</given-names>
</name>
<name>
<surname>Stein</surname> <given-names>J. C.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Pasternak</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2009</year>). <article-title>The B73 maize genome: complexity, diversity, and dynamics</article-title>. <source>Science</source> <volume>326</volume>, <fpage>1112</fpage>&#x2013;<lpage>1115</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1126/science.1178534</pub-id>
</citation>
</ref>
<ref id="B89">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schrauf</surname> <given-names>M. F.</given-names>
</name>
<name>
<surname>Martini</surname> <given-names>J. W. R.</given-names>
</name>
<name>
<surname>Simianer</surname> <given-names>H.</given-names>
</name>
<name>
<surname>de los Campos</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Cantet</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Freudenthal</surname> <given-names>J.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Phantom epistasis in genomic selection: on the predictive ability of Epistatic models</article-title>. <source>G3 Genes|Genomes|Genetics</source> <volume>10</volume>, <fpage>3137</fpage>&#x2013;<lpage>3145</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1534/g3.120.401300</pub-id>
</citation>
</ref>
<ref id="B90">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Solberg</surname> <given-names>T. R.</given-names>
</name>
<name>
<surname>Sonesson</surname> <given-names>A. K.</given-names>
</name>
<name>
<surname>Woolliams</surname> <given-names>J. A.</given-names>
</name>
<name>
<surname>Meuwissen</surname> <given-names>T. H. E.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Genomic selection using different marker types and densities</article-title>. <source>J. Anim. Sci.</source> <volume>86</volume>, <fpage>2447</fpage>&#x2013;<lpage>2454</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2527/jas.2007-0010</pub-id>
</citation>
</ref>
<ref id="B91">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Soleimani</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lehnert</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Keilwagen</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Plieske</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Ordon</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Naseri Rad</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>Comparison between core set selection methods using different illumina marker platforms: A case study of assessment of diversity in wheat</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2020.01040</pub-id>
</citation>
</ref>
<ref id="B92">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Hyten</surname> <given-names>D. L.</given-names>
</name>
<name>
<surname>Jia</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Quigley</surname> <given-names>C. V.</given-names>
</name>
<name>
<surname>Fickus</surname> <given-names>E. W.</given-names>
</name>
<name>
<surname>Nelson</surname> <given-names>R. L.</given-names>
</name>
<etal/>
</person-group>. (<year>2013</year>). <article-title>Development and evaluation of soySNP50K, a high-density genotyping array for soybean</article-title>. <source>PloS One</source> <volume>8</volume>, <elocation-id>e54985</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0054985</pub-id>
</citation>
</ref>
<ref id="B93">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Technow</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Riedelsheimer</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Schrag</surname> <given-names>T. A.</given-names>
</name>
<name>
<surname>Melchinger</surname> <given-names>A. E.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>Genomic prediction of hybrid performance in maize with models incorporating dominance and population specific marker effects</article-title>. <source>Theor. Appl. Genet.</source> <volume>125</volume>, <fpage>1181</fpage>&#x2013;<lpage>1194</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-012-1905-8</pub-id>
</citation>
</ref>
<ref id="B94">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Terraillon</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Frisch</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Falke</surname> <given-names>K. C.</given-names>
</name>
<name>
<surname>Jaiser</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Spiller</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Csel&#xe9;nyi</surname> <given-names>L.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Genomic prediction can provide precise estimates of the genotypic value of barley lines evaluated in unreplicated trials</article-title>. <source>Front. Plant Sci.</source> <volume>13</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2022.735256</pub-id>
</citation>
</ref>
<ref id="B95">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>van Binsbergen</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Calus</surname> <given-names>M. P. L.</given-names>
</name>
<name>
<surname>Bink</surname> <given-names>M. C. A. M.</given-names>
</name>
<name>
<surname>van Eeuwijk</surname> <given-names>F. A.</given-names>
</name>
<name>
<surname>Schrooten</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Veerkamp</surname> <given-names>R. F.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Genomic prediction using imputed whole-genome sequence data in Holstein Friesian cattle</article-title>. <source>Genet. Selection Evol.</source> <volume>47</volume>, <elocation-id>71</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-015-0149-x</pub-id>
</citation>
</ref>
<ref id="B96">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanLiere</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Rosenberg</surname> <given-names>N. A.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Mathematical properties of the r2 measure of linkage disequilibrium</article-title>. <source>Theor. Population Biol.</source> <volume>74</volume>, <fpage>130</fpage>&#x2013;<lpage>137</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1016/j.tpb.2008.05.006</pub-id>
</citation>
</ref>
<ref id="B97">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanRaden</surname> <given-names>P. M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Efficient methods to compute genomic predictions</article-title>. <source>J. Dairy Sci.</source> <volume>91</volume>, <fpage>4414</fpage>&#x2013;<lpage>4423</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3168/jds.2007-0980</pub-id>
</citation>
</ref>
<ref id="B98">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>VanRaden</surname> <given-names>P. M.</given-names>
</name>
<name>
<surname>Van Tassell</surname> <given-names>C. P.</given-names>
</name>
<name>
<surname>Wiggans</surname> <given-names>G. R.</given-names>
</name>
<name>
<surname>Sonstegard</surname> <given-names>T. S.</given-names>
</name>
<name>
<surname>Schnabel</surname> <given-names>R. D.</given-names>
</name>
<name>
<surname>Taylor</surname> <given-names>J. F.</given-names>
</name>
<etal/>
</person-group>. (<year>2009</year>). <article-title>Invited Review: Reliability of genomic predictions for North American Holstein bulls</article-title>. <source>J. Dairy Sci.</source> <volume>92</volume>, <fpage>16</fpage>&#x2013;<lpage>24</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3168/jds.2008-1514</pub-id>
</citation>
</ref>
<ref id="B99">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Villumsen</surname> <given-names>T. M.</given-names>
</name>
<name>
<surname>Janss</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Bayesian genomic selection: the effect of haplotype length and priors</article-title>. <source>BMC Proc.</source> <volume>3</volume>, <elocation-id>S11</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1753-6561-3-S1-S11</pub-id>
</citation>
</ref>
<ref id="B100">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Villumsen</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Janss</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Lund</surname> <given-names>M. s.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>The importance of haplotype length and heritability using genomic selection in dairy cattle</article-title>. <source>J. Anim. Breed. Genet.</source> <volume>126</volume>, <fpage>3</fpage>&#x2013;<lpage>13</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/j.1439-0388.2008.00747.x</pub-id>
</citation>
</ref>
<ref id="B101">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vollrath</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Chawla</surname> <given-names>H. S.</given-names>
</name>
<name>
<surname>Alnajar</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Gabur</surname> <given-names>I.</given-names>
</name>
<name>
<surname>Lee</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Weber</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Dissection of quantitative blackleg resistance reveals novel variants of resistance gene rlm9 in elite Brassica napus</article-title>. <source>Front. Plant Sci.</source> <volume>12</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2021.749491</pub-id>
</citation>
</ref>
<ref id="B102">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Voss-Fels</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Stahl</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Wittkop</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lichthardt</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Nagler</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Rose</surname> <given-names>T.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Breeding improves wheat productivity under contrasting agrochemical input levels</article-title>. <source>Nat. Plants</source> <volume>5</volume>, <fpage>706</fpage>&#x2013;<lpage>714</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41477-019-0445-5</pub-id>
</citation>
</ref>
<ref id="B103">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Voss-Fels</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Wei</surname> <given-names>X.</given-names>
</name>
<name>
<surname>Ross</surname> <given-names>E. M.</given-names>
</name>
<name>
<surname>Frisch</surname> <given-names>M.</given-names>
</name>
<name>
<surname>Aitken</surname> <given-names>K. S.</given-names>
</name>
<name>
<surname>Cooper</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2021</year>). <article-title>Strategies and considerations for implementing genomic selection to improve traits with additive and non-additive genetic architectures in sugarcane breeding</article-title>. <source>Theor. Appl. Genet.</source> <volume>134</volume>, <fpage>1493</fpage>&#x2013;<lpage>1511</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-021-03785-3</pub-id>
</citation>
</ref>
<ref id="B104">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wainschtein</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Jain</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Cupples</surname> <given-names>L. A.</given-names>
</name>
<name>
<surname>Shadyab</surname> <given-names>A. H.</given-names>
</name>
<name>
<surname>McKnight</surname> <given-names>B.</given-names>
</name>
<etal/>
</person-group>. (<year>2022</year>). <article-title>Assessing the contribution of rare variants to complex trait heritability from whole-genome sequence data</article-title>. <source>Nat. Genet.</source> <volume>54</volume>, <fpage>263</fpage>&#x2013;<lpage>273</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/s41588-021-00997-7</pub-id>
</citation>
</ref>
<ref id="B105">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Akey</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Zhang</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Chakraborty</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Jin</surname> <given-names>L.</given-names>
</name>
</person-group> (<year>2002</year>). <article-title>Distribution of recombination crossovers and the origin of haplotype blocks: the interplay of population history, recombination, and mutation</article-title>. <source>Am. J. Hum. Genet.</source> <volume>71</volume>, <fpage>1227</fpage>&#x2013;<lpage>1234</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1086/344398</pub-id>
</citation>
</ref>
<ref id="B106">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname> <given-names>F.</given-names>
</name>
<name>
<surname>Moon</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Letsou</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Sapkota</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Im</surname> <given-names>C.</given-names>
</name>
<etal/>
</person-group>. (<year>2023</year>). <article-title>Genome-wide analysis of rare haplotypes associated with breast cancer risk</article-title>. <source>Cancer Res.</source> <volume>83</volume>, <fpage>332</fpage>&#x2013;<lpage>345</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1158/0008-5472.CAN-22-1888</pub-id>
</citation>
</ref>
<ref id="B107">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Werner</surname> <given-names>C. R.</given-names>
</name>
<name>
<surname>Gaynor</surname> <given-names>R. C.</given-names>
</name>
<name>
<surname>Gorjanc</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Hickey</surname> <given-names>J. M.</given-names>
</name>
<name>
<surname>Kox</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Abbadi</surname> <given-names>A.</given-names>
</name>
<etal/>
</person-group>. (<year>2020</year>). <article-title>How population structure impacts genomic selection accuracy in cross-validation: implications for practical breeding</article-title>. <source>Front. Plant Sci.</source> <volume>11</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2020.592977</pub-id>
</citation>
</ref>
<ref id="B108">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Werner</surname> <given-names>C. R.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>L.</given-names>
</name>
<name>
<surname>Voss-Fels</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Abbadi</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Leckband</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Frisch</surname> <given-names>M.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>a). <article-title>Genome-wide regression models considering general and specific combining ability predict hybrid performance in oilseed rape with similar accuracy regardless of trait architecture</article-title>. <source>Theor. Appl. Genet.</source> <volume>131</volume>, <fpage>299</fpage>&#x2013;<lpage>317</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1007/s00122-017-3002-5</pub-id>
</citation>
</ref>
<ref id="B109">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Werner</surname> <given-names>C. R.</given-names>
</name>
<name>
<surname>Voss-Fels</surname> <given-names>K. P.</given-names>
</name>
<name>
<surname>Miller</surname> <given-names>C. N.</given-names>
</name>
<name>
<surname>Qian</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Hua</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Guan</surname> <given-names>C.-Y.</given-names>
</name>
<etal/>
</person-group>. (<year>2018</year>b). <article-title>Effective genomic selection in a narrow-genepool crop with low-density markers: Asian rapeseed as an example</article-title>. <source>Plant Genome</source> <volume>11</volume>, <fpage>170084</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.3835/plantgenome2017.09.0084</pub-id>
</citation>
</ref>
<ref id="B110">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolc</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Stricker</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Arango</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Settar</surname> <given-names>P.</given-names>
</name>
<name>
<surname>Fulton</surname> <given-names>J. E.</given-names>
</name>
<name>
<surname>O&#x2019;Sullivan</surname> <given-names>N. P.</given-names>
</name>
<etal/>
</person-group>. (<year>2011</year>). <article-title>Breeding value prediction for production traits in layer chickens using pedigree or genomic relationships in a reduced animal model</article-title>. <source>Genet. Selection Evol.</source> <volume>43</volume>, <elocation-id>5</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/1297-9686-43-5</pub-id>
</citation>
</ref>
<ref id="B111">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wood</surname> <given-names>A. R.</given-names>
</name>
<name>
<surname>Tuke</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Nalls</surname> <given-names>M. A.</given-names>
</name>
<name>
<surname>Hernandez</surname> <given-names>D. G.</given-names>
</name>
<name>
<surname>Bandinelli</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Singleton</surname> <given-names>A. B.</given-names>
</name>
<etal/>
</person-group>. (<year>2014</year>). <article-title>Another explanation for apparent epistasis</article-title>. <source>Nature</source> <volume>514</volume>, <fpage>E3</fpage>&#x2013;<lpage>E5</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/nature13691</pub-id>
</citation>
</ref>
<ref id="B112">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>W&#xfc;rschum</surname> <given-names>T.</given-names>
</name>
<name>
<surname>Abel</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Zhao</surname> <given-names>Y.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Potential of genomic selection in rapeseed (Brassica napus L.) breeding</article-title>. <source>Plant Breed.</source> <volume>133</volume>, <fpage>45</fpage>&#x2013;<lpage>51</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1111/pbr.12137</pub-id>
</citation>
</ref>
<ref id="B113">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ye</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Gao</surname> <given-names>N.</given-names>
</name>
<name>
<surname>Zheng</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Chen</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Teng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Yuan</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2019</year>). <article-title>Strategies for obtaining and pruning imputed whole-genome sequence data for genomic prediction</article-title>. <source>Front. Genet.</source> <volume>10</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fgene.2019.00673</pub-id>
</citation>
</ref>
<ref id="B114">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yu</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Xie</surname> <given-names>W.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Xing</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Xu</surname> <given-names>C.</given-names>
</name>
<name>
<surname>Li</surname> <given-names>X.</given-names>
</name>
<etal/>
</person-group>. (<year>2011</year>). <article-title>Gains in QTL detection using an ultra-high density SNP map based on population sequencing relative to traditional RFLP/SSR markers</article-title>. <source>PLoS One</source> <volume>6</volume>, <elocation-id>e17595</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1371/journal.pone.0017595</pub-id>
</citation>
</ref>
<ref id="B115">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Z.</given-names>
</name>
<name>
<surname>Ersoz</surname> <given-names>E.</given-names>
</name>
<name>
<surname>Lai</surname> <given-names>C.-Q.</given-names>
</name>
<name>
<surname>Todhunter</surname> <given-names>R. J.</given-names>
</name>
<name>
<surname>Tiwari</surname> <given-names>H. K.</given-names>
</name>
<name>
<surname>Gore</surname> <given-names>M. A.</given-names>
</name>
<etal/>
</person-group>. (<year>2010</year>). <article-title>Mixed linear model approach adapted for genome-wide association studies</article-title>. <source>Nat. Genet.</source> <volume>42</volume>, <fpage>355</fpage>&#x2013;<lpage>360</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1038/ng.546</pub-id>
</citation>
</ref>
<ref id="B116">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>Q.</given-names>
</name>
<name>
<surname>Sahana</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Su</surname> <given-names>G.</given-names>
</name>
<name>
<surname>Guldbrandtsen</surname> <given-names>B.</given-names>
</name>
<name>
<surname>Lund</surname> <given-names>M. S.</given-names>
</name>
<name>
<surname>Calus</surname> <given-names>M. P. L.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Impact of rare and low-frequency sequence variants on reliability of genomic prediction in dairy cattle</article-title>. <source>Genet. Selection Evol.</source> <volume>50</volume>, <fpage>62</fpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.1186/s12711-018-0432-8</pub-id>
</citation>
</ref>
<ref id="B117">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname> <given-names>A.</given-names>
</name>
<name>
<surname>Wang</surname> <given-names>H.</given-names>
</name>
<name>
<surname>Beyene</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Semagn</surname> <given-names>K.</given-names>
</name>
<name>
<surname>Liu</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Cao</surname> <given-names>S.</given-names>
</name>
<etal/>
</person-group>. (<year>2017</year>). <article-title>Effect of trait heritability, training population size and marker density on genomic prediction accuracy estimation in 22 bi-parental tropical maize populations</article-title>. <source>Front. Plant Sci.</source> <volume>8</volume>. doi:&#xa0;<pub-id pub-id-type="doi">10.3389/fpls.2017.01916</pub-id>
</citation>
</ref>
<ref id="B118">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhao</surname> <given-names>Y.</given-names>
</name>
<name>
<surname>Zeng</surname> <given-names>J.</given-names>
</name>
<name>
<surname>Fernando</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Reif</surname> <given-names>J. C.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Genomic prediction of hybrid wheat performance</article-title>. <source>Crop Sci.</source> <volume>53</volume>, <fpage>802</fpage>&#x2013;<lpage>810</lpage>. doi:&#xa0;<pub-id pub-id-type="doi">10.2135/cropsci2012.08.0463</pub-id>
</citation>
</ref>
<ref id="B119">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zimin</surname> <given-names>A. V.</given-names>
</name>
<name>
<surname>Puiu</surname> <given-names>D.</given-names>
</name>
<name>
<surname>Hall</surname> <given-names>R.</given-names>
</name>
<name>
<surname>Kingan</surname> <given-names>S.</given-names>
</name>
<name>
<surname>Clavijo</surname> <given-names>B. J.</given-names>
</name>
<name>
<surname>Salzberg</surname> <given-names>S. L.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>The first near-complete assembly of the hexaploid bread wheat genome, Triticum aestivum</article-title>. <source>GigaScience</source> <volume>6</volume>, <elocation-id>gix097</elocation-id>. doi:&#xa0;<pub-id pub-id-type="doi">10.1093/gigascience/gix097</pub-id>
</citation>
</ref>
</ref-list>
</back>
</article>