<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="brief-report" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">856872</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.856872</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Brief Research Report</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Linear Mixed-Effect Models Through the Lens of Hardy&#x2013;Weinberg Disequilibrium</article-title>
<alt-title alt-title-type="left-running-head">Zhang&#x2009; and Sun&#x2009;</alt-title>
<alt-title alt-title-type="right-running-head">LMM Through the Lens of HWD</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Zhang&#x2009;</surname>
<given-names>Lin</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1677850/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Sun&#x2009;</surname>
<given-names>Lei</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/40108/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<sup>1</sup>
<institution>Department of Statistical Sciences</institution>, <institution>University of Toronto</institution>, <addr-line>Toronto</addr-line>, <addr-line>ON</addr-line>, <country>Canada</country>
</aff>
<aff id="aff2">
<sup>2</sup>
<institution>Division of Biostatistics</institution>, <institution>Dalla Lana School of Public Health</institution>, <institution>University of Toronto</institution>, <addr-line>Toronto</addr-line>, <addr-line>ON</addr-line>, <country>Canada</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/744713/overview">Lide Han</ext-link>, Vanderbilt University Medical Center, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/32750/overview">Guo-Bo Chen</ext-link>, Zhejiang Provincial People&#x2019;s Hospital, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/891131/overview">Chengsong Zhu</ext-link>, University of Texas Southwestern Medical Center, United States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Lei Sun&#x2009;, <email>sun@utstat.toronto.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>04</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>856872</elocation-id>
<history>
<date date-type="received">
<day>17</day>
<month>01</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>09</day>
<month>03</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Zhang&#x2009; and Sun&#x2009;.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhang&#x2009; and Sun&#x2009;</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>For genetic association studies with related individuals, the linear mixed-effect model is the most commonly used method. In this report, we show that contrary to the popular belief, this standard method can be sensitive to departure from Hardy&#x2013;Weinberg equilibrium (i.e., Hardy&#x2013;Weinberg disequilibrium) at the causal SNPs in two ways. First, when the trait heritability is treated as a nuisance parameter, although the association test has correct type I error control, the resulting heritability estimate can be biased, often upward, in the presence of Hardy&#x2013;Weinberg disequilibrium. Second, if the true heritability is used in the linear mixed-effect model, then the corresponding association test can be biased in the presence of Hardy&#x2013;Weinberg disequilibrium. We provide some analytical insights along with supporting empirical results from simulation and application studies.</p>
</abstract>
<kwd-group>
<kwd>genome-wide association study</kwd>
<kwd>dependent sample</kwd>
<kwd>robust association analysis</kwd>
<kwd>heritability estimate</kwd>
<kwd>Hardy&#x2013;Weinberg equilibrium</kwd>
</kwd-group>
<contract-num rid="cn001">RGPIN-04934 RGPAS-522594</contract-num>
<contract-sponsor id="cn001">Natural Sciences and Engineering Research Council of Canada<named-content content-type="fundref-id">10.13039/501100000038</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Genetic association tests are often derived from a regression model, regressing the phenotypic data of a complex trait (<italic>Y</italic>) on the genotypic data of a single-nucleotide polymorphism (SNP; <italic>G</italic>), as well as on the covariate data of important environmental factors (<italic>Z</italic>). When individuals in a sample are genetically related with each other, the linear mixed-effect model (LMM) is the most commonly used method for genome-wide association studies (GWAS) (<xref ref-type="bibr" rid="B6">Eu-Ahsunthornwattana et al., 2014</xref>). The variance&#x2013;covariance matrix of the regression model is partitioned into a weighted sum of the genetic correlation matrix and the correlation matrix due to shared environmental effects. The genetic correlation matrix is typically represented by the kinship coefficient matrix, which is either inferred from the (correctly) known pedigree structure or estimated based on the available genome-wide genetic data (<xref ref-type="bibr" rid="B32">Yang et al., 2011</xref>; <xref ref-type="bibr" rid="B5">Dimitromanolakis et al., 2019</xref>). The weight for the genetic correlation matrix is referred to as the heritability of the trait (<xref ref-type="bibr" rid="B24">Visscher et al., 2006</xref>; <xref ref-type="bibr" rid="B23">Visscher et al., 2008</xref>); <xref ref-type="bibr" rid="B7">Falconer (1985)</xref> gave a theoretical modeling of the variance partition, which sets the foundation for heritability.</p>
<p>It is commonly assumed that these regression-based association tests are robust to departure from Hardy&#x2013;Weinberg equilibrium (HWE) (<xref ref-type="bibr" rid="B16">Sasieni, 1997</xref>). HWE states that the two alleles in a genotype are independent draws from the same Bernoulli distribution, or, equivalently, genotype frequencies depend solely on the allele frequencies (<xref ref-type="bibr" rid="B8">Hardy et al., 1908</xref>; <xref ref-type="bibr" rid="B26">Weinberg, 1908</xref>). For a biallelic SNP with two possible alleles <italic>A</italic> and <italic>a</italic>, let <italic>p</italic> and 1 &#x2212; <italic>p</italic> be the population allele frequencies, respectively. Under HWE, <italic>p</italic>
<sub>
<italic>aa</italic>
</sub> &#x3d; (1 &#x2212; <italic>p</italic>)<sup>2</sup>, <italic>p</italic>
<sub>
<italic>Aa</italic>
</sub> &#x3d; 2<italic>p</italic>(1 &#x2212; <italic>p</italic>), and <italic>p</italic>
<sub>
<italic>AA</italic>
</sub> &#x3d; <italic>p</italic>
<sup>2</sup>, where <italic>p</italic>
<sub>
<italic>aa</italic>
</sub>, <italic>p</italic>
<sub>
<italic>Aa</italic>
</sub>, and <italic>p</italic>
<sub>
<italic>AA</italic>
</sub> are the population genotype frequencies of genotypes <italic>aa</italic>, <italic>Aa</italic>, and <italic>AA</italic>, respectively. To quantify the departure from HWE or the amount of Hardy&#x2013;Weinberg disequilibrium (HWD),<disp-formula id="e1">
<mml:math id="m1">
<mml:mi>&#x3b4;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(1)</label>
</disp-formula>is a widely used measure (<xref ref-type="bibr" rid="B27">Weir, 1996</xref>), and <italic>&#x3b4;</italic> &#x3d; 0 indicates HWE holds. We note that a) HWE is also known as Hardy&#x2013;Weinberg proportion and b) <italic>&#x3b4;</italic> is also known as <italic>p</italic>(1 &#x2212; <italic>p</italic>)<italic>F</italic>, where <italic>F</italic> is the inbreeding coefficient (<xref ref-type="bibr" rid="B14">Powell et al., 2010</xref>). Equivalently, instead of quantifying the genotype frequencies as <italic>p</italic>
<sub>
<italic>aa</italic>
</sub> &#x3d; (1 &#x2212; <italic>p</italic>)<sup>2</sup> &#x2b; <italic>&#x3b4;</italic>, <italic>p</italic>
<sub>
<italic>Aa</italic>
</sub> &#x3d; 2<italic>p</italic>(1 &#x2212; <italic>p</italic>) &#x2212; 2<italic>&#x3b4;</italic>, and <italic>p</italic>
<sub>
<italic>AA</italic>
</sub> &#x3d; <italic>p</italic>
<sup>2</sup> &#x2b; <italic>&#x3b4;</italic> based on <italic>&#x3b4;</italic> (<xref ref-type="bibr" rid="B27">Weir, 1996</xref>), we can define them based on <italic>F</italic> as <italic>p</italic>
<sub>
<italic>aa</italic>
</sub> &#x3d; (1 &#x2212; <italic>p</italic>)<sup>2</sup> &#x2b; <italic>p</italic>(1 &#x2212; <italic>p</italic>)<italic>F</italic>, <italic>p</italic>
<sub>
<italic>Aa</italic>
</sub> &#x3d; 2<italic>p</italic>(1 &#x2212; <italic>p</italic>)(1 &#x2212; <italic>F</italic>), and <italic>p</italic>
<sub>
<italic>AA</italic>
</sub> &#x3d; <italic>p</italic>
<sup>2</sup> &#x2b; <italic>p</italic>(1 &#x2212; <italic>p</italic>)<italic>F</italic> (<xref ref-type="bibr" rid="B14">Powell et al., 2010</xref>). As the classical Pearson <italic>&#x3c7;</italic>
<sup>2</sup> HWE testing is based on comparing the observed genotype counts with the expected under HWE (<xref ref-type="bibr" rid="B34">Zhang and Sun, 2021</xref>); we thus chose <italic>&#x3b4;</italic> for this work to be consistent with the GWAS literature.</p>
<p>A truly associated or causal SNP can be out of HWE (<xref ref-type="bibr" rid="B30">Wittke-Thompson et al., 2005</xref>; <xref ref-type="bibr" rid="B15">Ryckman and Williams, 2008</xref>; <xref ref-type="bibr" rid="B21">Turner et al., 2011</xref>), which is often overlooked but an important consideration when studying a method&#x2019;s robustness to HWD. Note that the HWD attributed to true association is typically not as extreme as the HWD caused by genotyping errors (<xref ref-type="bibr" rid="B35">Zhang and Sun, 2020</xref>). Thus, true HWD can remain in a &#x201c;cleaned&#x201d; dataset after applying the standard HWD-based quality control screening using a stringent <italic>p</italic>-value threshold [e.g., 10<sup>&#x2013;12</sup> for an application of the UK Biobank data by <xref ref-type="bibr" rid="B1">Bycroft et al. (2018)</xref>]. With a sample of independent individuals, both theoretical and empirical results support that genotype-based association tests are robust to HWD (<xref ref-type="bibr" rid="B16">Sasieni, 1997</xref>; <xref ref-type="bibr" rid="B17">Schaid and Jacobsen, 1999</xref>; <xref ref-type="bibr" rid="B33">Zhang, 2021</xref>). However, in the presence of sample dependency, little has been discussed.</p>
<p>In this report, we first provide some analytical insights on why the standard LMM can be sensitive to HWD in pedigree data in contrast to when analyzing a sample of unrelated individuals. We then demonstrate with a simple sib-pair design that 1) when the heritability is estimated from the data as in practice, although the empirical type I error rate of the LMM is well controlled, the estimated heritability is biased, often upward biased; 2) when the true heritability is known and used, the empirical type I error rate of the LMM is then inflated when <italic>&#x3b4;</italic> &#x3e; 0, and deflated if <italic>&#x3b4;</italic> &#x3c; 0. The result of 2) is novel, but it is mostly of an academic interest as the true heritability of a trait is often unknown in practice. On the other hand, the result of 1) has important practical implications because if the estimate of a trait heritability is larger than the true value, then it helps explain some of the &#x201c;missing heritability&#x201d; (<xref ref-type="bibr" rid="B10">Manolio et al., 2009</xref>); the insightful work of <xref ref-type="bibr" rid="B3">Chen (2014)</xref> &#x201c;discuss[es] the circumstances in which the HE [Haseman&#x2010;Elston] regression and the mixed linear model are equivalent.&#x201d;</p>
</sec>
<sec id="s2">
<title>2 Methods</title>
<sec id="s2-1">
<title>2.1 Traditional <italic>Y</italic> &#x223c; <italic>G</italic> Model With Independent Samples, <italic>T</italic> <sub>Indep</sub>, Is Robust to HWD</title>
<p>Let <italic>Y</italic> be a (continuous) trait of interest, and <italic>G</italic> &#x3d; 0, 1, and 2, respectively, for the genotypes <italic>aa</italic>, <italic>Aa</italic>, and <italic>AA</italic> of a SNP. Additionally, for notation simplicity but without loss of generality, we assume that there is only one additional covariate, denoted by <italic>Z</italic>. With a sample of <italic>n</italic> unrelated individuals, the traditional genotype-based association analysis assumes that<disp-formula id="e2">
<mml:math id="m2">
<mml:mi>y</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mtext>g</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mspace width="8.5359pt"/>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>0</mml:mtext>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi>I</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(2)</label>
</disp-formula>where y &#x3d; (<italic>y</italic>
<sub>1</sub>, <italic>y</italic>
<sub>2</sub>,&#x2026;, <italic>y</italic>
<sub>
<italic>n</italic>
</sub>) is a <italic>n</italic> &#xd7; 1 vector for the phenotypic values, 1 is a <italic>n</italic> &#xd7; 1 vector of 1&#x2019;s, g &#x3d; (<italic>g</italic>
<sub>1</sub>, <italic>g</italic>
<sub>2</sub>,&#x2026;, <italic>g</italic>
<sub>
<italic>n</italic>
</sub>) is a <italic>n</italic> &#xd7; 1 vector for the genotypes of the SNP, z &#x3d; (<italic>z</italic>
<sub>1</sub>, <italic>z</italic>
<sub>2</sub>,&#x2026;, <italic>z</italic>
<sub>
<italic>n</italic>
</sub>) is a <italic>n</italic> &#xd7; 1 vector for the covariate values, <italic>&#x3f5;</italic>&#x2a; is the error term with variance <italic>&#x3c3;</italic>&#x2a;<sup>2</sup>, and <italic>I</italic> is the identity matrix.</p>
<p>Score-based tests are often used for genetic association analyses (<xref ref-type="bibr" rid="B4">Derkach et al., 2015</xref>). In this case, the score statistic of testing <italic>H</italic>
<sub>0</sub>: <italic>&#x3b2;</italic>&#x2a; &#x3d; 0 can be easily derived as<disp-formula id="e3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>&#x2009;indep</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mo>&#x22c5;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="[" close="]">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="{" close="}">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:math>
<label>(3)</label>
</disp-formula>To observe <italic>T</italic> <sub>indep</sub>&#x2019;s connection with Hardy&#x2013;Weinberg disequilibrium, it is instructive to employ some algebraic tricks and show that<disp-formula id="equ1">
<mml:math id="m4">
<mml:mfrac>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>T</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mtext>var</mml:mtext>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
</disp-formula>
</p>
<p>Because <inline-formula id="inf1">
<mml:math id="m5">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>A</mml:mi>
<mml:mi>A</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> measures the amount of HWD present in the data (<xref ref-type="bibr" rid="B27">Weir, 1996</xref>), <italic>T</italic> <sub>indep</sub> inherently adjusts for departure from HWE through <inline-formula id="inf2">
<mml:math id="m6">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mtext>var</mml:mtext>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. As a result, the traditional genotype-based association test is robust to HWD in independent samples.</p>
<p>When <italic>Y</italic> is binary, the classic logistic regression is commonly used. However, <xref ref-type="bibr" rid="B2">Chen (1983)</xref> showed that under some regularity conditions, the score test statistics have an identical form for the exponential family in independent samples, which was recently validated by <xref ref-type="bibr" rid="B34">Zhang and Sun (2021)</xref> for genetic association studies. Additionally, <xref ref-type="bibr" rid="B4">Derkach et al. (2015)</xref> showed that for <italic>Y</italic>-dependent sampling, &#x201c;the score statistics are identical for conditional and full likelihood approaches, and are of the same form as those for ordinary random sampling.&#x201d; Thus, in terms of association testing (not genetic effect estimation), we can conclude that genotype-based association studies of binary traits in independent samples are also robust to HWD.</p>
</sec>
<sec id="s2-2">
<title>2.2 Linear Mixed-Effect Model With Dependent Samples, <italic>T</italic> <sub>LMM</sub>, Can Be Sensitive to HWD</title>
<p>Although a pedigree-based study design is rare for genome-wide association studies, individuals can be (cryptically) related with each other even in population-based GWAS (<xref ref-type="bibr" rid="B19">Sun et al., 2017</xref>). Omitting related individuals simplifies the association analysis but reduces the sample size and thus power. Instead, &#x3a3;<sub>&#x3a6;</sub>, the kinship coefficient matrix, can be estimated using the available genome-wide data to capture the sample relatedness between the <italic>n</italic> individuals (<xref ref-type="bibr" rid="B24">Visscher et al., 2006</xref>; <xref ref-type="bibr" rid="B32">Yang et al., 2011</xref>). The association analysis using the full sample can be conducted using the linear mixed-effect model.<disp-formula id="e4">
<mml:math id="m7">
<mml:mtext>y</mml:mtext>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mtext>g</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3b3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mtext>&#x2009;where&#x2009;</mml:mtext>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2217;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>0</mml:mtext>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="0.22em"/>
<mml:mspace width="0.22em"/>
<mml:mtext>and</mml:mtext>
<mml:mspace width="0.22em"/>
<mml:mspace width="0.22em"/>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mi>I</mml:mi>
<mml:mo>.</mml:mo>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>Compared with the linear model used for independent samples, var(<italic>&#x3f5;</italic>&#x2a;) &#x3d; <italic>&#x3c3;</italic>&#x2a;<sup>2</sup>
<italic>I</italic> in <xref ref-type="disp-formula" rid="e2">Eq. 2</xref> is replaced by <inline-formula id="inf3">
<mml:math id="m8">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> to reflect the sample dependence. The matrix &#x3a3;<sub>
<italic>y</italic>
</sub> is a weighted average of two components, where &#x3a3;<sub>&#x3a6;</sub> reflects the sample relatedness; naturally, the model is reduced to the linear model of <xref ref-type="disp-formula" rid="e2">Eq. 2</xref> for independent samples when &#x3a3;<sub>&#x3a6;</sub> &#x3d; <italic>I</italic>. The weight <italic>h</italic>
<sup>2</sup> is interpreted as the heritability of the trait (<xref ref-type="bibr" rid="B23">Visscher et al., 2008</xref>), <inline-formula id="inf4">
<mml:math id="m9">
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> as the phenotypic variation due to (additive) genetic variation, and <inline-formula id="inf5">
<mml:math id="m10">
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> as the phenotypic variation due to environmental variation. The matrix &#x3a3;<sub>&#x3a6;</sub> is the kinship matrix, where &#x3a3;<sub>&#x3a6;</sub>(<italic>i</italic>, <italic>j</italic>) &#x3d; 2<italic>&#x3d5;</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> and <italic>&#x3d5;</italic>
<sub>
<italic>i</italic>,<italic>j</italic>
</sub> is the kinship coefficient between the <italic>i</italic>th and <italic>j</italic>th samples.</p>
<p>By convention, <italic>h</italic>
<sup>2</sup> is defined as<disp-formula id="equ2">
<mml:math id="m11">
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mtext>var</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>G</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mtext>var</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>Y</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
</disp-formula>where there could be multiple causal SNPs, <italic>k</italic> &#x3d; 1,&#x2026;, <italic>S</italic>. In reality, <italic>h</italic>
<sup>2</sup> is estimated by the correlation between phenotypes of related individuals. Consider the simple case of sibling pairs, and let <italic>Y</italic>
<sub>1</sub> and <italic>Y</italic>
<sub>2</sub> be the phenotypes for sib 1 and sib 2, respectively. Allowing for HWD and adjusting for the kinship coefficient <italic>&#x3d5;</italic>, the estimated <italic>h</italic>
<sup>2</sup> is<disp-formula id="equ3">
<mml:math id="m12">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mtext>corr</mml:mtext>
</mml:mrow>
<mml:mo>&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>/</mml:mo>
<mml:mn>2</mml:mn>
<mml:mi>&#x3d5;</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
</disp-formula>where corr(<italic>Y</italic>
<sub>1</sub>, <italic>Y</italic>
<sub>2</sub>) depends on the correlation between <italic>G</italic>
<sub>1<italic>k</italic>
</sub> and <italic>G</italic>
<sub>2<italic>k</italic>
</sub> between the siblings; see <xref ref-type="bibr" rid="B34">Zhang and Sun (2021)</xref> for the derivation of corr(<italic>G</italic>
<sub>1<italic>k</italic>
</sub>, <italic>G</italic>
<sub>2<italic>k</italic>
</sub>) accounting for kinship coefficient and HWD. Thus,<disp-formula id="equ4">
<mml:math id="m13">
<mml:mfrac>
<mml:mrow>
<mml:mi>E</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
</disp-formula>and the bias of the <italic>h</italic>
<sup>2</sup> estimate is<disp-formula id="e5">
<mml:math id="m14">
<mml:mi>E</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x22c5;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mo>.</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>Under the simple case of one causal SNP, the bias is simplified to <italic>h</italic>
<sup>2</sup> &#x22c5; <italic>&#x3b4;</italic>/(<italic>p</italic>(1 &#x2212; <italic>p</italic>)).</p>
<p>Given the analytical insights provided so far, we then briefly examine the empirical properties of <italic>T</italic> <sub>LMM</sub> through both application and simulation studies.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Cystic Fibrosis Sib-Pair Data Application: <italic>T</italic>
<sub>LMM</sub> Has Correct Type I Error but <italic>h</italic>
<sup>2</sup> Appears to Be Overestimated</title>
<p>We extracted 65 sibling pairs from a cystic fibrosis (CF) gene modifier study (<xref ref-type="bibr" rid="B31">Wright et al., 2011</xref>; <xref ref-type="bibr" rid="B20">Sun et al., 2012</xref>). The phenotype <italic>Y</italic> of interest is the lung function measurements of the 130 related individuals with CF. In total, there were 570,539 SNPs genotyped using the Illumina 610-Quad Beadchip after applying the standard quality control, including minor allele frequency (MAF) greater than 2%. To stabilize the variance estimation, we additionally required SNPs to have MAF greater than 5%. We then applied <italic>T</italic> <sub>LMM</sub> to the remaining 505,172 SNPs. In the application, we treated <italic>h</italic>
<sup>2</sup> as unknown and estimated it based on the linear mixed-effect model of <xref ref-type="disp-formula" rid="e4">Eq. 4</xref> as in convention.</p>
<p>When <italic>h</italic>
<sup>2</sup> was estimated from the data, our association testing based on <italic>T</italic> <sub>LMM</sub> had good type I error control (results not shown), consistent with the empirical observations in the GWAS literature. However, the estimated <italic>h</italic>
<sup>2</sup>, obtained using the 65-pair sibling data, is <inline-formula id="inf6">
<mml:math id="m15">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.82</mml:mn>
</mml:math>
</inline-formula>. This value is substantially greater than 0.5, the commonly believed &#x201c;true&#x201d; heritability of lung function in CF obtained from the classic monozygous (MZ) vs. dizygous (DZ) twin-based estimation method (<xref ref-type="bibr" rid="B22">Vanscoy et al., 2007</xref>).</p>
<p>To verify if the large heritability estimate from the LMM method in our application was due to chance, we conducted a proof-of-principle simulation study. We assumed that only one causal SNP, <italic>G</italic>
<sub>causal</sub> with MAF of 0.2, affects <italic>Y</italic> with <italic>h</italic>
<sup>2</sup> &#x3d; 0.5. Genotype and phenotype values for 65 sibling pairs were then simulated under the assumption of HWE (i.e., without HWD). Among the 100,000 independently simulated replicates, only 4.24<italic>%</italic> of the heritability estimates were greater than <inline-formula id="inf7">
<mml:math id="m16">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.82</mml:mn>
</mml:math>
</inline-formula>. This suggests that <inline-formula id="inf8">
<mml:math id="m17">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.82</mml:mn>
</mml:math>
</inline-formula>, the value that was observed in the CF data application, was unlikely if the true heritability was 0.5 and without HWD at the causal SNP.</p>
<p>To verify if HWD at the causal SNP can lead to a biased heritability estimate, we then conducted additional simulation studies, following the same sib-pair design as mentioned previously. Our goal is to demonstrate that 1) when <italic>h</italic>
<sup>2</sup> is treated as a nuisance parameter, its estimate based on model (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) cross-reference can be biased in the presence of HWD; and 2) assuming the true <italic>h</italic>
<sup>2</sup> is known, the empirical type I error rate of LMM (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) cross-reference inflates when <italic>&#x3b4;</italic> &#x3e; 0 and deflates when <italic>&#x3b4;</italic> &#x3c; 0.</p>
</sec>
<sec id="s3-2">
<title>3.2 Simulated Sib-Pair Data in the Presence of HWD: <italic>h</italic>
<sup>2</sup> Estimate Is Biased</title>
<p>Consider a continuous trait <italic>Y</italic> with <italic>h</italic>
<sup>2</sup> &#x3d; 0.5 and influenced by one causal SNP, <italic>G</italic>
<sub>causal</sub>, with minor allele frequency of 0.2 and with HWD factor, <italic>&#x3b4;</italic>
<sub>causal</sub>, ranging from &#x2212;0.04 to 0.16. A non-associated SNP, <italic>G</italic>
<sub>tested</sub>, also has an MAF of 0.2 but with its own <italic>&#x3b4;</italic>
<sub>tested</sub>, which may not be the same as <italic>&#x3b4;</italic> <sub>causal</sub> in a specific simulation study. The sample size was 65 sibling pairs, chosen to match with the sample size of the cystic fibrosis application study in <xref ref-type="sec" rid="s3-1">Section 3.1</xref>.</p>
<p>Most practical implementations of the linear mixed-effect model (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) cross-reference treat <italic>h</italic>
<sup>2</sup> as a nuisance parameter, and no type I error issue has been reported. Indeed, when <italic>h</italic>
<sup>2</sup> was estimated in our simulation study conducted in <xref ref-type="sec" rid="s3-3">Section 3.3</xref>, the test size of <italic>T</italic> <sub>LMM</sub> was correct at the nominal level (black squares in <xref ref-type="fig" rid="F3">Figure 3</xref> shown in <xref ref-type="sec" rid="s3-3">Section 3.3</xref>) even if <italic>&#x3b4;</italic>
<sub>tested</sub> &#x2260; 0 (i.e., out of HWE) and across the range of <italic>&#x3b4;</italic>
<sub>causal</sub> values (from &#x2212;0.04 to 0.16).</p>
<p>However, in this situation, when <italic>h</italic>
<sup>2</sup> is treated as unknown, the impact of HWD is on the estimation of <italic>h</italic>
<sup>2</sup>. Specifically, <xref ref-type="fig" rid="F1">Figure 1</xref> shows that <inline-formula id="inf9">
<mml:math id="m18">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> is downward biased when <italic>&#x3b4;</italic>
<sub>causal</sub> &#x3c; 0, and upward biased if <italic>&#x3b4;</italic>
<sub>causal</sub> &#x3e; 0. The bias can be substantial. For example, when <italic>&#x3b4;</italic> <sub>causal</sub> &#x3d; 0.10, the estimated heritability <inline-formula id="inf10">
<mml:math id="m19">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> is centered at 0.80 as compared to the true value of 0.5, with a bias of 0.30. Indeed, based on our theoretical insight in <xref ref-type="sec" rid="s2-2">Section 2.2</xref>, the expected bias is <italic>h</italic>
<sup>2</sup> &#x22c5; <italic>&#x3b4;</italic>/(<italic>p</italic>(1 &#x2212; <italic>p</italic>)) &#x3d; 0.5 &#x22c5; 0.1/(0.2(1&#x2013;0.2)) &#x3d; 0.31.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Box plots of <inline-formula id="inf11">
<mml:math id="m20">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, estimated from the linear mixed-effect model (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) against <bold>
<italic>&#x3b4;</italic>
</bold>
<sub>
<bold>causal</bold>
</sub>. The true heritability of the phenotype is <italic>h</italic>
<sup>2</sup> &#x3d; 0.5. The minor allele frequencies <italic>p</italic>
<sub>causal</sub> &#x3d; <italic>p</italic>
<sub>tested</sub> &#x3d; 0.2 and 10,000 independent replicates of phenotypes and genotypes for 65 sibling pairs were simulated for each <italic>&#x3b4;</italic>
<sub>causal</sub> value. The empirical type 1 error rates are shown in <xref ref-type="fig" rid="F3">Figure 3</xref> as black squares.</p>
</caption>
<graphic xlink:href="fgene-13-856872-g001.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F1">Figure 1</xref>, it is notable that <inline-formula id="inf12">
<mml:math id="m21">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> can be greater than one. Since <italic>h</italic>
<sup>2</sup> is the proportion of variance in <italic>Y</italic> explained by additive genetic variation, 0 &#x2264; <italic>h</italic>
<sup>2</sup> &#x2264; 1 by definition. However, if <italic>&#x3b4;</italic>
<sub>causal</sub> &#x2260; 0, <inline-formula id="inf13">
<mml:math id="m22">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> based on the LMM, without additional truncation, is a biased estimate of <italic>h</italic>
<sup>2</sup> with a bias of <italic>h</italic>
<sup>2</sup> &#x22c5; <italic>&#x3b4;</italic>/(<italic>p</italic>(1 &#x2212; <italic>p</italic>)) for this sib-pair design as shown in <xref ref-type="sec" rid="s2-2">Section 2.2</xref>; the bias is 0 (i.e., no bias) under HWE when <italic>&#x3b4;</italic> &#x3d; 0.</p>
<p>Additionally, although a larger sample that consists of 5,000 sibling pairs shrinks the variance of the <italic>h</italic>
<sup>2</sup> estimate as expected, it does not shrink the bias, as shown in <xref ref-type="fig" rid="F2">Figure 2</xref>. However, we also note that, in practice, it is unlikely to have so many sibling pairs.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Box plots of <inline-formula id="inf14">
<mml:math id="m23">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, estimated from the linear mixed-effect model (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) against <italic>&#x3b4;</italic>
<sub>causal</sub> using 5,000 sibling pairs. The other parameter values are the same as in <xref ref-type="fig" rid="F1">Figure 1</xref>, where the true heritability of the phenotype is <italic>h</italic>
<sup>2</sup> &#x3d; 0.5, the minor allele frequencies <italic>p</italic>
<sub>causal</sub> &#x3d; <italic>p</italic>
<sub>tested</sub> &#x3d; 0.2, and 10,000 independent replicates were simulated for each <italic>&#x3b4;</italic>
<sub>causal</sub> value.</p>
</caption>
<graphic xlink:href="fgene-13-856872-g002.tif"/>
</fig>
</sec>
<sec id="s3-3">
<title>3.3 Simulated Sib-Pair Data in the Presence of HWD: When Using the True <italic>h</italic>
<sup>2</sup> Value <italic>T</italic> <sub>LMM</sub> Has Incorrect Test Size</title>
<p>Here, we conducted the association analysis between <italic>Y</italic> and the non-associated SNP, <italic>G</italic>
<sub>tested</sub>, using the LMM model of <xref ref-type="disp-formula" rid="e4">Eq. 4</xref> but assuming <italic>h</italic>
<sup>2</sup> &#x3d; 0.5 is known.</p>
<p>
<xref ref-type="fig" rid="F3">Figure 3A</xref> plots the empirical type I error rates (blue circles) of <italic>T</italic> <sub>LMM</sub> using the true <italic>h</italic>
<sup>2</sup> &#x3d; 0.5, for a nominal level of 0.05, estimated from independently simulated 10,000 replicates for each <italic>&#x3b4;</italic>
<sub>causal</sub> value. (An empirical type I error greater than 0.05 &#x2b; 3 &#x22c5; 0.002 &#x3d; 0.056 can be considered inflated as the standard error of the empirical type I error rate can be estimated as <inline-formula id="inf15">
<mml:math id="m24">
<mml:msqrt>
<mml:mrow>
<mml:mn>0.05</mml:mn>
<mml:mo>&#x22c5;</mml:mo>
<mml:mn>0.95</mml:mn>
<mml:mo>/</mml:mo>
<mml:mn>10000</mml:mn>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0.002</mml:mn>
</mml:math>
</inline-formula>.) In <xref ref-type="fig" rid="F3">Figure 3A</xref>, the trend of type I error inflation is clear as <italic>&#x3b4;</italic>
<sub>causal</sub> increases.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Empirical type I error rate of <italic>T</italic> <sub>LMM</sub> based on the linear mixed-effect model (<xref ref-type="disp-formula" rid="e4">Eq. 4</xref>) against <bold>
<italic>&#x3b4;</italic>
</bold>
<sub>
<bold>causal</bold>
</sub>. <bold>(A)</bold> When <italic>G</italic>
<sub>tested</sub> of tested SNPs is in HWD with <italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0.06. <bold>(B)</bold> When <italic>G</italic>
<sub>tested</sub> of tested SNPs is in HWE with <italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0. The true heritability of the phenotype is <italic>h</italic>
<sup>2</sup> &#x3d; 0.5, the minor allele frequencies <italic>p</italic>
<sub>causal</sub> &#x3d; <italic>p</italic>
<sub>tested</sub> &#x3d; 0.2, and 10,000 independent replicates of phenotypes and genotypes for 65 sibling pairs were simulated for each <italic>&#x3b4;</italic>
<sub>causal</sub> value. The blue circles are for <italic>T</italic> <sub>LMM</sub> using the true heritability <italic>h</italic>
<sup>2</sup> &#x3d; 0.5, and the black squares are for <italic>T</italic> <sub>LMM</sub> while estimating <italic>h</italic>
<sup>2</sup> (results of <inline-formula id="inf16">
<mml:math id="m25">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> are shown in <xref ref-type="fig" rid="F1">Figure 1</xref>).</p>
</caption>
<graphic xlink:href="fgene-13-856872-g003.tif"/>
</fig>
<p>In <xref ref-type="fig" rid="F3">Figure 3A</xref>, we set <italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0.06, but we note that the main cause of the type I error issue is <italic>&#x3b4;</italic>
<sub>causal</sub> &#x2260; 0 when using the LMM of <xref ref-type="disp-formula" rid="e4">Eq. 4</xref> with <italic>h</italic>
<sup>2</sup> &#x3d; 0.5 plugged in. Indeed, <xref ref-type="fig" rid="F3">Figure 3B</xref> shows that even if <italic>G</italic>
<sub>tested</sub> is in HWE (i.e., <italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0), the problem remains, albeit less severe, as long as <italic>&#x3b4;</italic>
<sub>causal</sub> &#x2260; 0.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>We used a sib-pair design to demonstrate that the linear mixed-effect model can be problematic in the presence of Hardy&#x2013;Weinberg disequilibrium at the causal SNP(s). To demonstrate that the LMM-based heritability estimate can be biased, as a proof-of-principle, our simulation study assumed that the phenotype <italic>Y</italic> has only one causal SNP, which is unrealistic for complex traits. However, the analytical insight shown in <xref ref-type="disp-formula" rid="e5">Eq. 5</xref> (i.e., bias expected to be <inline-formula id="inf17">
<mml:math id="m26">
<mml:msup>
<mml:mrow>
<mml:mi>h</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x22c5;</mml:mo>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>) suggests that the issue discussed here remains relevant in the case of multiple causal SNPs as <inline-formula id="inf18">
<mml:math id="m27">
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is unlikely to be zero, even while allowing the signs of <italic>&#x3b4;</italic>
<sub>
<italic>k</italic>
</sub> to differ.</p>
<p>Assuming the true heritability <italic>h</italic>
<sup>2</sup> is known, we also demonstrated the potential type I error issue of the LMM in the presence of HWD using data that consist of related individuals only. In practice, this issue diminishes if the sample includes a large number of independent individuals or the magnitude of HWD at the causal SNP is small. Additionally, in practice, <italic>h</italic>
<sup>2</sup> is treated as unknown, in which case, the type I error rate of the LMM is well controlled; indeed, no increased false positives of the LMM due to HWD have been reported in the literature to the best of our knowledge. However, the estimate of <italic>h</italic>
<sup>2</sup> can be upward biased and upwardly so if <italic>&#x3b4;</italic>
<sub>causal</sub> &#x3e; 0, as demonstrated in the simulation study in <xref ref-type="sec" rid="s3-2">Section 3.2</xref> and seen in the cystic fibrosis application study in <xref ref-type="sec" rid="s3-1">Section 3.1</xref>. This new observation offers a possible complementary explanation of the &#x201c;missing heritability&#x201d; discussed extensively in <xref ref-type="bibr" rid="B9">Maher (2008)</xref>.</p>
<p>In practice, SNPs out of HWE are typically not analyzed due to concerns for low genotyping quality (<xref ref-type="bibr" rid="B29">Wellcome Trust Case Control Consortium, 2007</xref>; <xref ref-type="bibr" rid="B1">Bycroft et al., 2018</xref>; <xref ref-type="bibr" rid="B11">Marees et al., 2018</xref>). However, the observation made here remains relevant as the heritability estimates in LMM-based models are biased when the causal SNPs are in HWD (which is unknown in practice) but not the tested SNPs. This is also supported by <xref ref-type="fig" rid="F3">Figure 3B</xref>. When there was HWD at the causal SNP (e.g., <italic>&#x3b4;</italic>
<sub>causal</sub> &#x3d; 0.10 on the X-axis), there was a type I error issue even if there was no HWD at the tested SNP (i.e., <italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0). Conversely, <xref ref-type="fig" rid="F3">Figure 3A</xref> shows that if there was no HWD at the causal SNP (i.e., <italic>&#x3b4;</italic>
<sub>causal</sub> &#x3d; 0 on the X-axis), then the test is accurate even if there was HWD at the tested SNP (<italic>&#x3b4;</italic>
<sub>tested</sub> &#x3d; 0.06).</p>
<p>Additionally, the HWE-based screening practice itself can be called into question because a truly associated SNP is often in HWD (<xref ref-type="bibr" rid="B30">Wittke-Thompson et al., 2005</xref>; <xref ref-type="bibr" rid="B15">Ryckman and Williams, 2008</xref>; <xref ref-type="bibr" rid="B21">Turner et al., 2011</xref>). The potential of leveraging the HWD expected at a causal SNP to increase the power of association testing has been explored by several groups (<xref ref-type="bibr" rid="B18">Song and Elston, 2006</xref>; <xref ref-type="bibr" rid="B25">Wang and Shete, 2008</xref>; <xref ref-type="bibr" rid="B35">Zhang and Sun, 2020</xref>).</p>
<p>We have not examined the implication of HWD combined with linkage disequilibrium (LD) (<xref ref-type="bibr" rid="B28">Weir, 2008</xref>) on the LMM, which is an important future research question. Additionally, recent work has shown that dominant genetic effect could complicate the LD measure and interpretation (<xref ref-type="bibr" rid="B13">Palmer et al., 2021</xref>), which in turn could affect our examination of the effect of HWD on the LMM.</p>
<p>Although the linear mixed-effect model is a popular and powerful method for GWAS, conceptually, the use of kinship coefficient matrix (i.e., &#x3a3;<sub>&#x3a6;</sub>), derived from <italic>G</italic>, as part of the variance&#x2013;covariance matrix (i.e., &#x3a3;<sub>
<italic>y</italic>
</sub>) of the LMM can be problematic because the response variable <italic>Y</italic> is the phenotype of interest. An alternative approach is to reverse the roles of <italic>Y</italic> and <italic>G</italic> in the regression model. Indeed, <xref ref-type="bibr" rid="B12">O&#x2019;Reilly et al. (2012)</xref> proposed MultiPhen, a method that treats the genotype <italic>G</italic> of an SNP as the response variable and phenotype values <italic>Y</italic> of multiple traits as predictors, and uses an ordinal logistic regression applicable to independent samples. More recently, <xref ref-type="bibr" rid="B33">Zhang (2021)</xref> (Chapter 2) proposed a generalized reverse (or retrospective) regression model that can be applied to dependent samples, which takes the form of<disp-formula id="e6">
<mml:math id="m28">
<mml:mi>g</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b2;</mml:mi>
<mml:mtext>y</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3b3;</mml:mi>
<mml:mtext>z</mml:mtext>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>,</mml:mo>
<mml:mspace width="8.5359pt"/>
<mml:mi>&#x3f5;</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mtext>0</mml:mtext>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mspace width="8.5359pt"/>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>g</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>&#x3b4;</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where &#x3a3;<sub>&#x3a6;</sub> is the kinship coefficient matrix as defined earlier and &#x3a3;<sub>
<italic>&#x3b4;</italic>
</sub> is a function of <italic>&#x3b4;</italic> that explicitly models the amount of HWD; the use of a linear model for the discrete genotype data <italic>G</italic> is motivated by the work of <xref ref-type="bibr" rid="B2">Chen (1983)</xref>.</p>
<p>Interestingly, if the variance and covariance matrices in <xref ref-type="disp-formula" rid="e4">Eqs 4</xref>, <xref ref-type="disp-formula" rid="e6">6</xref> of the LMM were the same, the resulting score test statistics are also the same. However, conceptually, the model <xref ref-type="disp-formula" rid="e6">Eq. 6</xref> correctly uses the kinship coefficient matrix to model the response variable <italic>G</italic>, in contrast to the LMM model of <xref ref-type="disp-formula" rid="e4">Eq. 4</xref>. Specifically, at a tested SNP, as the reverse regression is conditional on <italic>Y</italic>, the variance&#x2013;covariance matrix only concerns <italic>G</italic>
<sub>tested</sub>, that is, &#x3a3;<sub>
<italic>g</italic>
</sub>. The modeling and estimation of &#x3a3;<sub>
<italic>g</italic>
</sub> can account for potential HWD through &#x3a3;<sub>
<italic>&#x3b4;</italic>
</sub>, in addition to the genetic correlation captured by the kinship coefficient matrix of &#x3a3;<sub>&#x3a6;</sub>, resulting in a more robust association test for related individuals. Indeed, when the method was applied to the same simulated sib-pair data in <xref ref-type="sec" rid="s3-3">Section 3.3</xref>, it had correct type I error control [results shown in Figure 2.2 of Chapter 2 of <xref ref-type="bibr" rid="B33">Zhang (2021)</xref>]. However, how to model gene&#x2013;environment interaction through the reverse regression framework remains an open question.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The data analyzed in this study are subject to the following licenses/restrictions: The CF application data are available by application to the Cystic Fibrosis Canada National data registry for researchers who meet the criteria for access to confidential clinical data for the purpose of CF research. Requests to access these datasets should be directed to <ext-link ext-link-type="uri" xlink:href="http://cfregistry@cysticfibrosis.ca">cfregistry@cysticfibrosis.ca</ext-link>.</p>
</sec>
<sec id="s6">
<title>Ethics Statement</title>
<p>The studies involving human participants were reviewed and approved by the Canadian Gene Modifier Study (CGMS), the Research Ethics Board of the Hospital for Sick Children (&#x23; 0020020214 from 2012&#x2013;2019 and &#x23;1000065760 from 2019-present), and all participating subsites. Written informed consent to participate in this study was provided by the participants&#x2019; legal guardian/next of kin.</p>
</sec>
<sec id="s7">
<title>Author Contributions</title>
<p>LZ and LS proposed the method and wrote the manuscript. LZ performed the analysis. LS obtained the application data and funding.</p>
</sec>
<sec id="s8">
<title>Funding</title>
<p>This research was funded by the Natural Sciences and Engineering Research Council of Canada (NSERC; RGPIN-04934 and RGPAS-522594).</p>
</sec>
<sec sec-type="COI-statement" id="s9">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors, and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>We thank Dr. Lisa J. Strug and acknowledge her laboratory for providing the cystic fibrosis genotype data. LZ was a trainee and funding recipient of the CANSSI Ontario STAGE (Strategic Training for Advanced Genetic Epidemiology) program at the University of Toronto.</p>
</ack>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bycroft</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Freeman</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Petkova</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Band</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Elliott</surname>
<given-names>L. T.</given-names>
</name>
<name>
<surname>Sharp</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>The uk Biobank Resource with Deep Phenotyping and Genomic Data</article-title>. <source>Nature</source> <volume>562</volume>, <fpage>203</fpage>&#x2013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0579-z</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>C.-F.</given-names>
</name>
</person-group> (<year>1983</year>). <article-title>Score Tests for Regression Models</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>78</volume>, <fpage>158</fpage>&#x2013;<lpage>161</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1983.10477945</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chen</surname>
<given-names>G.-B.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Estimating Heritability of Complex Traits from Genome-wide Association Studies Using IBS-Based Haseman&#xe2;&#x20ac;"Elston Regression</article-title>. <source>Front. Genet.</source> <volume>5</volume>, <fpage>107</fpage>. <pub-id pub-id-type="doi">10.3389/fgene.2014.00107</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Derkach</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lawless</surname>
<given-names>J. F.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Score Tests for Association under Response-dependent Sampling Designs for Expensive Covariates</article-title>. <source>Biometrika</source> <volume>102</volume>, <fpage>988</fpage>&#x2013;<lpage>994</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/asv038</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dimitromanolakis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Paterson</surname>
<given-names>A. D.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Fast and Accurate Shared Segment Detection and Relatedness Estimation in Un-phased Genetic Data via Truffle</article-title>. <source>Am. J. Hum. Genet.</source> <volume>105</volume>, <fpage>78</fpage>&#x2013;<lpage>88</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2019.05.007</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Eu-Ahsunthornwattana</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Miller</surname>
<given-names>E. N.</given-names>
</name>
<name>
<surname>Fakiola</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jeronimo</surname>
<given-names>S. M. B.</given-names>
</name>
<name>
<surname>Blackwell</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Cordell</surname>
<given-names>H. J.</given-names>
</name>
<etal/>
</person-group> (<year>2014</year>). <article-title>Comparison of Methods to Account for Relatedness in Genome-wide Association Studies with Family-Based Data</article-title>. <source>Plos Genet.</source> <volume>10</volume>, <fpage>e1004445</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1004445</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Falconer</surname>
<given-names>D. S.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>A Note on Fisher&#x27;s &#x27;average Effect&#x27; and &#x27;average Excess&#x27;</article-title>. <source>Genet. Res.</source> <volume>46</volume>, <fpage>337</fpage>&#x2013;<lpage>347</lpage>. <pub-id pub-id-type="doi">10.1017/s0016672300022825</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hardy</surname>
<given-names>G. H.</given-names>
</name>
</person-group> (<year>1908</year>). <article-title>Mendelian Proportions in a Mixed Population</article-title>. <source>Science</source> <volume>28</volume>, <fpage>49</fpage>&#x2013;<lpage>50</lpage>. <pub-id pub-id-type="doi">10.1126/science.28.706.49</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Maher</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Personal Genomes: The Case of the Missing Heritability</article-title>. <source>Nature</source> <volume>456</volume>, <fpage>18</fpage>&#x2013;<lpage>21</lpage>. <pub-id pub-id-type="doi">10.1038/456018a</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manolio</surname>
<given-names>T. A.</given-names>
</name>
<name>
<surname>Collins</surname>
<given-names>F. S.</given-names>
</name>
<name>
<surname>Cox</surname>
<given-names>N. J.</given-names>
</name>
<name>
<surname>Goldstein</surname>
<given-names>D. B.</given-names>
</name>
<name>
<surname>Hindorff</surname>
<given-names>L. A.</given-names>
</name>
<name>
<surname>Hunter</surname>
<given-names>D. J.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Finding the Missing Heritability of Complex Diseases</article-title>. <source>Nature</source> <volume>461</volume>, <fpage>747</fpage>&#x2013;<lpage>753</lpage>. <pub-id pub-id-type="doi">10.1038/nature08494</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Marees</surname>
<given-names>A. T.</given-names>
</name>
<name>
<surname>de Kluiver</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Stringer</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Vorspan</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Curis</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Marie-Claire</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>A Tutorial on Conducting Genome-wide Association Studies: Quality Control and Statistical Analysis</article-title>. <source>Int. J. Methods Psychiatr. Res.</source> <volume>27</volume>, <fpage>e1608</fpage>. <pub-id pub-id-type="doi">10.1002/mpr.1608</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>O&#x27;Reilly</surname>
<given-names>P. F.</given-names>
</name>
<name>
<surname>Hoggart</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Pomyen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Calboli</surname>
<given-names>F. C.</given-names>
</name>
<name>
<surname>Elliott</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Jarvelin</surname>
<given-names>M. R.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Multiphen: Joint Model of Multiple Phenotypes Can Increase Discovery in Gwas</article-title>. <source>PLoS One</source> <volume>7</volume>, <fpage>e34861</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0034861</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Palmer</surname>
<given-names>D. S.</given-names>
</name>
<name>
<surname>Zhou</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Abbott</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Baya</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Churchhouse</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Seed</surname>
<given-names>C.</given-names>
</name>
<etal/>
</person-group> (<year>2021</year>). <article-title>Analysis of Genetic Dominance in the uk Biobank</article-title>. <comment>bioRxiv</comment> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Powell</surname>
<given-names>J. E.</given-names>
</name>
<name>
<surname>Visscher</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Goddard</surname>
<given-names>M. E.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Reconciling the Analysis of Ibd and Ibs in Complex Trait Studies</article-title>. <source>Nat. Rev. Genet.</source> <volume>11</volume>, <fpage>800</fpage>&#x2013;<lpage>805</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2865</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ryckman</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>S. M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Calculation and Use of the hardy-weinberg Model in Association Studies</article-title>. <source>Curr. Protoc. Hum. Genet.</source> <volume>Chapter 1</volume>, <fpage>Unit</fpage>&#x2013;<lpage>18</lpage>. <pub-id pub-id-type="doi">10.1002/0471142905.hg0118s57</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sasieni</surname>
<given-names>P. D.</given-names>
</name>
</person-group> (<year>1997</year>). <article-title>From Genotypes to Genes: Doubling the Sample Size</article-title>. <source>Biometrics</source> <volume>53</volume>, <fpage>1253</fpage>&#x2013;<lpage>1261</lpage>. <pub-id pub-id-type="doi">10.2307/2533494</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Schaid</surname>
<given-names>D. J.</given-names>
</name>
<name>
<surname>Jacobsen</surname>
<given-names>S. J.</given-names>
</name>
</person-group> (<year>1999</year>). <article-title>Blased Tests of Association: Comparisons of Allele Frequencies when Departing from Hardy-Weinberg Proportions</article-title>. <source>Am. J. Epidemiol.</source> <volume>149</volume>, <fpage>706</fpage>&#x2013;<lpage>711</lpage>. <pub-id pub-id-type="doi">10.1093/oxfordjournals.aje.a009878</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Song</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Elston</surname>
<given-names>R. C.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>A Powerful Method of Combining Measures of Association and Hardy-Weinberg Disequilibrium for fine-mapping in Case-Control Studies</article-title>. <source>Statist. Med.</source> <volume>25</volume>, <fpage>105</fpage>&#x2013;<lpage>126</lpage>. <pub-id pub-id-type="doi">10.1002/sim.2350</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Dimitromanolakis</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>W.-M.</given-names>
</name>
</person-group> (<year>2017</year>). &#x201c;<article-title>Identifying Cryptic Relationships</article-title>,&#x201d; in <source>Statistical Human Genetics</source> (<publisher-name>Springer</publisher-name>), <fpage>45</fpage>&#x2013;<lpage>60</lpage>. <pub-id pub-id-type="doi">10.1007/978-1-4939-7274-6_4</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Rommens</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Corvol</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Chiang</surname>
<given-names>T. A.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Multiple Apical Plasma Membrane Constituents Are Associated with Susceptibility to Meconium Ileus in Individuals with Cystic Fibrosis</article-title>. <source>Nat. Genet.</source> <volume>44</volume>, <fpage>562</fpage>&#x2013;<lpage>569</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2221</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Turner</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Armstrong</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Bradford</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Carlson</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Crawford</surname>
<given-names>D. C.</given-names>
</name>
<name>
<surname>Crenshaw</surname>
<given-names>A. T.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Quality Control Procedures for Genome-wide Association Studies</article-title>. <source>Curr. Protoc. Hum. Genet.</source> <volume>Chapter 1</volume>, <fpage>Unit1</fpage>&#x2013;<lpage>19</lpage>. <pub-id pub-id-type="doi">10.1002/0471142905.hg0119s68</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Vanscoy</surname>
<given-names>L. L.</given-names>
</name>
<name>
<surname>Blackman</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Collaco</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Bowers</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Lai</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Naughton</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>Heritability of Lung Disease Severity in Cystic Fibrosis</article-title>. <source>Am. J. Respir. Crit. Care Med.</source> <volume>175</volume>, <fpage>1036</fpage>&#x2013;<lpage>1043</lpage>. <pub-id pub-id-type="doi">10.1164/rccm.200608-1164oc</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Visscher</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Hill</surname>
<given-names>W. G.</given-names>
</name>
<name>
<surname>Wray</surname>
<given-names>N. R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Heritability in the Genomics Era - Concepts and Misconceptions</article-title>. <source>Nat. Rev. Genet.</source> <volume>9</volume>, <fpage>255</fpage>&#x2013;<lpage>266</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2322</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Visscher</surname>
<given-names>P. M.</given-names>
</name>
<name>
<surname>Medland</surname>
<given-names>S. E.</given-names>
</name>
<name>
<surname>Ferreira</surname>
<given-names>M. A. R.</given-names>
</name>
<name>
<surname>Morley</surname>
<given-names>K. I.</given-names>
</name>
<name>
<surname>Zhu</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Cornes</surname>
<given-names>B. K.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>Assumption-free Estimation of Heritability from Genome-wide Identity-By-Descent Sharing between Full Siblings</article-title>. <source>Plos Genet.</source> <volume>2</volume>, <fpage>e41</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.0020041</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Shete</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>A Test for Genetic Association that Incorporates Information about Deviation from hardy-weinberg Proportions in Cases</article-title>. <source>Am. J. Hum. Genet.</source> <volume>83</volume>, <fpage>53</fpage>&#x2013;<lpage>63</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2008.06.010</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Weinberg</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>1908</year>). <source>On the Demonstration of Heredity in Man</source>. in <source>(1963) Papers on Human Genetics</source>. </citation>
</ref>
<ref id="B27">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Weir</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>1996</year>). <source>Genetic Data Analysis II: Methods for Discrete Population Genetic Data</source>. <publisher-loc>Sunderland, Massachusetts</publisher-loc>: <publisher-name>Sinauer Series (Sinauer)</publisher-name>. </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Weir</surname>
<given-names>B. S.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Linkage Disequilibrium and Association Mapping</article-title>. <source>Annu. Rev. Genom. Hum. Genet.</source> <volume>9</volume>, <fpage>129</fpage>&#x2013;<lpage>142</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.genom.9.081307.164347</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<collab>Wellcome Trust Case Control Consortium</collab> (<year>2007</year>). <article-title>Genome-wide Association Study of 14,000 Cases of Seven Common Diseases and 3,000 Shared Controls</article-title>. <source>Nature</source> <volume>447</volume>, <fpage>661</fpage>&#x2013;<lpage>678</lpage>. <pub-id pub-id-type="doi">10.1038/nature05911</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wittke-Thompson</surname>
<given-names>J. K.</given-names>
</name>
<name>
<surname>Pluzhnikov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Cox</surname>
<given-names>N. J.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Rational Inferences about Departures from hardy-weinberg Equilibrium</article-title>. <source>Am. J. Hum. Genet.</source> <volume>76</volume>, <fpage>967</fpage>&#x2013;<lpage>986</lpage>. <pub-id pub-id-type="doi">10.1086/430507</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wright</surname>
<given-names>F. A.</given-names>
</name>
<name>
<surname>Strug</surname>
<given-names>L. J.</given-names>
</name>
<name>
<surname>Doshi</surname>
<given-names>V. K.</given-names>
</name>
<name>
<surname>Commander</surname>
<given-names>C. W.</given-names>
</name>
<name>
<surname>Blackman</surname>
<given-names>S. M.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Genome-wide Association and Linkage Identify Modifier Loci of Lung Disease Severity in Cystic Fibrosis at 11p13 and 20q13.2</article-title>. <source>Nat. Genet.</source> <volume>43</volume>, <fpage>539</fpage>&#x2013;<lpage>546</lpage>. <pub-id pub-id-type="doi">10.1038/ng.838</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Goddard</surname>
<given-names>M. E.</given-names>
</name>
<name>
<surname>Visscher</surname>
<given-names>P. M.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Gcta: a Tool for Genome-wide Complex Trait Analysis</article-title>. <source>Am. J. Hum. Genet.</source> <volume>88</volume>, <fpage>76</fpage>&#x2013;<lpage>82</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2010.11.011</pub-id> </citation>
</ref>
<ref id="B33">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <source>A General Study of Genetic Association Tests and the Test of Hardy&#x2013;Weinberg Equilibrium</source>. <comment>Ph.D. thesis</comment>. <publisher-name>University of Toronto</publisher-name>. </citation>
</ref>
<ref id="B34">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2021</year>). <source>A Generalized Robust Allele-Based Genetic Association Test</source>. <publisher-loc>Oxford, United Kingdom</publisher-loc>: <publisher-name>Biometrics</publisher-name>. </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Sun</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Leveraging hardy-weinberg Disequilibrium for Association Testing in Case-Control Studies</article-title>. <comment>bioRxiv</comment>. </citation>
</ref>
</ref-list>
</back>
</article>