<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="research-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">867724</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2022.867724</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>A Comparison of Methods for Gene-Based Testing That Account for Linkage Disequilibrium</article-title>
<alt-title alt-title-type="left-running-head">Cinar and Viechtbauer</alt-title>
<alt-title alt-title-type="right-running-head">Methods for Gene-Based Testing</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Cinar</surname>
<given-names>Ozan</given-names>
</name>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1011291/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Viechtbauer</surname>
<given-names>Wolfgang</given-names>
</name>
<uri xlink:href="https://loop.frontiersin.org/people/564047/overview"/>
</contrib>
</contrib-group>
<aff>
<institution>Department of Psychiatry and Neuropsychology</institution>, <institution>Maastricht University</institution>, <addr-line>Maastricht</addr-line>, <country>Netherlands</country>
</aff>
<author-notes>
<corresp id="c001">&#x2a;Correspondence: Ozan Cinar, <email>ozancinar86@gmail.com</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/744713/overview">Lide Han</ext-link>, Vanderbilt University Medical Center, United States</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/496329/overview">Jianbo He</ext-link>, Nanjing Agricultural University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1035276/overview">Judong Shen</ext-link>, Merck &#x26; Co., Inc., United States</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>05</day>
<month>05</month>
<year>2022</year>
</pub-date>
<pub-date pub-type="collection">
<year>2022</year>
</pub-date>
<volume>13</volume>
<elocation-id>867724</elocation-id>
<history>
<date date-type="received">
<day>01</day>
<month>02</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>04</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2022 Cinar and Viechtbauer.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Cinar and Viechtbauer</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p>
</license>
</permissions>
<abstract>
<p>Controlling the type I error rate while retaining sufficient power is a major concern in genome-wide association studies, which nowadays often examine more than a million single-nucleotide polymorphisms (SNPs) simultaneously. Methods such as the Bonferroni correction can lead to a considerable decrease in power due to the large number of tests conducted. Shifting the focus to higher functional structures (e.g., genes) can reduce the loss of power. This can be accomplished via the combination of <italic>p</italic>-values of SNPs that belong to the same structural unit to test their joint null hypothesis. However, standard methods for this purpose (e.g., Fisher&#x2019;s method) do not account for the dependence among the tests due to linkage disequilibrium (LD). In this paper, we review various adjustments to methods for combining <italic>p</italic>-values that take LD information explicitly into consideration and evaluate their performance in a simulation study based on data from the HapMap project. The results illustrate the importance of incorporating LD information into the methods for controlling the type I error rate at the desired level. Furthermore, some methods are more successful in controlling the type I error rate than others. Among them, Brown&#x2019;s method was the most robust technique with respect to the characteristics of the genes and outperformed the Bonferroni method in terms of power in many scenarios. Examining the genetic factors of a phenotype of interest at the gene-rather than SNP-level can provide researchers benefits in terms of the power of the study. While doing so, one should be careful to account for LD in SNPs belonging to the same gene, for which Brown&#x2019;s method seems the most robust technique.</p>
</abstract>
<kwd-group>
<kwd>genome-wide association studies</kwd>
<kwd>gene-based testing</kwd>
<kwd>combining p-values</kwd>
<kwd>correlated tests</kwd>
<kwd>linkage disequilibrium</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Genome-wide association (GWA) studies are commonly used to investigate the contribution of genetic variants to the risk of developing certain diseases (<xref ref-type="bibr" rid="B45">Manolio, 2010</xref>). In a typical GWA study, large quantities of single-nucleotide polymorphisms (SNPs) are genotyped to examine their association with some phenotype of interest (e.g., the presence or absence of a disease) or their interaction with some environmental factor (<xref ref-type="bibr" rid="B3">Baranzini et al., 2009</xref>; <xref ref-type="bibr" rid="B29">Jiao et al., 2015</xref>). However, the availability of genotype information for such a large number of SNPs will either lead to a high rate of type I errors or requires stringent corrections for multiple testing, which in turn inflates the number of type II errors (<xref ref-type="bibr" rid="B30">Johnson et al., 2010</xref>).</p>
<p>In particular, the probability of falsely rejecting an individual null hypothesis (e.g., that a SNP is unrelated to the outcome) is set a priori to a specific value by the researcher. This <italic>pointwise error rate</italic> (or error rate per hypothesis) is conventionally set to <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub> &#x3d; 0.05. However, the <italic>familywise error rate</italic>, <inline-formula id="inf1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> (i.e., the probability of falsely rejecting at least one of <italic>k</italic> true null hypotheses) quickly increases when testing a large number of independent hypotheses. A common method to control the familywise error rate is the Bonferroni correction (<xref ref-type="bibr" rid="B7">Bland and Altman, 1995</xref>) that sets the pointwise error rate to <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub>/<italic>k</italic>, which in turn keeps <italic>&#x3b1;</italic>
<sub>
<italic>s</italic>
</sub> below the desired type I error rate (<xref ref-type="bibr" rid="B55">Shaffer, 1995</xref>). Considering that nowadays around a million SNPs are genotyped in a typical GWA study (<xref ref-type="bibr" rid="B45">Manolio, 2010</xref>), the commonly used significance threshold of 5 &#xd7; 10<sup>&#x2013;8</sup> in such studies is loosely based on the Bonferroni correction (<xref ref-type="bibr" rid="B30">Johnson et al., 2010</xref>; <xref ref-type="bibr" rid="B26">Huang et al., 2012</xref>).</p>
<p>As a consequence of the decreased significance threshold, rejection of a null hypothesis becomes more difficult, whether it be a true null hypothesis or not. Therefore, the Bonferroni correction also increases the type II error rate (i.e., the probability of failing to reject a false null hypothesis), which in turn decreases power. Although other multiple testing correction methods have been developed that lead to less severe reductions in power (<xref ref-type="bibr" rid="B24">Holm, 1979</xref>; <xref ref-type="bibr" rid="B57">Simes, 1986</xref>; <xref ref-type="bibr" rid="B23">Hochberg, 1988</xref>; <xref ref-type="bibr" rid="B25">Hommel, 1988</xref>; <xref ref-type="bibr" rid="B6">Benjamini and Hochberg, 1995</xref>; <xref ref-type="bibr" rid="B15">Conneely and Boehnke, 2007</xref>), the reduction can still be severe due to the large number of SNPs considered in a typical GWA study (<xref ref-type="bibr" rid="B49">Narum, 2006</xref>).</p>
<p>A promising approach for mitigating this severe loss of power is to shift the focus of the analyses to higher functional structures such as genes (known as gene-based testing) or sets of genes that belong to common pathways (<xref ref-type="bibr" rid="B34">Lehne et al., 2011</xref>). As a result, the number of hypotheses tested declines dramatically (e.g., to 25,000&#x2013;30,000 when testing at the gene level) and hence power is not as severely impacted when a correction for multiple testing is then applied. Furthermore, by aggregating signals from multiple SNPs, gene-based testing can be more appropriate for understanding the genetic structure of complex diseases (<xref ref-type="bibr" rid="B42">Liu et al., 2010</xref>; <xref ref-type="bibr" rid="B11">Chung et al., 2019</xref>).</p>
<p>Although the joint contribution of the SNPs in a gene can be examined with multi-locus tests, such as Hotelling&#x2019;s <italic>T</italic>
<sup>2</sup> (<xref ref-type="bibr" rid="B9">Chapman and Whittaker, 2008</xref>; <xref ref-type="bibr" rid="B48">Moskvina et al., 2012</xref>), such approaches require access to the raw genomic data which may not be available. In the absence of the raw data, we can test the joint null hypothesis of the SNPs that belong to a gene by combining their individual <italic>p</italic>-values into an overall <italic>p</italic>-value. A wide variety of methods have been described in the literature for combining independent tests of hypotheses (<xref ref-type="bibr" rid="B52">Pearson, 1938</xref>; <xref ref-type="bibr" rid="B33">Lancaster, 1949</xref>; <xref ref-type="bibr" rid="B59">Stouffer et al., 1949</xref>; <xref ref-type="bibr" rid="B66">Wilkinson, 1951</xref>; <xref ref-type="bibr" rid="B39">Lipt&#xe1;k, 1958</xref>; <xref ref-type="bibr" rid="B5">Becker, 1994</xref>). Among these, Fisher&#x2019;s method (<xref ref-type="bibr" rid="B19">Fisher, 1932</xref>) may be the best-known one, which also has high relative efficiency asymptotically when compared to other methods (<xref ref-type="bibr" rid="B40">Littell and Folks, 1971</xref>; <xref ref-type="bibr" rid="B41">Littell and Folks, 1973</xref>). However, Fisher&#x2019;s method, like many other methods for combining tests of hypotheses, assumes that the <italic>p</italic>-values are independent of each other. This assumption is known to be violated in the present context as SNPs are often in linkage disequilibrium (LD), that is, the alleles at different loci exhibit non-random associations (<xref ref-type="bibr" rid="B58">Slatkin, 2008</xref>). As a consequence, Fisher&#x2019;s method does not provide nominal results, typically leading to an inflation in the type I error rate (<xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>).</p>
<p>Several attempts have been made to adjust methods for combining <italic>p</italic>-values such that they take dependence into consideration. <xref ref-type="bibr" rid="B8">Brown (1975)</xref> proposed an adjustment to Fisher&#x2019;s method for combining the results of dependent tests that has been used for gene-based testing (<xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>; <xref ref-type="bibr" rid="B70">Zhang et al., 2020</xref>), whereas several other authors have described the use of principal component analysis (PCA) on the LD correlation matrix to estimate the effective number of tests (<xref ref-type="bibr" rid="B10">Cheverud, 2001</xref>; <xref ref-type="bibr" rid="B51">Nyholt, 2004</xref>; <xref ref-type="bibr" rid="B35">Li and Ji, 2005</xref>; <xref ref-type="bibr" rid="B21">Gao et al., 2008</xref>; <xref ref-type="bibr" rid="B20">Galwey, 2009</xref>), which in turn can be combined with various multiple testing correction procedures. In addition, several authors have applied permutation tests or other permutation-type procedures to account for the dependence (<xref ref-type="bibr" rid="B38">Lin, 2005</xref>; <xref ref-type="bibr" rid="B42">Liu et al., 2010</xref>). However, while permutation tests are often considered a &#x2018;gold standard&#x2019; approach, such methods are computationally very demanding especially in GWA studies. Furthermore, proper permutation tests require access to the raw data which can be another limitation. A promising way to mimic the results of permutation tests (without needing the raw data and requiring a fraction of the time) is to generate pseudo replicates of the test statistics assuming they follow a multivariate normal distribution under the null hypothesis. These pseudo test statistics are then converted into SNP-level <italic>p</italic>-values which can be used to generate an empirical distribution of the combined <italic>p</italic>-value under the null hypothesis that takes the degree of LD into consideration (<xref ref-type="bibr" rid="B42">Liu et al., 2010</xref>; <xref ref-type="bibr" rid="B36">Li et al., 2011</xref>).</p>
<p>The statistical properties (i.e., type I error rate and power) of these methods have been examined in previous research (<xref ref-type="bibr" rid="B38">Lin, 2005</xref>; <xref ref-type="bibr" rid="B15">Conneely and Boehnke, 2007</xref>; <xref ref-type="bibr" rid="B9">Chapman and Whittaker, 2008</xref>; <xref ref-type="bibr" rid="B30">Johnson et al., 2010</xref>; <xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>; <xref ref-type="bibr" rid="B65">Wen and Lu, 2011</xref>; <xref ref-type="bibr" rid="B1">Alves and Yu, 2014</xref>). However, there are still several points that have not been considered so far. First, none of the studies have performed an extensive comparison among all methods on a genome-wide scale simultaneously. Furthermore, PCA-based approaches have only been combined with the Bonferroni correction (or with Tippett&#x2019;s method; see below), although they can also be used to modify other tests (e.g., Fisher&#x2019;s method) to account for dependence among the <italic>p</italic>-values. In addition, it is unknown whether the statistical properties of the correlations (e.g., their central tendency or spread) used to quantify the degree of LD might affect the performance of the methods. Moreover, some theoretical properties of the methods have not received sufficient attention. Most importantly, Brown&#x2019;s generalization of Fisher&#x2019;s method only applies to one-sided tests (<xref ref-type="bibr" rid="B8">Brown, 1975</xref>). This property is especially problematic in GWA studies, since tests of the association between the SNPs and the phenotype of interest are typically two-sided (<xref ref-type="bibr" rid="B32">Laird and Lange, 2010</xref>). An extension of Brown&#x2019;s method to two-sided tests has been described (<xref ref-type="bibr" rid="B69">Yang et al., 2016</xref>); however, its performance in the present context has yet to be investigated.</p>
<p>In this article, we review a variety of methods for combining <italic>p</italic>-values that can be used for gene-based testing and describe how LD can be directly incorporated into these methods. While doing so, an important goal is to provide a more complete description of how methods for combining <italic>p</italic>-values and adjustment techniques can be combined. For example, we will describe how an estimate of the effective number of tests can be used to adjust Fisher&#x2019;s method. Furthermore, we describe how all methods for combining <italic>p</italic>-values can be adjusted with an empirical distribution obtained using a pseudo-permutation approach. We also discuss the generalization of Brown&#x2019;s method to two-sided tests. Finally, we compare the type I error rate and power of the methods based on a genome-wide Monte Carlo simulation study using LD matrices derived from the International HapMap Project (<xref ref-type="bibr" rid="B61">The International HapMap Consortium, 2003</xref>).</p>
</sec>
<sec id="s2">
<title>2 Methods</title>
<p>For a collection of <italic>i</italic> &#x3d; 1, &#x2026; , <italic>k</italic> SNPs that belong to a gene (or pathway), let <italic>p</italic>
<sub>1</sub>, &#x2026; , <italic>p</italic>
<sub>
<italic>k</italic>
</sub> denote the <italic>p</italic>-values obtained when testing the association of each SNP with some phenotype of interest (or the interaction of each SNP with some other variable). We use <italic>H</italic>
<sub>0<italic>i</italic>
</sub> to denote the null hypothesis corresponding to the <italic>i</italic>th SNP. Since we are only interested in testing for association regardless of directionality, we assume that the <italic>p</italic>-values are derived from two-sided tests. Moreover, we assume that the tests have nominal properties, so that <italic>p</italic>
<sub>
<italic>i</italic>
</sub> &#x223c;Uniform (0, 1) when <italic>H</italic>
<sub>0<italic>i</italic>
</sub> is true. Depending on the type of test used for deriving the <italic>p</italic>-values, this assumption may only be true asymptotically (i.e., if the sample size underlying the tests is large). For the purposes of describing the methods, we still make this assumption, but return to this issue in the discussion section.</p>
<p>Instead of considering each of the <italic>p</italic>-values and null hypotheses individually, the goal is to combine the information from the individual tests into one that tests the gene as a whole. To be precise, the goal is to test the joint null hypothesis that none of the SNPs in the gene are associated with the phenotype (i.e., <italic>H</italic>
<sub>0<italic>i</italic>
</sub> is true for all tests) against the alternative that at least one SNP is associated. We will now describe a variety of methods for this purpose.</p>
<sec id="s2-1">
<title>2.1 The Bonferroni Method</title>
<p>The Bonferroni correction (<xref ref-type="bibr" rid="B7">Bland and Altman, 1995</xref>) is a method that was originally developed to control the familywise error rate when conducting multiple hypothesis tests. In order to apply the correction, the threshold for significance is adjusted by dividing the pointwise error rate, <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub>, by the number of simultaneous tests, <italic>k</italic>. Alternatively, we can adjust the individual <italic>p</italic>-values by multiplying them with <italic>k</italic>. Any test whose adjusted <italic>p</italic>-value is then equal to or less than <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub> is declared significant (<xref ref-type="bibr" rid="B57">Simes, 1986</xref>).</p>
<p>Although not typically described in this manner, the Bonferroni method can also be used as a method for combining <italic>p</italic>-values. In particular, if any one of the adjusted <italic>p</italic>-values is significant, then the joint null hypothesis is automatically rejected. In the context of GWA studies, this means that if at least one SNP is significantly associated with the phenotype of interest, then the gene that this SNP belongs to is considered significant. Accordingly, the combined <italic>p</italic>-value for a gene is given by<disp-formula id="e1">
<mml:math id="m2">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(1)</label>
</disp-formula>where min(1, &#x2026; ) simply ensures that the combined <italic>p</italic>-value cannot exceed 1.</p>
<p>Contrary to popular belief, the Bonferroni correction does not make any assumptions about the degree of dependence among the <italic>p</italic>-values (<xref ref-type="bibr" rid="B22">Goeman and Solari, 2014</xref>). In other words, regardless of the degree of dependence among the tests from which the <italic>p</italic>-values are derived, the method guarantees that the type I error rate is no larger than the desired nominal rate. This makes the methods particularly interesting for gene-based testing, where we know that the tests are likely to be dependent due to LD.</p>
</sec>
<sec id="s2-2">
<title>2.2 Methods Assuming Independence</title>
<p>In this subsection, we will describe methods that assume that the tests, and hence the <italic>p</italic>-values to be combined, are independent. Adjustments thereof will be considered later.</p>
<sec id="s2-3">
<title>2.2.1 Tippett&#x2019;s Method</title>
<p>Tippett&#x2019;s method (<xref ref-type="bibr" rid="B62">Tippett, 1931</xref>), also known as the Dunn-&#x160;id&#xe1;k correction for multiple testing (<xref ref-type="bibr" rid="B56">&#x160;id&#xe1;k, 1957</xref>; <xref ref-type="bibr" rid="B16">Dunn, 1958</xref>), follows from the fact that the familywise type I error rate for <italic>k</italic> independent tests, <inline-formula id="inf2">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>, will equal a desired nominal rate, <italic>&#x3b1;</italic>, if we set <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub> &#x3d; 1 &#x2212; (1 &#x2212; <italic>&#x3b1;</italic>)<sup>1/<italic>k</italic>
</sup>. Hence, the joint null distribution can be rejected if min(<italic>p</italic>
<sub>
<italic>i</italic>
</sub>) &#x2264; 1 &#x2212; (1 &#x2212; <italic>&#x3b1;</italic>)<sup>1/<italic>k</italic>
</sup>. Analogously, we can use<disp-formula id="e2">
<mml:math id="m4">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(2)</label>
</disp-formula>
</p>
<p>as the <italic>p</italic>-value for the gene. As opposed to the Bonferroni method, which is slightly conservative even when all tests are independent, the method provides exact control of the type I error rate, but only under independence.</p>
</sec>
<sec id="s2-3-1">
<title>2.2.2 Binomial Test</title>
<p>Under the joint null hypothesis, <italic>r</italic> &#x223c;Binomial (<italic>k</italic>, <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub>) where <italic>r</italic> denotes the number of tests that are significant at <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub>. Therefore, we can reject the joint null hypothesis if<disp-formula id="e3">
<mml:math id="m5">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mfenced open="(" close=")">
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mfenced>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(3)</label>
</disp-formula>
</p>
<p>is equal to or less than the desired type I error rate (<xref ref-type="bibr" rid="B66">Wilkinson, 1951</xref>). Intuitively, we can interpret this method as a test of &#x201c;excess significance&#x201d; of the SNPs within a gene. For example, the chances of finding <italic>r</italic> &#x3d; 10 or more significant SNPs at <italic>&#x3b1;</italic>
<sub>
<italic>p</italic>
</sub> &#x3d; 0.05 in a gene with 100 independent SNPs is approximately <italic>p</italic> &#x3d; 0.028, which would be significant at <italic>&#x3b1;</italic> &#x3d; 0.05.</p>
</sec>
<sec id="s2-3-2">
<title>2.2.3 Fisher&#x2019;s Method</title>
<p>Assuming that <italic>p</italic>
<sub>
<italic>i</italic>
</sub> &#x223c;Uniform (0, 1) under the null hypothesis of no association, it is easy to show that &#x2212;2&#x2009;ln(<italic>p</italic>
<sub>
<italic>i</italic>
</sub>) is chi-square distributed with 2 degrees of freedom. Hence, the combined test statistic<disp-formula id="e4">
<mml:math id="m6">
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(4)</label>
</disp-formula>
</p>
<p>follows a chi-squared distribution with 2<italic>k</italic> degrees of freedom under the joint null hypothesis (<xref ref-type="bibr" rid="B19">Fisher, 1932</xref>). The <italic>p</italic>-value for the gene can therefore be computed with <italic>p</italic> &#x3d; 1 &#x2212; <italic>F</italic> (<italic>X</italic>
<sup>2</sup>, 2<italic>k</italic>), where <italic>F</italic> (&#x22c5;, 2<italic>k</italic>) denotes the cumulative distribution function of a chi-square distribution with 2<italic>k</italic> degrees of freedom.</p>
</sec>
<sec id="s2-3-3">
<title>2.2.4 Stouffer&#x2019;s Method</title>
<p>Let &#x3a6;(&#x22c5;) denote the cumulative distribution function of the standard normal distribution and &#x3a6;<sup>&#x2212;1</sup>(&#x22c5;) its inverse. Since &#x3a6;<sup>&#x2212;1</sup> (1 &#x2212; <italic>p</italic>
<sub>
<italic>i</italic>
</sub>) follows a standard normal distribution under <italic>H</italic>
<sub>0<italic>i</italic>
</sub>, <inline-formula id="inf3">
<mml:math id="m7">
<mml:mi>z</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#x223c;</mml:mo>
<mml:mtext>Normal</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0,1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> under the joint null hypothesis (<xref ref-type="bibr" rid="B59">Stouffer et al., 1949</xref>). The <italic>p</italic>-value for a gene is then computed with <italic>p</italic> &#x3d; 1 &#x2212; &#x3a6;(<italic>z</italic>).</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.3 Incorporating Linkage Disequilibrium</title>
<p>Except for the Bonferroni method, all methods described in the previous section assume that the tests are independent. Therefore, under this assumption (and the assumptions stated at the beginning of this section), these methods are guaranteed to have a type I error rate equal to the desired <italic>&#x3b1;</italic> level (for the binomial test, the type I error rate is <inline-formula id="inf4">
<mml:math id="m8">
<mml:mo>&#x2264;</mml:mo>
<mml:mi>&#x3b1;</mml:mi>
</mml:math>
</inline-formula> due to the discrete nature of the binomial distribution). On the other hand, when the independence assumption is violated, the true type I error rate may deviate from <italic>&#x3b1;</italic> in either direction, but usually leading to inflation (i.e., the joint null is rejected too often). In comparison, the Bonferroni method makes no assumptions about the degree of dependence among the tests and is guaranteed to have a rejection rate that is no larger than <italic>&#x3b1;</italic>, but can be quite conservative under dependence. We will therefore now consider adjustments to the methods that account for dependence among the tests and that can bring their type I error rate closer to <italic>&#x3b1;</italic>.</p>
<sec id="s2-4-1">
<title>2.3.1 Effective Number of Tests</title>
<p>One potential approach to adjust the previous methods is to quantify the degree of dependence between the tests, estimate the effective number of independent tests based on this information, and incorporate this estimate into the methods described above.</p>
<p>The degree of dependence between the tests is closely related to the strength of the association between the SNPs. The latter can be quantified with various statistics (e.g., <italic>D</italic>, <italic>D</italic>&#x2032;, <italic>r</italic>, or <italic>r</italic>
<sup>2</sup>) expressing the degree of LD between pairs of SNPs (<xref ref-type="bibr" rid="B32">Laird and Lange, 2010</xref>). We can use one of these measures to construct a <italic>k</italic> &#xd7; <italic>k</italic> association matrix for all SNPs, sometimes called an &#x201c;LD map&#x201d;. The effective number of tests can then be estimated based on this association matrix. A variety of approaches have been described in the literature for this purpose (<xref ref-type="bibr" rid="B10">Cheverud, 2001</xref>; <xref ref-type="bibr" rid="B51">Nyholt, 2004</xref>; <xref ref-type="bibr" rid="B35">Li and Ji, 2005</xref>; <xref ref-type="bibr" rid="B21">Gao et al., 2008</xref>; <xref ref-type="bibr" rid="B20">Galwey, 2009</xref>). A common feature of all methods is that they start by applying PCA to the association matrix. We use <italic>&#x3bb;</italic>
<sub>
<italic>i</italic>
</sub> to denote the <italic>i</italic>th eigenvalue extracted from the PCA.</p>
<p>The method proposed by <xref ref-type="bibr" rid="B10">Cheverud (2001)</xref> and <xref ref-type="bibr" rid="B51">Nyholt (2004)</xref> estimates the effective number of tests with<disp-formula id="e5">
<mml:math id="m9">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>CN</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>Var</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(5)</label>
</disp-formula>where Var(<italic>&#x3bb;</italic>) is the variance of the <italic>k</italic> eigenvalues. On the other hand, <xref ref-type="bibr" rid="B35">Li and Ji (2005)</xref> suggested the formula<disp-formula id="e6">
<mml:math id="m10">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>LJ</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
</mml:math>
<label>(6)</label>
</disp-formula>where<disp-formula id="e7">
<mml:math id="m11">
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>I</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2265;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x2b;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
<label>(7)</label>
</disp-formula>and &#x230a;&#x22c5;&#x230b; is the floor function. According to the method by <xref ref-type="bibr" rid="B21">Gao et al. (2008)</xref>, we first sort the eigenvalues in decreasing order, letting <italic>&#x3bb;</italic>
<sub>(1)</sub> denote the largest and <italic>&#x3bb;</italic>
<sub>(<italic>k</italic>)</sub> the smallest eigenvalue. Then the effective number of tests is defined as<disp-formula id="e8">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAO</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mfenced>
<mml:mspace width="0.28em"/>
<mml:mtext>such&#x2009;that</mml:mtext>
<mml:mspace width="0.28em"/>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>C</mml:mi>
<mml:mo>,</mml:mo>
</mml:math>
<label>(8)</label>
</disp-formula>where <italic>C</italic> is a user-defined parameter and usually chosen to be 0.995. Finally, <xref ref-type="bibr" rid="B20">Galwey (2009)</xref> proposed to estimate the effective number of tests with<disp-formula id="e9">
<mml:math id="m13">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAL</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msqrt>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(9)</label>
</disp-formula>where <inline-formula id="inf5">
<mml:math id="m14">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>max</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bb;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>.</p>
<p>All of the methods described above have the following desirable properties. When applied to an identity matrix (i.e., when there is no association between any pair of SNPs), then <italic>k</italic>
<sub>eff</sub> &#x3d; <italic>k</italic>, so that the effective number of tests is equal to the number of SNPs. An exception to this property can occur with <inline-formula id="inf6">
<mml:math id="m15">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAO</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. Depending on the value of <italic>C</italic> and the number of tests, it can happen that the effective number of tests is then estimated to be less than <italic>k</italic> (i.e., when <italic>k</italic> (1 &#x2212; <italic>C</italic>) &#x3e; 1 then <inline-formula id="inf7">
<mml:math id="m16">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>O</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3c;</mml:mo>
<mml:mi>k</mml:mi>
</mml:math>
</inline-formula>). On the other hand, when all of the SNPs are perfectly associated (i.e., the correlation matrix is equal to a <italic>k</italic> &#xd7; <italic>k</italic> matrix of 1&#x2019;s), then <italic>k</italic>
<sub>eff</sub> &#x3d; 1. In essence, the same test is then repeated <italic>k</italic> times, yielding identical results, so that effectively only a single test has been carried out. However, the methods differ in how association matrices that fall in between these two extremes are handled, yielding varying estimates of the effective number of tests between 1 and <italic>k</italic>.</p>
<p>Once <italic>k</italic>
<sub>eff</sub> has been estimated with one of these approaches, it can be used to adjust each of the methods for combining <italic>p</italic>-values described earlier. For the Bonferroni and Tippett&#x2019;s methods, we substitute <italic>k</italic>
<sub>eff</sub> for <italic>k</italic> so that<disp-formula id="e10">
<mml:math id="m17">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>,</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>.</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>and<disp-formula id="e11">
<mml:math id="m18">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>min</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(11)</label>
</disp-formula>are then the <italic>p</italic>-values for the gene. For the binomial test, we first define <inline-formula id="inf8">
<mml:math id="m19">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo>&#x230a;</mml:mo>
<mml:mrow>
<mml:mi>r</mml:mi>
<mml:mo>&#xd7;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
<mml:mo>&#x230b;</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> as the adjusted (i.e., effective) number of significant SNPs within the gene. Then the <italic>p</italic>-value for the gene is computed with<disp-formula id="e12">
<mml:math id="m20">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munderover>
<mml:mfenced open="(" close=")">
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:mfenced>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
<mml:mo>.</mml:mo>
</mml:math>
<label>(12)</label>
</disp-formula>Use of the floor function for computing <inline-formula id="inf9">
<mml:math id="m21">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula> may be conservative, but we consider this preferable over rounding and the risk of a too liberal test. Fisher&#x2019;s method can be adjusted by replacing the degrees of freedom of the chi-square distribution with 2<italic>k</italic>
<sub>eff</sub> and adjusting the test statistic with <inline-formula id="inf10">
<mml:math id="m22">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>. Hence, the <italic>p</italic>-value for the gene is then computed with <inline-formula id="inf11">
<mml:math id="m23">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>F</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>,</mml:mo>
<mml:mn>2</mml:mn>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>. Finally, for Stouffer&#x2019;s method, we let <inline-formula id="inf12">
<mml:math id="m24">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msqrt>
<mml:mrow>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:mfrac>
</mml:mrow>
</mml:msqrt>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>z</mml:mi>
</mml:math>
</inline-formula> denote the adjusted test statistic and hence <inline-formula id="inf13">
<mml:math id="m25">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mo>&#x303;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> is then the <italic>p</italic>-value for the gene.</p>
</sec>
<sec id="s2-4-2">
<title>2.3.2 Methods Based on Empirically-Derived Null Distributions</title>
<p>Another approach to account for dependence is to make use of permutation testing (<xref ref-type="bibr" rid="B30">Johnson et al., 2010</xref>; <xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>). The idea is to empirically derive the null distribution of the test statistic of interest by reshuffling the data in such a way that relevant features of the data structure are preserved except for the actual association being tested. For example, when testing for the association between each SNP and case-control status, reshuffling the status variable breaks any existing associations, but keeps the LD structure of the SNPs intact. Hence, any dependence among the <italic>p</italic>-values to be combined using one of the methods described earlier is automatically incorporated into the null distribution. The <italic>p</italic>-value for a gene is then computed from the percentile of the actually observed test statistic under the empirical null distribution. Note that in the present case, the test statistic of interest is actually a <italic>p</italic>-value itself, so letting <italic>p</italic>
<sub>
<italic>j</italic>
</sub> denote the combined <italic>p</italic>-value based on the <italic>j</italic>th permutation of the data (with <italic>j</italic> &#x3d; 1, &#x2026; , <italic>s</italic>) and <italic>p</italic>
<sub>obs</sub> the observed combined <italic>p</italic>-value, the <italic>p</italic>-value for a gene is then given by <inline-formula id="inf14">
<mml:math id="m26">
<mml:mi>p</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mi>I</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2264;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>obs</mml:mtext>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mi>s</mml:mi>
</mml:math>
</inline-formula>.</p>
<p>Permuting the data in the manner described above requires access to the raw data, so that the phenotype variable can be reshuffled and the test of association can be conducted for each SNP. In addition, repeatedly computing the test of association for each SNP within a gene can be computationally demanding. We can reduce the computational burden and eliminate the dependence upon the raw data by directly generating <italic>p</italic>-values based on an association matrix that reflects the degree of LD among the SNPs (<xref ref-type="bibr" rid="B42">Liu et al., 2010</xref>) which may be obtained from a reference population and not necessarily the given data.</p>
<p>In particular, let <italic>R</italic> denote the LD association matrix constructed from the correlations among the SNPs. We can then quickly generate a large number (<italic>s</italic>) of samples from a multivariate normal distribution with a true mean vector equal to zeros and covariance matrix <italic>R</italic>. Let <italic>Z</italic> denote the <italic>s</italic> &#xd7; <italic>k</italic> matrix of these values and <italic>P</italic> &#x3d; 2 (1 &#x2212; &#x3a6;(&#x7c;<italic>Z</italic>&#x7c;)) the matrix of two-sided <italic>p</italic>-values obtained by applying &#x3a6;(&#x22c5;) element-wise. For each row in <italic>P</italic>, we then apply one of the methods for combining <italic>p</italic>-values, yielding <italic>p</italic>
<sub>
<italic>j</italic>
</sub>. The <italic>p</italic>-value for a gene is then again computed as described above.</p>
</sec>
<sec id="s2-4-3">
<title>2.3.3 Methods Derived Under Dependence</title>
<p>The last set of methods we will consider are modifications of Fisher&#x2019;s and Stouffer&#x2019;s method so that dependence among the tests is directly taken into consideration.</p>
<sec id="s2-4-3-1">
<title>2.3.3.1 Brown&#x2019;s Method</title>
<p>The first adjustment is based on <xref ref-type="bibr" rid="B8">Brown (1975)</xref> who proposed a modification of Fisher&#x2019;s method for combining the results of correlated one-sided <italic>z</italic>-tests. If the <italic>p</italic>-values are not independent, <italic>X</italic>
<sup>2</sup> has expected value E(<italic>X</italic>
<sup>2</sup>) &#x3d; 2<italic>k</italic> and variance <inline-formula id="inf15">
<mml:math id="m27">
<mml:mi>Var</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>4</mml:mn>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mtext>Cov</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>, where the covariance between two &#x2212;2&#x2009;ln (&#x22c5;)-transformed <italic>p</italic>-values is given by<disp-formula id="e13">
<mml:math id="m28">
<mml:mtext>Cov</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>4</mml:mn>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x222c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>4</mml:mn>
<mml:mo>,</mml:mo>
</mml:math>
<label>(13)</label>
</disp-formula>where (<italic>z</italic>
<sub>
<italic>i</italic>
</sub>, <italic>z</italic>
<sub>
<italic>j</italic>
</sub>) is assumed to follow a bivariate standard normal distribution with correlation equal to the correlation among the two SNPs and <italic>f</italic>(<italic>z</italic>
<sub>
<italic>i</italic>
</sub>, <italic>z</italic>
<sub>
<italic>j</italic>
</sub>) denotes the joint probability density function of this distribution. The covariance term can be computed using numerical integration, although <xref ref-type="bibr" rid="B8">Brown (1975)</xref> also proposed a closed-form approximation that avoids this step. Next, we assume that <italic>X</italic>
<sup>2</sup> follows a scaled chi-squared distribution, i.e., <inline-formula id="inf16">
<mml:math id="m29">
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>c</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> (or equivalently, <inline-formula id="inf17">
<mml:math id="m30">
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>/</mml:mo>
<mml:mi>c</mml:mi>
<mml:mo>&#x223c;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>), where <inline-formula id="inf18">
<mml:math id="m31">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denotes a chi-squared distributed random variable with <italic>f</italic> degrees of freedom, and then approximate this distribution by equating its first two moments to the expected value and variance of <italic>X</italic>
<sup>2</sup> as calculated above. That is, for <inline-formula id="inf19">
<mml:math id="m32">
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>c</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c7;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, it follows that E (<italic>X</italic>
<sup>2</sup>) &#x3d; <italic>cf</italic> and Var(<italic>X</italic>
<sup>2</sup>) &#x3d; 2<italic>c</italic>
<sup>2</sup>
<italic>f</italic>, which implies <inline-formula id="inf20">
<mml:math id="m33">
<mml:mi>f</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>2</mml:mn>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mtext>E</mml:mtext>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>/</mml:mo>
<mml:mi>Var</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> and <italic>c</italic> &#x3d; Var(<italic>X</italic>
<sup>2</sup>)/2E (<italic>X</italic>
<sup>2</sup>). The <italic>p</italic>-value for a gene is then computed with <italic>p</italic> &#x3d; 1 &#x2212; <italic>F</italic>(<italic>X</italic>
<sup>2</sup>/<italic>c</italic>, <italic>f</italic>), where <italic>F</italic>(&#x22c5;, <italic>f</italic>) denotes the cumulative distribution function of a chi-square distribution with <italic>f</italic> degrees of freedom.</p>
<p>As given above, the method is only applicable to one-sided tests. However, in GWA studies, the association between the phenotype and the SNPs is typically examined with two-sided tests. We can easily extend Brown&#x2019;s method to two-sided tests by computing the covariance with<disp-formula id="e14">
<mml:math id="m34">
<mml:mtext>Cov</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mo>&#x2061;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>4</mml:mn>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x222c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:mi>ln</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>4</mml:mn>
<mml:mo>,</mml:mo>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>with (<italic>z</italic>
<sub>
<italic>i</italic>
</sub>, <italic>z</italic>
<sub>
<italic>j</italic>
</sub>) and <italic>f</italic> (<italic>z</italic>
<sub>
<italic>i</italic>
</sub>, <italic>z</italic>
<sub>
<italic>j</italic>
</sub>) as defined above (<xref ref-type="bibr" rid="B69">Yang et al., 2016</xref>). The remaining steps of the method are unchanged.</p>
</sec>
<sec id="s2-4-3-2">
<title>2.3.3.2 Strube&#x2019;s Method</title>
<p>Finally, Stouffer&#x2019;s method can also be generalized to consider the dependence among tests (<xref ref-type="bibr" rid="B60">Strube, 1985</xref>). To do so, we assume (as in Brown&#x2019;s method) that the test statistics that generated the <italic>p</italic>-values follow a multivariate normal distribution where the correlations among the test statistics are given by the correlations among the SNPs. We then compute<disp-formula id="e15">
<mml:math id="m35">
<mml:mi>z</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mi>Var</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(15)</label>
</disp-formula>where <inline-formula id="inf21">
<mml:math id="m36">
<mml:mi>Var</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>k</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mn>2</mml:mn>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3e;</mml:mo>
<mml:mi>i</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mtext>Cov</mml:mtext>
<mml:mfenced open="" close=")">
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:mfenced>
</mml:math>
</inline-formula>. The challenge is again the computation of the covariance term, which in this case is given by<disp-formula id="e16">
<mml:math id="m37">
<mml:mtext>Cov</mml:mtext>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#x3d;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x222c;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>&#x221e;</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mo>&#xd7;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi mathvariant="normal">&#x3a6;</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
<mml:mi>f</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi mathvariant="normal">d</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>z</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
<label>(16)</label>
</disp-formula>
</p>
<p>and which can again be computed using numerical integration. Then the combined <italic>p</italic>-value is calculated with <italic>p</italic> &#x3d; 1 &#x2212; &#x3a6;(<italic>z</italic>).</p>
</sec>
</sec>
</sec>
<sec id="s2-5">
<title>2.4 Illustrative Example</title>
<p>The methods for combining <italic>p</italic>-values described in the previous section can yield conflicting conclusions. In particular, while the (unadjusted) Bonferroni method controls the type I error rate even under dependence, it may fail to detect significant associations when combining non-independent <italic>p</italic>-values due to its conservative behaviour in such contexts. We present an example to illustrate this point.</p>
<p>
<xref ref-type="bibr" rid="B63">Assche et al. (2017)</xref> reported the results of a candidate gene study based on a sample of 982 Caucasian adolescents, analyzing 4,947 SNPs clustered in 263 genes known to be involved in neurotransmission. The outcome of interest was the (log-transformed) score on the Center for Epidemiologic Studies Depression Scale (<xref ref-type="bibr" rid="B54">Radloff, 1977</xref>). The association between each SNP and the outcome was tested using an additive model (<xref ref-type="bibr" rid="B32">Laird and Lange, 2010</xref>). The resulting <italic>p</italic>-values were then combined within each gene using Brown&#x2019;s method. The results showed that a small number of genes were significantly associated with the phenotype at <italic>&#x3b1;</italic> &#x3d; 0.05.</p>
<p>For illustration purposes, we obtained the combined <italic>p</italic>-values for two genes (<italic>GRID2IP</italic> and <italic>ARNTL2</italic>) with all of the methods described above. LD maps were calculated using the LD() function of the genetics package in R (<xref ref-type="bibr" rid="B64">Warnes et al., 2013</xref>), using the allelic correlation to measure the degree of association between the SNPs within each gene. The combined <italic>p</italic>-values were then obtained using the poolr package in R (<xref ref-type="bibr" rid="B12">Cinar and Viechtbauer, 2020</xref>). Empirical null distributions were generated as described earlier using <italic>s</italic> &#x3d; 10<sup>6</sup> samples.</p>
</sec>
<sec id="s2-6">
<title>2.5 Simulation Study</title>
<p>To compare the performance of the various methods more systematically, we conducted a simulation study based on HapMap phase II &#x2b; III data (<xref ref-type="bibr" rid="B61">The International HapMap Consortium, 2003</xref>) so that the results are representative of real genotype and LD information across the whole genome. Since genetic recombination breaks down disequilibria among the SNPs over time, LD tends to be weaker in older populations (<xref ref-type="bibr" rid="B31">Koch et al., 2013</xref>). We therefore used information from the TSI (Italian) sample from a somewhat younger population to avoid LD maps overwhelmed with negligible pairwise LD values. The sample contained <italic>n</italic> &#x3d; 102 individuals and 1,421,526 SNPs with their chromosome and position information.</p>
<p>We focus on autosomal chromosomes and excluded the sex chromosomes. Furthermore, insertions and deletions (INDELs) in the data (<xref ref-type="bibr" rid="B46">Mills et al., 2006</xref>) were removed. SNPs were assigned to genes using the biomaRt package in R through the Ensembl database (<xref ref-type="bibr" rid="B27">Hubbard et al., 2002</xref>; <xref ref-type="bibr" rid="B17">Durinck et al., 2005</xref>, <xref ref-type="bibr" rid="B18">2009</xref>). SNPs that were not assigned to a gene were excluded while SNPs that were assigned to multiple genes (due to overlapping genes) were kept in the study. After the assignment of SNPs to genes, the data included 915,259 SNPs in 30,910 genes. The number of SNPs per gene ranged from 2 to 3,178 with a mean (SD) of 29.61 (68.68) and a median of 12. Missing genotypes were then imputed using the MaCH software (<xref ref-type="bibr" rid="B37">Li et al., 2010</xref>). Next, LD maps (again using allelic correlations) were computed as described above. For LD maps that were not positive definite, the nearest positive definite correlation matrices were obtained with the nearPD() function of the Matrix package (<xref ref-type="bibr" rid="B4">Bates and Maechler, 2015</xref>). Finally, genotypes were coded in an additive manner (i.e., 0/1/2 coding), corresponding to the number of minor alleles at a locus (<xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>).</p>
<p>For each gene, we examined the type I error rate of the methods by 1) simulating a dichotomous phenotype variable (e.g., case-control status) for the <italic>n</italic> &#x3d; 102 individuals based on a Bernoulli distribution with <italic>&#x3c0;</italic> &#x3d; 0.50, 2) testing the association between each SNP within the gene and the phenotype variable using the Cochran-Armitage trend test (<xref ref-type="bibr" rid="B14">Cochran, 1954</xref>; <xref ref-type="bibr" rid="B2">Armitage, 1955</xref>), 3) combining the resulting <italic>p</italic>-values with each method described earlier, 4) repeating this process 1,000 times, and 5) calculating the proportion of times that the gene is declared significant at <italic>&#x3b1;</italic> &#x3d; 0.05 according to each method. For methods that make use of empirical distributions, we generated these distributions as described earlier using <italic>s</italic> &#x3d; 10<sup>5</sup> samples. Since the Cochran-Armitage trend test does not require that the Hardy-Weinberg equilibrium (HWE) assumption holds (<xref ref-type="bibr" rid="B32">Laird and Lange, 2010</xref>), we did not filter out SNPs that violate this assumption. Also, we note here that our goal was to examine the performance of the various methods based on individual genes, not at a whole genome-wide level. Therefore, we tested each gene at <italic>&#x3b1;</italic> &#x3d; 0.05, not at some level corrected for multiple testing. However, if a particular methods controls the type I error rate on individual genes, then it will also do so when testing all genes at <italic>&#x3b1;</italic> &#x3d; 0.05/<italic>g</italic>, where <italic>g</italic> denotes the total number of genes tested.</p>
<p>To examine the power of the methods, the same steps as described above were repeated, but the probability of having case status was now made a logistic function of one or multiple SNPs within the gene, that is, for the <italic>j</italic>th individual in the data set, we set the probability of having case status equal to<disp-formula id="e17">
<mml:math id="m38">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3c0;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>exp</mml:mi>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(17)</label>
</disp-formula>where <italic>x</italic>
<sub>
<italic>ij</italic>
</sub> denotes the number of minor alleles for the <italic>i</italic>th SNP of the <italic>j</italic>th individual, <italic>&#x3b2;</italic>
<sub>
<italic>i</italic>
</sub> determines how strongly the <italic>i</italic>th SNP is related to case-control status, and<disp-formula id="e18">
<mml:math id="m39">
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mo movablelimits="false" form="prefix">&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(18)</label>
</disp-formula>so that approximately half of the <italic>n</italic> individuals were assigned to case status and the other half were controls.</p>
<p>This part of the simulation involved several scenarios with different features. In the first set of conditions, a single SNP within a given gene was chosen in each iteration and the corresponding <italic>&#x3b2;</italic>
<sub>
<italic>i</italic>
</sub> value was set to either 0.2 or 0.5 (two different conditions). In the remaining conditions, we either allowed 5% or 20% of the SNPs within each gene to be associated with the phenotype. We again examined two different effect sizes (0.2 or 0.5) and either set <italic>&#x3b2;</italic>
<sub>
<italic>i</italic>
</sub> to the effect size value for each selected SNP (non-distributed effect) or distributed the effect over all selected SNPs (e.g., <italic>&#x3b2;</italic>
<sub>
<italic>i</italic>
</sub> &#x3d; 0.2/10 for an effect size of 0.2 and 10 selected SNPs). Finally, the selected SNPs were either positioned on a compact region of the gene (for this, a single SNP was randomly chosen among the first <italic>k</italic> &#xd7; (1 &#x2212; 0.05) or <italic>k</italic> &#xd7; (1 &#x2212; 0.20) SNPs and all consecutive SNPs were then also selected) or were dispersed throughout the gene (for this, significant SNPs were equally spaced throughout the gene). Hence, in addition to the first two conditions where a single SNP was associated with the phenotype with either an effect size of 0.2 or 0.5, we examined another 16 conditions, as all factors (5 vs. 20% of SNPs selected, effect size of 0.2 vs. 0.5, non-distributed vs. distributed effect, non-compact vs. compact SNP selection) were fully crossed.</p>
<p>To examine how sample size impacts the type I error rate and power of the methods, the same steps were repeated but in each iteration, a bootstrap sample of size 102, 500, or 1,000 was generated from the original data (together with the non-bootstrap conditions, we therefore examined 4 different sample size conditions). We included a condition with a bootstrap sample size of 102 to examine whether the performance of the methods differed whether the original data or a bootstrap sample of the same size was used. In total, we therefore examined a total of (1 <sub>type I error</sub> &#x2b; 2 <sub>Single SNP</sub> &#x2b; 16 <sub>Multiple SNPs</sub>) &#xd7; 4 <sub>Sample Size</sub> &#x3d; 76 different conditions.</p>
<p>The simulation was carried out using R (<xref ref-type="bibr" rid="B53">R Core Team, 2020</xref>) and was run on a cluster computer, making use of 144 cores (12 Intel Xeon E5-2,650 2.20&#xa0;GHz CPUs with 12 cores each) using parallel/multicore processing. Total computation time for the simulation was approximately 20,000 core hours.</p>
</sec>
</sec>
<sec id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Illustrative Example</title>
<p>For the illustrative example, <xref ref-type="table" rid="T1">Table 1</xref> presents the combined <italic>p</italic>-values for the two genes. Heat maps corresponding to the LD structure for these two genes (and the individual <italic>p</italic>-values for the SNPs) are provided in <xref ref-type="sec" rid="s10">Supplementary Figure S9</xref> as part of the supplementary materials. The table shows that the Bonferroni method fails to detect a significant association between both genes and the phenotype, whereas other approaches (including Brown&#x2019;s method) suggest a significant association. Interestingly, adjusting the Bonferroni method with two of the PCA-based methods (i.e., <inline-formula id="inf22">
<mml:math id="m40">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>LJ</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf23">
<mml:math id="m41">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAL</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>) leads to a significant finding at least for <italic>GRID2IP</italic>.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Combined <italic>p</italic>-values for the <italic>GRID2IP</italic> and <italic>ARNTL2</italic> genes based on the methods presented in <xref ref-type="sec" rid="s2">Section 2</xref> (combined <italic>p</italic>-values that show <italic>non-significant</italic> associations are accentuated in <italic>italic</italic>).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">
<italic>GRID2IP</italic>
</th>
<th rowspan="2" align="center">Unadjusted</th>
<th align="center">Cheverud-Nyholt</th>
<th align="center">Li and ji</th>
<th align="center">Gao</th>
<th align="center">Galwey</th>
<th rowspan="2" align="center">Empirically derived</th>
<th rowspan="2" align="center">Under dependence</th>
</tr>
<tr>
<th align="left">
<italic>k</italic> &#x3d; 23</th>
<th align="center">
<inline-formula id="inf24">
<mml:math id="m42">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>CN</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>20</mml:mn>
</mml:math>
</inline-formula>
</th>
<th align="center">
<inline-formula id="inf25">
<mml:math id="m43">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>LJ</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>15</mml:mn>
</mml:math>
</inline-formula>
</th>
<th align="center">
<inline-formula id="inf26">
<mml:math id="m44">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAO</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>18</mml:mn>
</mml:math>
</inline-formula>
</th>
<th align="center">
<inline-formula id="inf27">
<mml:math id="m45">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAL</mml:mtext>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>13</mml:mn>
</mml:math>
</inline-formula>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Bonferroni</td>
<td align="center">
<italic>0.068</italic>
</td>
<td align="center">
<italic>0.060</italic>
</td>
<td align="center">0.045</td>
<td align="center">
<italic>0.054</italic>
</td>
<td align="center">0.039</td>
<td align="center">
<italic>0.052</italic>
</td>
<td rowspan="3" align="left"/>
</tr>
<tr>
<td align="left">Tippett</td>
<td align="center">
<italic>0.066</italic>
</td>
<td align="center">
<italic>0.058</italic>
</td>
<td align="center">0.044</td>
<td align="center">
<italic>0.052</italic>
</td>
<td align="center">0.038</td>
<td align="center">
<italic>0.051</italic>
</td>
</tr>
<tr>
<td align="left">Binomial</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
</tr>
<tr>
<td align="left">Fisher</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">0.002</td>
<td align="center">0.001</td>
</tr>
<tr>
<td align="left">Stouffer</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">0.002</td>
<td align="center">&#x3c;0.001</td>
</tr>
</tbody>
</table>
<table>
<thead valign="top">
<tr>
<th align="left">&#xa0;<italic>
<bold>ARNTL2</bold>
</italic>
</th>
<th rowspan="2" align="center">
<bold>Unadjusted</bold>
</th>
<th align="center">
<bold>Cheverud-Nyholt</bold>
</th>
<th align="center">
<bold>Li and Ji</bold>
</th>
<th align="center">
<bold>Gao</bold>
</th>
<th align="center">
<bold>Galwey</bold>
</th>
<th rowspan="2" align="center">
<bold>Empirically Derived</bold>
</th>
<th rowspan="2" align="center">
<bold>Under Dependence</bold>
</th>
</tr>
<tr>
<td align="left">&#xa0;<bold>
<italic>k</italic> &#x3d; 24</bold>
</td>
<td align="center">
<inline-formula id="inf28">
<mml:math id="m46">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>C</mml:mi>
<mml:mi>N</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>22</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="center">
<inline-formula id="inf29">
<mml:math id="m47">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>L</mml:mi>
<mml:mi>J</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>14</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="center">
<inline-formula id="inf30">
<mml:math id="m48">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>O</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>18</mml:mn>
</mml:math>
</inline-formula>
</td>
<td align="center">
<inline-formula id="inf31">
<mml:math id="m49">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>e</mml:mi>
<mml:mi>f</mml:mi>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>G</mml:mi>
<mml:mi>A</mml:mi>
<mml:mi>L</mml:mi>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>14</mml:mn>
</mml:math>
</inline-formula>
</td>
</tr>
</thead>
<tbody>
<tr>
<td align="left">&#xa0;Bonferroni</td>
<td align="center">
<italic>0.112</italic>
</td>
<td align="center">
<italic>0.103</italic>
</td>
<td align="center">
<italic>0.065</italic>
</td>
<td align="center">
<italic>0.084</italic>
</td>
<td align="center">
<italic>0.065</italic>
</td>
<td align="center">
<italic>0.082</italic>
</td>
<td rowspan="3" align="left"/>
</tr>
<tr>
<td align="left">&#xa0;Tippett</td>
<td align="center">
<italic>0.106</italic>
</td>
<td align="center">
<italic>0.098</italic>
</td>
<td align="center">
<italic>0.063</italic>
</td>
<td align="center">
<italic>0.081</italic>
</td>
<td align="center">
<italic>0.063</italic>
</td>
<td align="center">
<italic>0.082</italic>
</td>
</tr>
<tr>
<td align="left">&#xa0;Binomial</td>
<td align="center">&#x3c;0.001</td>
<td align="center">0.001</td>
<td align="center">0.004</td>
<td align="center">0.002</td>
<td align="center">0.004</td>
<td align="center">0.011</td>
</tr>
<tr>
<td align="left">&#xa0;Fisher</td>
<td align="center">&#x3c;0.001</td>
<td align="center">&#x3c;0.001</td>
<td align="center">0.003</td>
<td align="center">0.001</td>
<td align="center">0.003</td>
<td align="center">0.020</td>
<td align="center">0.016</td>
</tr>
<tr>
<td align="left">&#xa0;Stouffer</td>
<td align="center">0.001</td>
<td align="center">0.001</td>
<td align="center">0.009</td>
<td align="center">0.003</td>
<td align="center">0.009</td>
<td align="center">0.030</td>
<td align="center">0.023</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn>
<p>The number of SNPs in the genes are denoted by <italic>k</italic>, whereas <inline-formula id="inf32">
<mml:math id="m50">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>CN</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf33">
<mml:math id="m51">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>LJ</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf34">
<mml:math id="m52">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAO</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and <inline-formula id="inf35">
<mml:math id="m53">
<mml:msubsup>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mtext>eff</mml:mtext>
</mml:mrow>
<mml:mrow>
<mml:mtext>GAL</mml:mtext>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> denote the effective number of tests estimated by the methods specified in the column header.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>While this example demonstrates that conclusions can differ depending on the method used, we do not know which of the conclusions drawn above are correct. In other words, the (non-significant) results of the Bonferroni method may be Type II errors (which are then avoided by using other methods) or they may be true negatives (with other methods then leading to Type I errors). The results of the simulation study will provide further insights into the performance of the methods when the true status of each gene is known.</p>
</sec>
<sec id="s3-2">
<title>3.2 Simulation Study</title>
<sec id="s3-2-1">
<title>3.2.1 Type I Error Rates</title>
<p>
<xref ref-type="fig" rid="F1">Figure 1</xref> shows boxplots of the type I error rates of all methods observed on all 30,910 genes applied to the original HapMap dataset (i.e., based on the non-bootstrapped data). Individual genes are indicated as points when their rate was more than 1.5 times the interquartile range above the third or below the first quartile. The mean rejection rates (and SDs) are also indicated in the figure.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Type I error rates of methods for gene-based testing when applied to the original HapMap data. The numbers above the boxplots show the mean (SD) rejection rates of the methods. The horizontal grey dashed line corresponds to the nominal rejection rate of <italic>&#x3b1;</italic> &#x3d; 0.05. Brown (1) and Brown (2) refer to the one- and two-sided versions of Brown&#x2019;s method (i.e., <xref ref-type="disp-formula" rid="e13">Eqs 13</xref>,<xref ref-type="disp-formula" rid="e14">14</xref>, respectively). Individual genes are indicated as points when their rejection rate was more than 1.5 times the interquartile range above the third or below the first quartile.</p>
</caption>
<graphic xlink:href="fgene-13-867724-g001.tif"/>
</fig>
<p>None of the unadjusted methods could achieve on average a nominal rejection rate of <italic>&#x3b1;</italic> &#x3d; 0.05. As expected, the Bonferroni method tended to be conservative, as was Tippett&#x2019;s method, which produced very similar results to the former method throughout the entire simulation study and which will therefore not be further discussed (note that rates above 0.05 for the Bonferroni method&#x2013;which occurred for about 1.5% of the genes&#x2013;reflect simulation error, since we know that the method guarantees that the type I error rate is equal to or less than <italic>&#x3b1;</italic> regardless of the degree of dependence). On the other hand, the remaining methods were generally liberal, at times dramatically so, with the binomial test at least providing an average rejection rate closest to the nominal level.</p>
<p>The adjustments for addressing the dependence did bring the rejection rates closer to the nominal level with varying degrees of success. In fact, when adjusted with the Li &#x26; Ji method, the Bonferroni method had a nominal average rejection rate, although this came at the cost of increased variability in the type I error rates, and the occurrence of rates well above the nominal level for particular genes. For the binomial test, the average rates fluctuated around the nominal level, being slightly conservative with the Li &#x26; Ji and Galwey adjustments and slightly liberal with the Nyholt and Gao adjustments. In contrast, none of the PCA-based adjustments could bring the average type I error rates of the Fisher and Stouffer methods sufficiently close to <italic>&#x3b1;</italic> &#x3d; 0.05.</p>
<p>Using empirically-derived null distributions produced rejection rates that were on average reasonably close to the nominal level, especially for Fisher&#x2019;s and Stouffer&#x2019;s methods. Moreover, the type I error rates of individual genes had much lower variability than the rates obtained with the PCA-based adjustments. This was also true for the Bonferroni method and the binomial test, but these methods were slightly conservative on average when adjusted in this manner.</p>
<p>The (two-sided) generalization of Fisher&#x2019;s method to dependent tests (i.e., Brown&#x2019;s method) yielded a nominal rejection rate on average. Furthermore, the variability (i.e., SD) of the rates for individual genes was lowest compared to all other methods. Quite importantly (mis)application of the one-sided version of Brown&#x2019;s method (since the <italic>p</italic>-values for the SNPs were computed from two-sided tests) resulted in worse performance (further references to Brown&#x2019;s method will therefore pertain to the two-sided version unless otherwise stated). On the other hand, the generalization of Stouffer&#x2019;s method to dependent tests (i.e., Strube&#x2019;s method) performed reasonably well, although its type I error rate was on average slightly inflated.</p>
<p>
<xref ref-type="sec" rid="s10">Supplementary Figures S1&#x2013;S3</xref> show the type I error rates of the methods based on bootstrap samples of sizes 102, 500, and 1,000, respectively. A comparison of <xref ref-type="fig" rid="F1">Figure 1</xref>; <xref ref-type="sec" rid="s10">Supplementary Figure S1</xref> (both with <italic>n</italic> &#x3d; 102) shows that the performance of the methods was similar regardless of whether they were applied to the original data or to bootstrap samples of the same size. The only exception to this was Stouffer&#x2019;s method, which became slightly more conservative for the bootstrapped data. Also, the patterns in the type I error rates of the methods were not fundamentally altered when applied to larger sample sizes. Using empirical distributions in combination with Fisher&#x2019;s and Stouffer&#x2019;s methods and Brown&#x2019;s method generally resulted in adequate control of the type I error rate on average and comparatively low variability in the rates for individual genes.</p>
<p>To examine whether the performance of the methods was affected by certain characteristics of the genes, we examined their type I error rates as a function of the (log transformed) number of SNPs in the genes, the average correlation in the LD maps, the degree of variability (i.e., standard deviation) of the correlations, the square-root of the mean squared correlations (SRMSC), the average minor allele frequencies (MAF) of the SNPs in the genes, and the standard deviation (SD) of the MAFs. The SRMSC was of particular interest as it distinguishes genes whose SNPs are independent (SRMSC equal to 0) from genes with SNPs in strong LD regardless of the directionality of the association (SRMSC close to 1). We used locally estimated scatterplot smoothing to visualize the relationship between these characteristics and the rejection rates for each method.</p>
<p>
<xref ref-type="fig" rid="F2">Figure 2</xref> shows that the type I error rate of many methods was affected by the number of SNPs in the genes. In particular, the Bonferroni method became increasingly conservative as the number of SNPs increased. Interestingly, this dependence on <italic>k</italic> was essentially removed when using the adjustment of Li &#x26; Ji and, to a slightly lesser extent, the adjustment of Gao. In contrast, the binomial test and Fisher&#x2019;s method became increasingly liberal as a function of <italic>k</italic>, whereas Stouffer&#x2019;s method displayed non-monotonic behaviour. The PCA-based adjustments helped to reduce the inflation in the type I error rates of these methods, but could not eliminate the dependence on <italic>k</italic>. Furthermore, all methods adjusted based on empirical distributions became increasingly conservative as the number of SNPs increased. Finally, Brown&#x2019;s method yielded essentially nominal rates regardless of <italic>k</italic>, except for very large genes, where the method became slightly conservative.</p>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Type I error rates of methods for gene-based testing as a function of the number of SNPs in the genes (log-transformed). The horizontal grey dashed line corresponds to the nominal rejection rate of <italic>&#x3b1;</italic> &#x3d; 0.05. Brown (1) and Brown (2) refer to the one- and two-sided versions of Brown&#x2019;s method (i.e., <xref ref-type="disp-formula" rid="e13">Eqs 13</xref>,<xref ref-type="disp-formula" rid="e14">14</xref>, respectively).</p>
</caption>
<graphic xlink:href="fgene-13-867724-g002.tif"/>
</fig>
<p>
<xref ref-type="fig" rid="F3">Figure 3</xref> shows the type I error rates as a function of the SRMSC values. As expected, the figure points out the increasingly conservative behavior of the Bonferroni method as the SNPs within the genes become more dependent, while Fisher&#x2019;s and Stouffer&#x2019;s methods then become liberal. The conservative behavior of the binomial test under independence is also expected (due to the discrete nature of the binomial distribution, the type I error rate of the test will not exceed <italic>&#x3b1;</italic> &#x3d; 0.05, but will often fall well below it). More surprising is the fact that the type I error rate of the method was essentially nominal for genes with very strong LD. To understand this phenomenon, consider a gene with <italic>k</italic> SNPs in perfect LD. In that case, all (two-sided) <italic>p</italic>-values are identical and hence either none or all <italic>k</italic> SNPs are significant. Since the latter will happen (under the joint null) with probability <italic>&#x3b1;</italic>, the test will exhibit nominal performance under this extreme scenario. With respect to the adjustments, the various PCA-based approaches again had the effect of counteracting the conservativeness of the Bonferroni method, while leading to a reduction in the type I error rates of the other methods. For all methods, adjusting with the use of empirical distributions was most successful when LD is strong, while Brown&#x2019;s method, although slightly conservative under independence, and performed well over the range of SRMSC values. Strube&#x2019;s method performed similarly, but with some slight inflation for larger SRMSC values.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Type I error rates of methods for gene-based testing as a function of the square-root mean squared correlation (SRMSC). The horizontal grey dashed line corresponds to the nominal rejection rate of <italic>&#x3b1;</italic> &#x3d; 0.05. Brown (1) and Brown (2) refer to the one- and two-sided versions of Brown&#x2019;s method (i.e., <xref ref-type="disp-formula" rid="e13">Eqs 13</xref>,<xref ref-type="disp-formula" rid="e14">14</xref>, respectively).</p>
</caption>
<graphic xlink:href="fgene-13-867724-g003.tif"/>
</fig>
<p>
<xref ref-type="sec" rid="s10">Supplementary Figures S4, S5</xref> show the type I error rates of the methods as a function of the average correlation and the SD of the correlations of the SNPs within the genes. Here, we again find that changes in these characteristics have essentially no impact on the performance of Brown&#x2019;s method, as well as when using empirical distributions in combination with Fisher&#x2019;s and Strube&#x2019;s methods. Interestingly, the one-sided version of Brown&#x2019;s method performed similarly to the two-sided one when the average LD was larger than <inline-formula id="inf36">
<mml:math id="m54">
<mml:mo>&#x2248;</mml:mo>
<mml:mn>0.3</mml:mn>
</mml:math>
</inline-formula>.</p>
<p>
<xref ref-type="sec" rid="s10">Supplementary Figures S6, S7</xref> display the type I error rates of the methods as a function of the average MAFs and their SD within the genes. The performance of Brown&#x2019;s method and Fisher&#x2019;s method with the empirical distribution adjustment was again not affected substantially by these factors except that these methods were slightly conservative when the average MAF was below 0.1 within the genes. Stouffer&#x2019;s method adjusted based on empirical distributions and its generalization to dependence was also relatively robust to both factors.</p>
</sec>
<sec id="s3-2-2">
<title>3.2.2 Statistical Power</title>
<p>
<xref ref-type="fig" rid="F4">Figure 4</xref> illustrates the power of the methods (averaged over genes) for the 54 conditions where the joint null hypothesis was false. Each panel corresponds to one of three sample size conditions (i.e., 102, 500, and 1,000), while the <italic>x</italic>-axis indicates the condition, starting with the two &#x201c;single SNP&#x201d; conditions (with effect sizes of 0.2 vs. 0.5) followed by the 16 &#x201c;multiple SNPs&#x201d; conditions (with either 5% or 20% of SNPs selected, an effect sizes of 0.2 or 0.5, a non-distributed or distributed effect, and either non-compact or compact SNP positions). We only show the power rates for the (unadjusted) Bonferroni method and those method and adjustment combinations that could control the type I error rate on average (i.e., the Bonferroni method with the Li &#x26; Ji adjustment, Brown&#x2019;s method, and use of empirical distributions in combination with Fisher&#x2019;s and Stouffer&#x2019;s methods).</p>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Power comparison of five methods for gene-based testing for three different sample sizes (<italic>n</italic> &#x3d; 102, 500, and 1000). The points represent the average rejection rate of a method under a specific scenario as indicated by the labels/symbols under the <italic>x</italic>-axis (number/percent of SNPs that are associated with the phenotype, the effect size, whether the effect was distributed over the significant SNPs, and whether significant SNPs were compactly positioned within the genes). Brown (2) refers to the two-sided version of Brown&#x2019;s method (i.e., <xref ref-type="disp-formula" rid="e13">Eq. 13</xref>).</p>
</caption>
<graphic xlink:href="fgene-13-867724-g004.tif"/>
</fig>
<p>As expected, power increased with the sample size, with the effect size, when a higher percentage of SNPs was associated with the phenotype, and when the effect was not distributed across these SNPs. Whether the selected SNPs fell into a compact region of the gene or were distributed throughout had comparatively little influence on the results. The (unadjusted) Bonferroni method typically had lower power compared to the other methods especially for lower sample sizes. Also, Stouffer&#x2019;s method adjusted with the use of empirical distributions tended to have slightly lower power compared to the other adjusted methods. While differences between the other methods were often negligible, a slight power advantage could be observed for the Bonferroni method adjusted with the Li &#x26; Ji correction when a single SNP or a low percentage of them contained a strong signal. Otherwise, Brown&#x2019;s method tended to have a slight power advantage.</p>
<p>To compare the computational efficiency of the methods, <xref ref-type="sec" rid="s10">Supplementary Figure S8</xref> presents the average computation times (based on 500 iterations) of the unadjusted methods along with the Li &#x26; Ji, empirical (using 10<sup>4</sup> samples), and dependence adjustments (the latter only for the Fisher and Stouffer methods) for a set of genes with {10, 25, 50, 100, 199, 254, 492, 897, and 1150} SNPs. While the unadjusted methods show no noteworthy increase in computational times as a function of the gene size, the results show that the use of the adjustments does come at the cost of increased computational times, less so for the Li &#x26; Ji adjustment and more so when using the Brown and Strube methods. Finally, as expected, the empirical methods demand the highest amount of computer time (although this can be mitigated to some extent; see <xref ref-type="bibr" rid="B13">Cinar and Viechtbauer (2022)</xref> for details). The computation times for the different test types, however, did not appear to vary substantially.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<title>4 Discussion</title>
<p>In this paper, we described some common methods for gene-based testing that combine the <italic>p</italic>-values of individual SNPs within genes (or that are clustered within some other higher-level functional structure) to test the joint null hypothesis that none of the SNPs within a gene are associated with the phenotype of interest. Along the way, we described a variety of adjustment techniques to incorporate LD information into this process. To examine and compare the type I error rates and power of the methods, we conducted an extensive simulation study based on HapMap data. While the (unadjusted) Bonferroni method guarantees that the type I error rate is never larger than the chosen significance level for all genes, the results show that this comes at the cost of a decrease in power for detecting genes that contain SNPs associated with the phenotype of interest.</p>
<p>Other methods for gene-based testing require adjustments based on the LD structure to ensure that their type I error rate is close to the nominal level on average. Doing so can increase the power for detecting &#x201c;significant genes&#x201d;, but this in turn can lead to an inflated type I error rate for some of the individual genes. We would consider this an acceptable risk under two conditions. First, the variability in the rates for individual genes should be low (to avoid excessively inflated type I error rates for particular genes). Moreover, the method should provide adequate control of the type I error rate regardless of the characteristics of the genes.</p>
<p>Among the various methods examined, the extension of Brown&#x2019;s method to two-sided tests comes closest to fulfilling these requirements. It had a nominal type I error rate on average and the lowest variability in the rates for individual genes (also when compared against the unadjusted Bonferroni method). The highest type I error rate observed across all 30,910 genes was 0.091, but this value might reflect at least in part simulation error, as the Bonferroni method also had inflated rates for 470 (1.5%) of the genes, with a maximum rate equal to 0.069. To further examine this, we repeated the simulation for these 470 genes using 10<sup>6</sup> iterations (see <xref ref-type="sec" rid="s10">Supplementary Figure S10</xref> for boxplots of the type I error rates of the Bonferroni and Brown&#x2019;s method). Now, only 34 of these genes still had a type I error rate above 0.05 with the Bonferroni method, with a maximum rate of 0.056. In contrast, the highest type I error rate of Brown&#x2019;s method was then 0.058, although (as expected) a higher number (188 out of these 470 genes) still had a rate above 0.05. Finally, the results based on all 30,910 genes showed that the Bonferroni method became increasingly conservative for genes with a larger number of SNPs or SNPs that were in stronger LD, while the performance of Brown&#x2019;s method was essentially independent of the various gene characteristics examined (except for some slight conservativeness when the degree of LD was very weak).</p>
<p>Another consideration in this context is the relative performance of the methods depending on whether the &#x2018;signal&#x2019; is concentrated in a single SNP or distributed over a larger number of them. The Bonferroni method&#x2013;which focuses on the lowest <italic>p</italic>-value among the SNPs within a gene&#x2013;might be at an advantage under the former scenario, while methods that can aggregate signals across multiple SNPs (such as Fisher&#x2019;s and Stouffer&#x2019;s method and versions thereof adjusted to account for dependence) would be expected to be more powerful in the latter case. However, under the conditions studied, the (unadjusted) Bonferroni method was never able to outperform Brown&#x2019;s method even when only a single SNP was strongly associated with the phenotype. Only when combined with the adjustment by Li &#x26; Ji did the Bonferroni method show a slight power advantage under this scenario. Brown&#x2019;s method may therefore be particularly advantageous when studying complex diseases where relatively small associations are likely to be spread across many SNPs and multiple genes (<xref ref-type="bibr" rid="B50">Neale and Sham, 2004</xref>; <xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>).</p>
<p>We also considered how an estimate of the effective number of tests can be used to adjust other methods besides the Bonferroni or Tippett methods (to which such adjustments are typically applied). However, none of these generalizations yielded nominal type I error rates on average. On the other hand, combining the Bonferroni method with the estimate of <xref ref-type="bibr" rid="B35">Li and Ji (2005)</xref> did perform adequately and, as mentioned above, may be of interest when the signal is concentrated in a single SNP. Our findings are in line with those by <xref ref-type="bibr" rid="B65">Wen and Lu (2011)</xref> who showed that the method by <xref ref-type="bibr" rid="B35">Li and Ji (2005)</xref> performs better than other effective number of tests adjustments.</p>
<p>Finally, we explored methods that mimic &#x201c;proper&#x201d; permutation tests by using pseudo replicates of the <italic>p</italic>-values to construct the empirical distributions needed for such tests. This approach greatly reduces the computation time (and can even be used when the raw data are not available) and produces results that are quite similar to those of conventional permutation techniques (<xref ref-type="bibr" rid="B38">Lin, 2005</xref>; <xref ref-type="bibr" rid="B15">Conneely and Boehnke, 2007</xref>; <xref ref-type="bibr" rid="B42">Liu et al., 2010</xref>; <xref ref-type="bibr" rid="B47">Moskvina et al., 2011</xref>). However, the results of our simulation study show that the performance of this approach depends on the method used for combining the <italic>p</italic>-values. Moreover, the type I error rate of these pseudo permutation tests either tended to be slightly conservative or, when the type I error rate was nominal on average, they offered no power advantage over Brown&#x2019;s method.</p>
<p>There are, however, a few issues that require further discussion. First, as mentioned at the beginning of <xref ref-type="sec" rid="s2">Section 2</xref>, the methods discussed in the present manuscript assume that the <italic>p</italic>-values follow a Uniform (0, 1) distribution under the null hypothesis. In our simulation study, the association between the SNPs and case-control status was tested using the Cochran-Armitage trend test (as is often done in practice when assuming an additive model). When using the typical normal approximation for conducting this test, the <italic>p</italic>-values only follow a uniform distribution asymptotically. While an exact version of this test is also available (<xref ref-type="bibr" rid="B67">Williams, 1988</xref>), the discrete nature of the test (since it is based on the frequency counts in a contingency table) can still make the exact <italic>p</italic>-values slightly conservative under the null hypothesis. In general though, the uniform assumption should hold when the sample size underlying the <italic>p</italic>-values is sufficiently large. Moreover, as this is a common issue for all of the methods described, it should not affect the relative performance of the methods.</p>
<p>Furthermore, in this paper, we focused on methods for combining <italic>p</italic>-values that can explicitly incorporate information from the LD matrix into their computation. Other recently proposed techniques for combining <italic>p</italic>-values, such as the Cauchy combination test (<xref ref-type="bibr" rid="B43">Liu et al., 2019</xref>; <xref ref-type="bibr" rid="B44">Liu and Xie, 2020</xref>) and the harmonic mean <italic>p</italic>-value method (<xref ref-type="bibr" rid="B68">Wilson, 2019</xref>), do not make use of LD information, but still provide control of the type I error rate under dependence. Now that the present results indicate the most advantageous methods that directly make use of the LD matrix, a further step will be a comparison of these method with those that do not.</p>
<p>Similarly, gene-based testing can also be conducted using modeling techniques (see <xref ref-type="bibr" rid="B9">Chapman and Whittaker, 2008</xref>; <xref ref-type="bibr" rid="B28">Ionita-Laza et al., 2013</xref>; <xref ref-type="bibr" rid="B48">Moskvina et al., 2012</xref>, for examples); however, such techniques require access to the raw genotype data. The focus of the present paper was on methods that avoid this requirement, but the relative performances of model-based methods and methods for combining <italic>p</italic>-values is an important subject to be examined in the future.</p>
<p>Finally, in our simulation, we focused on the gene regions in the HapMap data. As is well-known, SNPs in intergenic regions may play an important role in gene regulation and therefore may also be associated with a phenotype of interest (<xref ref-type="bibr" rid="B28">Ionita-Laza et al., 2013</xref>). Methods for combining <italic>p</italic>-values can also be utilized for synthesizing information from such genome regions as long as the <italic>p</italic>-values and LD matrices are derived accordingly. One could, for example, treat such regions as separate sets, or include intergenic SNPs with their neighboring genes.</p>
<p>In conclusion, the present results indicate that the two-sided version of Brown&#x2019;s method is a potentially attractive alternative to the use of the Bonferroni correction and other methods for gene-based testing. It is generally able to control the type I error rate and can lead to increased power, especially when associations are spread across multiple SNPs and genes. Those are the circumstances characterized by complex diseases where shifting the focus to higher functional structures may in fact be particularly advantageous.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>The original contributions presented in the study are included in the article/<xref ref-type="sec" rid="s10">Supplementary Material</xref>, further inquiries can be directed to the corresponding author.</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>WV conceived the initial idea. All authors contributed to writing the simulation code, carrying out the simulation study, processing the results, and writing the manuscript. All authors read and approved the final manuscript.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>Parts of this work were carried out on the Dutch national e-infrastructure with the support of SURF Cooperative.</p>
</ack>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2022.867724/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2022.867724/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.PDF" id="SM1" mimetype="application/PDF" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Alves</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Yu</surname>
<given-names>Y.-K.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Accuracy Evaluation of the Unified P-Value from Combining Correlated P-Values</article-title>. <source>PLoS One</source> <volume>9</volume>, <fpage>e91225</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.010366210.1371/journal.pone.0091225</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Armitage</surname>
<given-names>P.</given-names>
</name>
</person-group> (<year>1955</year>). <article-title>Tests for Linear Trends in Proportions and Frequencies</article-title>. <source>Biometrics</source> <volume>11</volume>, <fpage>375</fpage>&#x2013;<lpage>386</lpage>. <pub-id pub-id-type="doi">10.2307/3001775</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baranzini</surname>
<given-names>S. E.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Gibson</surname>
<given-names>R. A.</given-names>
</name>
<name>
<surname>Galwey</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Naegelin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Barkhof</surname>
<given-names>F.</given-names>
</name>
<etal/>
</person-group> (<year>2009</year>). <article-title>Genome-wide Association Analysis of Susceptibility and Clinical Phenotype in Multiple Sclerosis</article-title>. <source>Hum. Mol. Genet.</source> <volume>18</volume>, <fpage>767</fpage>&#x2013;<lpage>778</lpage>. <pub-id pub-id-type="doi">10.1093/hmg/ddn388</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bates</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Maechler</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>Matrix: Sparse and Dense Matrix Classes and Methods</article-title>. <comment>R package version 1.2-3</comment>. </citation>
</ref>
<ref id="B5">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Becker</surname>
<given-names>B. J.</given-names>
</name>
</person-group> (<year>1994</year>). &#x201c;<article-title>Combining Significance Levels</article-title>,&#x201d; in <source>The Handbook of Research Synthesis</source>. Editors <person-group person-group-type="editor">
<name>
<surname>Cooper</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Hedges</surname>
<given-names>L. V.</given-names>
</name>
</person-group> (<publisher-loc>New York</publisher-loc>: <publisher-name>Russell Sage Foundation</publisher-name>), <fpage>215</fpage>&#x2013;<lpage>230</lpage>. </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Benjamini</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Hochberg</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing</article-title>. <source>J. R. Stat. Soc. Ser. B (Methodological)</source> <volume>57</volume>, <fpage>289</fpage>&#x2013;<lpage>300</lpage>. <pub-id pub-id-type="doi">10.1111/j.2517-6161.1995.tb02031.x</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bland</surname>
<given-names>J. M.</given-names>
</name>
<name>
<surname>Altman</surname>
<given-names>D. G.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Multiple Significance Tests: The Bonferroni Method</article-title>. <source>Br. Med. J.</source> <volume>310</volume>, <fpage>170</fpage>. <pub-id pub-id-type="doi">10.1136/bmj.310.6973.170</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Brown</surname>
<given-names>M. B.</given-names>
</name>
</person-group> (<year>1975</year>). <article-title>400: A Method for Combining Non-independent, One-Sided Tests of Significance</article-title>. <source>Biometrics</source> <volume>31</volume>, <fpage>987</fpage>&#x2013;<lpage>992</lpage>. <pub-id pub-id-type="doi">10.2307/2529826</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chapman</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Whittaker</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Analysis of Multiple SNPs in a Candidate Gene or Region</article-title>. <source>Genet. Epidemiol.</source> <volume>32</volume>, <fpage>560</fpage>&#x2013;<lpage>566</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.20330</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cheverud</surname>
<given-names>J. M.</given-names>
</name>
</person-group> (<year>2001</year>). <article-title>A Simple Correction for Multiple Comparisons in Interval Mapping Genome Scans</article-title>. <source>Heredity</source> <volume>87</volume>, <fpage>52</fpage>&#x2013;<lpage>58</lpage>. <pub-id pub-id-type="doi">10.1046/j.1365-2540.2001.00901.x</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Chung</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Jun</surname>
<given-names>G. R.</given-names>
</name>
<name>
<surname>Dupuis</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Farrer</surname>
<given-names>L. A.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Comparison of Methods for Multivariate Gene-Based Association Tests for Complex Diseases Using Common Variants</article-title>. <source>Eur. J. Hum. Genet.</source> <volume>27</volume>, <fpage>811</fpage>&#x2013;<lpage>823</lpage>. <pub-id pub-id-type="doi">10.1038/s41431-018-0327-8</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cinar</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Viechtbauer</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Poolr: Methods for Pooling P-Values from (Dependent) Tests</article-title>. <comment>R package version 0.9-6</comment>. </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cinar</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Viechtbauer</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2022</year>). <article-title>The Poolr Package for Combining Independent and Dependent P Values</article-title>. <source>J. Stat. Softw.</source> <volume>12</volume>, <fpage>1</fpage>&#x2013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.18637/jss.v101.i01</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cochran</surname>
<given-names>W. G.</given-names>
</name>
</person-group> (<year>1954</year>). <article-title>Some Methods for Strengthening the Common <italic>&#x3c7;</italic>
<sup>2</sup> Tests</article-title>. <source>Biometrics</source> <volume>10</volume>, <fpage>417</fpage>&#x2013;<lpage>451</lpage>. <pub-id pub-id-type="doi">10.2307/3001616</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Conneely</surname>
<given-names>K. N.</given-names>
</name>
<name>
<surname>Boehnke</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2007</year>). <article-title>So many Correlated Tests, So Little Time! Rapid Adjustment of P Values for Multiple Correlated Tests</article-title>. <source>Am. J. Hum. Genet.</source> <volume>81</volume>, <fpage>1158</fpage>&#x2013;<lpage>1168</lpage>. <pub-id pub-id-type="doi">10.1086/522036</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dunn</surname>
<given-names>O. J.</given-names>
</name>
</person-group> (<year>1958</year>). <article-title>Estimation of the Means of Dependent Variables</article-title>. <source>Ann. Math. Stat.</source> <volume>29</volume>, <fpage>1095</fpage>&#x2013;<lpage>1111</lpage>. <pub-id pub-id-type="doi">10.1214/aoms/1177706443</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Durinck</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Moreau</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kasprzyk</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Davis</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Moor</surname>
<given-names>B. D.</given-names>
</name>
<name>
<surname>Brazma</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2005</year>). <article-title>BioMart and Bioconductor: A Powerful Link between Biological Databases and Microarray Data Analysis</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>3439</fpage>&#x2013;<lpage>3440</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti525</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Durinck</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Spellman</surname>
<given-names>P. T.</given-names>
</name>
<name>
<surname>Birney</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Huber</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>Mapping Identifiers for the Integration of Genomic Datasets with the R/Bioconductor Package biomaRt</article-title>. <source>Nat. Protoc.</source> <volume>4</volume>, <fpage>1184</fpage>&#x2013;<lpage>1191</lpage>. <pub-id pub-id-type="doi">10.1038/nprot.2009.97</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Fisher</surname>
<given-names>R. A.</given-names>
</name>
</person-group> (<year>1932</year>). <source>Statistical Methods for Researchers</source>. <edition>4th. ed</edition>. <publisher-loc>Edinburgh, UK</publisher-loc>: <publisher-name>Oliver &#x26; Boyd</publisher-name>. </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Galwey</surname>
<given-names>N. W.</given-names>
</name>
</person-group> (<year>2009</year>). <article-title>A New Measure of the Effective Number of Tests, a Practical Tool for Comparing Families of Non-independent Significance Tests</article-title>. <source>Genet. Epidemiol.</source> <volume>33</volume>, <fpage>559</fpage>&#x2013;<lpage>568</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.20408</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gao</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Starmer</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Martin</surname>
<given-names>E. R.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>A Multiple Testing Correction Method for Genetic Association Studies Using Correlated Single Nucleotide Polymorphisms</article-title>. <source>Genet. Epidemiol.</source> <volume>32</volume>, <fpage>361</fpage>&#x2013;<lpage>369</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.20310</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Goeman</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Solari</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2014</year>). <article-title>Multiple Hypothesis Testing in Genomics</article-title>. <source>Stat. Med.</source> <volume>33</volume>, <fpage>1946</fpage>&#x2013;<lpage>1978</lpage>. <pub-id pub-id-type="doi">10.1002/sim.6082</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hochberg</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>A Sharper Bonferroni Procedure for Multiple Tests of Significance</article-title>. <source>Biometrika</source> <volume>75</volume>, <fpage>800</fpage>&#x2013;<lpage>802</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/75.4.800</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Holm</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>1979</year>). <article-title>A Simple Sequentially Rejective Multiple Test Procedure</article-title>. <source>Scand. J. Stat.</source> <volume>6</volume>, <fpage>65</fpage>&#x2013;<lpage>70</lpage>. </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hommel</surname>
<given-names>G.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>A Stagewise Rejective Multiple Test Procedure Based on a Modified Bonferroni Test</article-title>. <source>Biometrika</source> <volume>75</volume>, <fpage>383</fpage>&#x2013;<lpage>386</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/75.2.383</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Huang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ellinghaus</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Franke</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Howie</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>1000 Genomes-Based Imputation Identifies Novel and Refined Associations for the Wellcome Trust Case Control Consortium Phase 1 Data</article-title>. <source>Eur. J. Hum. Genet.</source> <volume>20</volume>, <fpage>801</fpage>&#x2013;<lpage>805</lpage>. <pub-id pub-id-type="doi">10.1038/ejhg.2012.3</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Hubbard</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Barker</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Birney</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Cameron</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Clark</surname>
<given-names>L.</given-names>
</name>
<etal/>
</person-group> (<year>2002</year>). <article-title>The Ensembl Genome Database Project</article-title>. <source>Nucleic Acids Res.</source> <volume>30</volume>, <fpage>38</fpage>&#x2013;<lpage>41</lpage>. <pub-id pub-id-type="doi">10.1093/nar/30.1.38</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ionita-Laza</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Makarov</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Buxbaum</surname>
<given-names>J. D.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Sequence Kernel Association Tests for the Combined Effect of Rare and Common Variants</article-title>. <source>Am. J. Hum. Genet.</source> <volume>92</volume>, <fpage>841</fpage>&#x2013;<lpage>853</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2013.04.015</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jiao</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Peters</surname>
<given-names>U.</given-names>
</name>
<name>
<surname>Berndt</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Bezieau</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Brenner</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Campbell</surname>
<given-names>P. T.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Powerful Set-Based Gene-Environment Interaction Testing Framework for Complex Diseases</article-title>. <source>Genet. Epidemiol.</source> <volume>39</volume>, <fpage>609</fpage>&#x2013;<lpage>618</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.21908</pub-id> </citation>
</ref>
<ref id="B30">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Johnson</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Nelson</surname>
<given-names>G. W.</given-names>
</name>
<name>
<surname>Troyer</surname>
<given-names>J. L.</given-names>
</name>
<name>
<surname>Lautenberger</surname>
<given-names>J. A.</given-names>
</name>
<name>
<surname>Kessing</surname>
<given-names>B. D.</given-names>
</name>
<name>
<surname>Winkler</surname>
<given-names>C. A.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>Accounting for Multiple Comparisons in a Genome-wide Association Study (GWAS)</article-title>. <source>BMC Genomics</source> <volume>11</volume>, <fpage>724</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2164-11-724</pub-id> </citation>
</ref>
<ref id="B31">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Koch</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Ristroph</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kirkpatrick</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Long Range Linkage Disequilibrium across the Human Genome</article-title>. <source>PLoS One</source> <volume>8</volume>, <fpage>e80754</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0080754</pub-id> </citation>
</ref>
<ref id="B32">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Laird</surname>
<given-names>N. M.</given-names>
</name>
<name>
<surname>Lange</surname>
<given-names>C.</given-names>
</name>
</person-group> (<year>2010</year>). <source>The Fundamentals of Modern Statistical Genetics</source>. <publisher-loc>New York</publisher-loc>: <publisher-name>Springer</publisher-name>. </citation>
</ref>
<ref id="B33">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lancaster</surname>
<given-names>H. O.</given-names>
</name>
</person-group> (<year>1949</year>). <article-title>The Combination of Probabilities Arising from Data in Discrete Distributions</article-title>. <source>Biometrika</source> <volume>36</volume>, <fpage>370</fpage>&#x2013;<lpage>382</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/36.3-4.370</pub-id> </citation>
</ref>
<ref id="B34">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lehne</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lewis</surname>
<given-names>C. M.</given-names>
</name>
<name>
<surname>Schlitt</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>From SNPs to Genes: Disease Association at the Gene Level</article-title>. <source>PLoS One</source> <volume>6</volume>, <fpage>e20133</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pone.0020133</pub-id> </citation>
</ref>
<ref id="B35">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Ji</surname>
<given-names>L.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>Adjusting Multiple Testing in Multilocus Analyses Using the Eigenvalues of a Correlation Matrix</article-title>. <source>Heredity</source> <volume>95</volume>, <fpage>221</fpage>&#x2013;<lpage>227</lpage>. <pub-id pub-id-type="doi">10.1038/sj.hdy.6800717</pub-id> </citation>
</ref>
<ref id="B36">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>M. X.</given-names>
</name>
<name>
<surname>Gui</surname>
<given-names>H. S.</given-names>
</name>
<name>
<surname>Kwan</surname>
<given-names>J. S.</given-names>
</name>
<name>
<surname>Sham</surname>
<given-names>P. C.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>GATES: A Rapid and Powerful Gene-Based Association Test Using Extended Simes Procedure</article-title>. <source>Am. J. Hum. Genet.</source> <volume>88</volume>, <fpage>283</fpage>&#x2013;<lpage>293</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2011.01.019</pub-id> </citation>
</ref>
<ref id="B37">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Willer</surname>
<given-names>C. J.</given-names>
</name>
<name>
<surname>Ding</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Scheet</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Abecasis</surname>
<given-names>G. R.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>MaCH: Using Sequence and Genotype Data to Estimate Haplotypes and Unobserved Genotypes</article-title>. <source>Genet. Epidemiol.</source> <volume>434</volume>, <fpage>816</fpage>&#x2013;<lpage>834</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.20533</pub-id> </citation>
</ref>
<ref id="B38">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lin</surname>
<given-names>D. Y.</given-names>
</name>
</person-group> (<year>2005</year>). <article-title>An Efficient Monte Carlo Approach to Assessing Statistical Significance in Genomic Studies</article-title>. <source>Bioinformatics</source> <volume>21</volume>, <fpage>781</fpage>&#x2013;<lpage>787</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bti053</pub-id> </citation>
</ref>
<ref id="B39">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lipt&#xe1;k</surname>
<given-names>T.</given-names>
</name>
</person-group> (<year>1958</year>). <article-title>On the Combination of Independent Tests</article-title>. <source>Magyar Tud Akad Mat Kutato Int. Kozl</source> <volume>3</volume>, <fpage>171</fpage>&#x2013;<lpage>197</lpage>. </citation>
</ref>
<ref id="B40">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Littell</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Folks</surname>
<given-names>J. L.</given-names>
</name>
</person-group> (<year>1971</year>). <article-title>Asymptotic Optimality of Fisher&#x2019;s Method of Combining Independent Tests</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>66</volume>, <fpage>802</fpage>&#x2013;<lpage>806</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1971.10482347</pub-id> </citation>
</ref>
<ref id="B41">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Littell</surname>
<given-names>R. C.</given-names>
</name>
<name>
<surname>Folks</surname>
<given-names>J. L.</given-names>
</name>
</person-group> (<year>1973</year>). <article-title>Asymptotic Optimality of Fisher&#x2019;s Method of Combining Independent Tests II</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>68</volume>, <fpage>193</fpage>&#x2013;<lpage>194</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.1973.10481362</pub-id> </citation>
</ref>
<ref id="B42">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>J. Z.</given-names>
</name>
<name>
<surname>Mcrae</surname>
<given-names>A. F.</given-names>
</name>
<name>
<surname>Nyholt</surname>
<given-names>D. R.</given-names>
</name>
<name>
<surname>Medland</surname>
<given-names>S. E.</given-names>
</name>
<name>
<surname>Wray</surname>
<given-names>N. R.</given-names>
</name>
<name>
<surname>Brown</surname>
<given-names>K. M.</given-names>
</name>
<etal/>
</person-group> (<year>2010</year>). <article-title>A Versatile Gene-Based Test for Genome-wide Association Studies</article-title>. <source>Am. J. Hum. Genet.</source> <volume>87</volume>, <fpage>139</fpage>&#x2013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2010.06.009</pub-id> </citation>
</ref>
<ref id="B43">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>Z.</given-names>
</name>
<name>
<surname>Morrison</surname>
<given-names>A. C.</given-names>
</name>
<name>
<surname>Boerwinkle</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Lin</surname>
<given-names>X.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Acat: a Fast and Powerful P Value Combination Method for Rare-Variant Analysis in Sequencing Studies</article-title>. <source>Am. J. Hum. Genet.</source> <volume>104</volume>, <fpage>410</fpage>&#x2013;<lpage>421</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2019.01.002</pub-id> </citation>
</ref>
<ref id="B44">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Liu</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Xie</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Cauchy Combination Test: a Powerful Test with Analytic P-Value Calculation under Arbitrary Dependency Structures</article-title>. <source>J. Am. Stat. Assoc.</source> <volume>115</volume>, <fpage>393</fpage>&#x2013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1080/01621459.2018.1554485</pub-id> </citation>
</ref>
<ref id="B45">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Manolio</surname>
<given-names>T. A.</given-names>
</name>
</person-group> (<year>2010</year>). <article-title>Genomewide Association Studies and Assessment of the Risk of Disease</article-title>. <source>New Engl. J. Med.</source> <volume>363</volume>, <fpage>166</fpage>&#x2013;<lpage>176</lpage>. <pub-id pub-id-type="doi">10.1056/nejmra0905980</pub-id> </citation>
</ref>
<ref id="B46">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mills</surname>
<given-names>R. E.</given-names>
</name>
<name>
<surname>Luttig</surname>
<given-names>C. T.</given-names>
</name>
<name>
<surname>Larkins</surname>
<given-names>C. E.</given-names>
</name>
<name>
<surname>Beauchamp</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Tsui</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Pittard</surname>
<given-names>W. S.</given-names>
</name>
<etal/>
</person-group> (<year>2006</year>). <article-title>An Initial Map of Insertion and Deletion (INDEL) Variation in the Human Genome</article-title>. <source>Genome Res.</source> <volume>16</volume>, <fpage>1182</fpage>&#x2013;<lpage>1190</lpage>. <pub-id pub-id-type="doi">10.1101/gr.4565806</pub-id> </citation>
</ref>
<ref id="B47">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moskvina</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>O&#x2019;Dushlaine</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Purcell</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Craddock</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Holmans</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>O&#x2019;Donovan</surname>
<given-names>M. C.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Evaluation of an Approximation Method for Assessment of Overall Significance of Multiple Dependent Tests in a Genomewide Association Study</article-title>. <source>Genet. Epidemiol.</source> <volume>35</volume>, <fpage>861</fpage>&#x2013;<lpage>866</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.20636</pub-id> </citation>
</ref>
<ref id="B48">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Moskvina</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Schmidt</surname>
<given-names>K. M.</given-names>
</name>
<name>
<surname>Vedernikov</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Owen</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Craddock</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Holmans</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2012</year>). <article-title>Permutation-based Approaches Do Not Adequately Allow for Linkage Disequilibrium in Gene-wide Multi-Locus Association Analysis</article-title>. <source>Eur. J. Hum. Genet.</source> <volume>20</volume>, <fpage>890</fpage>&#x2013;<lpage>896</lpage>. <pub-id pub-id-type="doi">10.1038/ejhg.2012.8</pub-id> </citation>
</ref>
<ref id="B49">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Narum</surname>
<given-names>S. R.</given-names>
</name>
</person-group> (<year>2006</year>). <article-title>Beyond Bonferroni: Less Conservative Analyses for Conservation Genetics</article-title>. <source>Conservation Genet.</source> <volume>7</volume>, <fpage>783</fpage>&#x2013;<lpage>787</lpage>. <pub-id pub-id-type="doi">10.1007/s10592-006-9189-710.1007/s10592-005-9056-y</pub-id> </citation>
</ref>
<ref id="B50">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Neale</surname>
<given-names>B. M.</given-names>
</name>
<name>
<surname>Sham</surname>
<given-names>P. C.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>The Future of Association Studies: Gene-Based Analysis and Replication</article-title>. <source>Am. J. Hum. Genet.</source> <volume>75</volume>, <fpage>353</fpage>&#x2013;<lpage>362</lpage>. <pub-id pub-id-type="doi">10.1086/423901</pub-id> </citation>
</ref>
<ref id="B51">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Nyholt</surname>
<given-names>D. R.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>A Simple Correction for Multiple Testing for Single-Nucleotide Polymorphisms in Linkage Disequilibrium with Each Other</article-title>. <source>Am. J. Hum. Genet.</source> <volume>74</volume>, <fpage>765</fpage>&#x2013;<lpage>769</lpage>. <pub-id pub-id-type="doi">10.1086/383251</pub-id> </citation>
</ref>
<ref id="B52">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pearson</surname>
<given-names>E. S.</given-names>
</name>
</person-group> (<year>1938</year>). <article-title>The Probability Integral Transformation for Testing Goodness of Fit and Combining Independent Tests of Significance</article-title>. <source>Biometrika</source> <volume>30</volume>, <fpage>134</fpage>&#x2013;<lpage>148</lpage>. <pub-id pub-id-type="doi">10.2307/233222910.1093/biomet/30.1-2.134</pub-id> </citation>
</ref>
<ref id="B53">
<citation citation-type="book">
<collab>R Core Team</collab> (<year>2020</year>). <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna, Austria</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name>. </citation>
</ref>
<ref id="B54">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Radloff</surname>
<given-names>L. S.</given-names>
</name>
</person-group> (<year>1977</year>). <article-title>The CES-D Scale: A Self-Report Depression Scale for Research in the General Population</article-title>. <source>Appl. Psychol. Meas.</source> <volume>1</volume>, <fpage>385</fpage>&#x2013;<lpage>401</lpage>. <pub-id pub-id-type="doi">10.1177/014662167700100306</pub-id> </citation>
</ref>
<ref id="B55">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Shaffer</surname>
<given-names>J. P.</given-names>
</name>
</person-group> (<year>1995</year>). <article-title>Multiple Hypothesis Testing</article-title>. <source>Annu. Rev. Psychol.</source> <volume>46</volume>, <fpage>561</fpage>&#x2013;<lpage>584</lpage>. <pub-id pub-id-type="doi">10.1146/annurev.ps.46.020195.003021</pub-id> </citation>
</ref>
<ref id="B56">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>&#x160;id&#xe1;k</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>1957</year>). <article-title>Rectangular Confidence Regions for the Means of Multivariate normal Distributions</article-title>. <source>J. Am. Stat. Associations</source> <volume>62</volume>, <fpage>626</fpage>&#x2013;<lpage>633</lpage>. <pub-id pub-id-type="doi">10.2307/2283989</pub-id> </citation>
</ref>
<ref id="B57">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Simes</surname>
<given-names>R. J.</given-names>
</name>
</person-group> (<year>1986</year>). <article-title>An Improved Bonferroni Procedure for Multiple Tests of Significance</article-title>. <source>Biometrika</source> <volume>73</volume>, <fpage>751</fpage>&#x2013;<lpage>754</lpage>. <pub-id pub-id-type="doi">10.1093/biomet/73.3.751</pub-id> </citation>
</ref>
<ref id="B58">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Slatkin</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2008</year>). <article-title>Linkage Disequilibrium: Understanding the Evolutionary Past and Mapping the Medical Future</article-title>. <source>Nat. Rev. Genet.</source> <volume>9</volume>, <fpage>477</fpage>&#x2013;<lpage>485</lpage>. <pub-id pub-id-type="doi">10.1038/nrg2361</pub-id> </citation>
</ref>
<ref id="B59">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Stouffer</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Suchman</surname>
<given-names>E. A.</given-names>
</name>
<name>
<surname>Devinney</surname>
<given-names>L. C.</given-names>
</name>
<name>
<surname>Star</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>R. M.</given-names>
</name>
</person-group> (<year>1949</year>). <source>The American Soldier: Adjustment during Army Life (Studies in Social Psychology in World War II</source>, <volume>1</volume>. <publisher-loc>Princeton</publisher-loc>: <publisher-name>Princeton University Press</publisher-name>. </citation>
</ref>
<ref id="B60">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Strube</surname>
<given-names>M. J.</given-names>
</name>
</person-group> (<year>1985</year>). <article-title>Combining and Comparing Significance Levels from Nonindependent Hypothesis Tests</article-title>. <source>Psychol. Bull.</source> <volume>97</volume>, <fpage>334</fpage>&#x2013;<lpage>341</lpage>. <pub-id pub-id-type="doi">10.1037/0033-2909.97.2.334</pub-id> </citation>
</ref>
<ref id="B61">
<citation citation-type="journal">
<collab>The International HapMap Consortium</collab> (<year>2003</year>). <article-title>The International HapMap Project</article-title>. <source>Nature</source> <volume>426</volume>, <fpage>789</fpage>&#x2013;<lpage>796</lpage>. <pub-id pub-id-type="doi">10.1038/nature02168</pub-id> </citation>
</ref>
<ref id="B62">
<citation citation-type="book">
<person-group person-group-type="author">
<name>
<surname>Tippett</surname>
<given-names>L. H. C.</given-names>
</name>
</person-group> (<year>1931</year>). <source>The Methods of Statistics</source>. <publisher-loc>London</publisher-loc>: <publisher-name>Williams &#x26; Norgate</publisher-name>. </citation>
</ref>
<ref id="B63">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Van Assche</surname>
<given-names>E.</given-names>
</name>
<name>
<surname>Moons</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Cinar</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Viechtbauer</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Oldehinkel</surname>
<given-names>A. J.</given-names>
</name>
<name>
<surname>Van Leeuwen</surname>
<given-names>K.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Gene-based Interaction Analysis Shows GABA Ergic Genes Interacting with Parenting in Adolescent Depressive Symptoms</article-title>. <source>J. Child Psychol. Psychiatry</source> <volume>58</volume>, <fpage>1301</fpage>&#x2013;<lpage>1309</lpage>. <pub-id pub-id-type="doi">10.1111/jcpp.12766</pub-id> </citation>
</ref>
<ref id="B64">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Warnes</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Gorjanc</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Leisch</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Man</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2013</year>). <article-title>Genetics: Population Genetics</article-title>. <comment>R package version 1.3.8.1</comment>. </citation>
</ref>
<ref id="B65">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wen</surname>
<given-names>S. H.</given-names>
</name>
<name>
<surname>Lu</surname>
<given-names>Z. S.</given-names>
</name>
</person-group> (<year>2011</year>). <article-title>Factors Affecting the Effective Number of Tests in Genetic Association Studies: A Comparative Study of Three PCA-Based Methods</article-title>. <source>J. Hum. Genet.</source> <volume>56</volume>, <fpage>428</fpage>&#x2013;<lpage>435</lpage>. <pub-id pub-id-type="doi">10.1038/jhg.2011.34</pub-id> </citation>
</ref>
<ref id="B66">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wilkinson</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>1951</year>). <article-title>A Statistical Consideration in Psychological Research</article-title>. <source>Psychol. Bull.</source> <volume>48</volume>, <fpage>156</fpage>&#x2013;<lpage>158</lpage>. <pub-id pub-id-type="doi">10.1037/h0059111</pub-id> </citation>
</ref>
<ref id="B67">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Williams</surname>
<given-names>D. A.</given-names>
</name>
</person-group> (<year>1988</year>). <article-title>Tests for Differences between Several Small Proportions</article-title>. <source>J. R. Stat. Soc. Ser. C</source> <volume>37</volume>, <fpage>421</fpage>&#x2013;<lpage>434</lpage>. <pub-id pub-id-type="doi">10.2307/2347316</pub-id> </citation>
</ref>
<ref id="B68">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wilson</surname>
<given-names>D. J.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>The Harmonic Mean P-Value for Combining Dependent Tests</article-title>. <source>Proc. Natl. Acad. Sci.</source> <volume>116</volume>, <fpage>1195</fpage>&#x2013;<lpage>1200</lpage>. <pub-id pub-id-type="doi">10.1073/pnas.1814092116</pub-id> </citation>
</ref>
<ref id="B69">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Yang</surname>
<given-names>J. J.</given-names>
</name>
<name>
<surname>Li</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Williams</surname>
<given-names>L. K.</given-names>
</name>
<name>
<surname>Buu</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>An Efficient Genome-wide Association Test for Multivariate Phenotypes Based on the Fisher Combination Function</article-title>. <source>BMC Bioinformatics</source> <volume>17</volume>, <fpage>19</fpage>. <pub-id pub-id-type="doi">10.1186/s12859-015-0868-6</pub-id> </citation>
</ref>
<ref id="B70">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhang</surname>
<given-names>H.</given-names>
</name>
<name>
<surname>Tong</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Landers</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>Z.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>TFisher: A Powerful Truncation and Weighting Procedure for Combining <italic>P</italic>-Values</article-title>. <source>Ann. Appl. Stat.</source> <volume>14</volume>, <fpage>178</fpage>&#x2013;<lpage>201</lpage>. <pub-id pub-id-type="doi">10.1214/19-AOAS1302</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>