<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article article-type="review-article" dtd-version="2.3" xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Genet.</journal-id>
<journal-title>Frontiers in Genetics</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Genet.</abbrev-journal-title>
<issn pub-type="epub">1664-8021</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">745901</article-id>
<article-id pub-id-type="doi">10.3389/fgene.2021.745901</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Genetics</subject>
<subj-group>
<subject>Methods</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Using Summary Statistics to Model Multiplicative Combinations of Initially Analyzed Phenotypes With a Flexible Choice of Covariates</article-title>
<alt-title alt-title-type="left-running-head">Wolf&#x2009; et&#x20;al.</alt-title>
<alt-title alt-title-type="right-running-head">Modeling Multiplicative Combinations With PCSS</alt-title>
</title-group>
<contrib-group>
<contrib contrib-type="author">
<name>
<surname>Wolf&#x2009;</surname>
<given-names>Jack M.</given-names>
</name>
<xref ref-type="aff" rid="aff1">
<sup>1</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/1302098/overview"/>
</contrib>
<contrib contrib-type="author">
<name>
<surname>Westra</surname>
<given-names>Jason</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<uri xlink:href="https://loop.frontiersin.org/people/463371/overview"/>
</contrib>
<contrib contrib-type="author" corresp="yes">
<name>
<surname>Tintle</surname>
<given-names>Nathan</given-names>
</name>
<xref ref-type="aff" rid="aff2">
<sup>2</sup>
</xref>
<xref ref-type="aff" rid="aff3">
<sup>3</sup>
</xref>
<xref ref-type="corresp" rid="c001">&#x2a;</xref>
<uri xlink:href="https://loop.frontiersin.org/people/82507/overview"/>
</contrib>
</contrib-group>
<aff id="aff1">
<label>
<sup>1</sup>
</label>Division of Biostatistics, School of Public Health, University of Minnesota, <addr-line>Minneapolis</addr-line>, <addr-line>MN</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff2">
<label>
<sup>2</sup>
</label>Department of Mathematics, Computer Science, and Statistics, Dordt University, <addr-line>Sioux Center</addr-line>, <addr-line>IA</addr-line>, <country>United&#x20;States</country>
</aff>
<aff id="aff3">
<label>
<sup>3</sup>
</label>Department of Population Health Nursing Science, College of Nursing, University of Illinois Chicago, <addr-line>Chicago</addr-line>, <addr-line>IL</addr-line>, <country>United&#x20;States</country>
</aff>
<author-notes>
<fn fn-type="edited-by">
<p>
<bold>Edited by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/277268/overview">Pingzhao Hu</ext-link>, University of Manitoba, Canada</p>
</fn>
<fn fn-type="edited-by">
<p>
<bold>Reviewed by:</bold> <ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/659695/overview">Wenjian Bi</ext-link>, Peking University, China</p>
<p>
<ext-link ext-link-type="uri" xlink:href="https://loop.frontiersin.org/people/1236975/overview">Rounak Dey</ext-link>, Harvard University, United&#x20;States</p>
</fn>
<corresp id="c001">&#x2a;Correspondence: Nathan Tintle, <email>ntintle@uic.edu</email>
</corresp>
<fn fn-type="other">
<p>This article was submitted to Statistical Genetics and Methodology, a section of the journal Frontiers in Genetics</p>
</fn>
</author-notes>
<pub-date pub-type="epub">
<day>12</day>
<month>10</month>
<year>2021</year>
</pub-date>
<pub-date pub-type="collection">
<year>2021</year>
</pub-date>
<volume>12</volume>
<elocation-id>745901</elocation-id>
<history>
<date date-type="received">
<day>22</day>
<month>07</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>23</day>
<month>09</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#xa9; 2021 Wolf&#x2009;, Westra and Tintle.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Wolf&#x2009;, Westra and Tintle</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/">
<p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these&#x20;terms.</p>
</license>
</permissions>
<abstract>
<p>While the promise of electronic medical record and biobank data is large, major questions remain about patient privacy, computational hurdles, and data access. One promising area of recent development is pre-computing non-individually identifiable summary statistics to be made publicly available for exploration and downstream analysis. In this manuscript we demonstrate how to utilize pre-computed linear association statistics between individual genetic variants and phenotypes to infer genetic relationships between products of phenotypes (e.g., ratios; logical combinations of binary phenotypes using &#x201c;and&#x201d; and &#x201c;or&#x201d;) with customized covariate choices. We propose a method to approximate covariate adjusted linear models for products and logical combinations of phenotypes using only pre-computed summary statistics. We evaluate our method&#x2019;s accuracy through several simulation studies and an application modeling ratios of fatty acids using data from the Framingham Heart Study. These studies show consistent ability to recapitulate analysis results performed on individual level data including maintenance of the Type I error rate, power, and effect size estimates. An implementation of this proposed method is available in the publicly available R package <monospace>pcsstools</monospace>.</p>
</abstract>
<kwd-group>
<kwd>summary statistics</kwd>
<kwd>covariate adjustment</kwd>
<kwd>linear models</kwd>
<kwd>phenotype</kwd>
<kwd>multiplication</kwd>
</kwd-group>
<contract-num rid="cn001">2R15HG006915-03</contract-num>
<contract-sponsor id="cn001">National Human Genome Research Institute<named-content content-type="fundref-id">10.13039/100000051</named-content>
</contract-sponsor>
</article-meta>
</front>
<body>
<sec id="s1">
<title>1 Introduction</title>
<p>Researchers now have readily available access to massive quantities of genotypic and phenotypic data (<xref ref-type="bibr" rid="B4">Cox, 2018</xref>; <xref ref-type="bibr" rid="B22">Simell et&#x20;al., 2019</xref>). For example, <italic>via</italic> the Electronic Medical Records and Genomics {eMERGE Network<xref ref-type="fn" rid="fn1">
<sup>1</sup>
</xref>, the UKBiobank (<xref ref-type="bibr" rid="B2">Bycroft et&#x20;al., 2018</xref>) other initiatives and repositories [e.g., 23andMe, MGI<xref ref-type="fn" rid="fn2">
<sup>2</sup>
</xref> (<xref ref-type="bibr" rid="B8">Gagliano Taliun et&#x20;al., 2020</xref>), FINRISK, CHOP (<xref ref-type="bibr" rid="B5">Diogo et&#x20;al., 2018</xref>), among others]}, researchers can access a wide variety of phenotypic and genomics data on hundreds of thousands of individuals. However, important questions remain about how to best leverage these repositories. For example, the size of biobank datasets makes it challenging to transfer, store, and analyze data locally. While cloud computing minimizes some of these issues, it brings its own challenges related to cost (storage and computation), transfer, and access. Furthermore, data security and privacy issues are of paramount importance throughout all aspects of the data access, storage, and analysis pipeline (<xref ref-type="bibr" rid="B13">Jones et&#x20;al., 2012</xref>; <xref ref-type="bibr" rid="B11">Heatherly, 2016</xref>; <xref ref-type="bibr" rid="B22">Simell et&#x20;al., 2019</xref>).</p>
<p>A key innovation in this field is pre-computing non-individually identifiable summary statistics on biobank data and maximizing access to this data (<xref ref-type="bibr" rid="B20">Pasaniuc and Price, 2017</xref>). For example, GeneAtlas provides basic summary statistics for simple linear regression models of single nucleotide variants (SNVs) with 1,000s of available phenotypic variables across hundreds of thousands of individuals in the UK Biobank (<xref ref-type="bibr" rid="B3">Canela-Xandri et&#x20;al., 2018</xref>), which also provides access to phenotype-phenotype correlations, single nucleotide polymorphism (SNP) minor allele frequencies (MAFs) and Hardy Weinberg Equilibrium (HWE) <italic>p</italic>-values. Likewise, PheWeb<xref ref-type="fn" rid="fn3">
<sup>3</sup>
</xref> is a software toolkit which provides access to UK Biobank and Michigan Genomics Initiative data <italic>via</italic> a series of easy-to-navigate visualization and summary tools (<xref ref-type="bibr" rid="B8">Gagliano Taliun et&#x20;al., 2020</xref>). Others (e.g., The Lee Lab for Statistical Genetics and Data Science<xref ref-type="fn" rid="fn4">
<sup>4</sup>
</xref>) simply provide access to sets of pre-computed summary statistics (PCSS) from large datasets. These resources mitigate many of the privacy and security concerns mentioned above since no individual participant data (IPD) is shared. In addition, the size of these repositories are only fractions of the size of IPD, making transfer and storage of the data much more efficient. Finally, pre-computing these summary statistics alleviates much of the computational burden on researchers who would otherwise have to calculate this information on their own. Despite these advantages, significant limitations currently exist when using these repositories of&#x20;PCSS.</p>
<p>For example, researchers may want to modify a phenotype with available PCSS to one that is of greater clinical interest or use different sets of covariates than those considered in pre-computed analyses. Recent work is beginning to address these limitations. In two recent papers by our group (<xref ref-type="bibr" rid="B9">Gasdaska et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B26">Wolf et&#x20;al., 2020</xref>), we demonstrated how to use standard PCSS (only means, variances, and correlations of all predictors and responses) to calculate the coefficients and standard errors for the linear model for a linear combination of phenotypes with an arbitrary set of covariates. This can then be used to perform Principal Component Analysis (PCA) on a set of phenotypes since principal component scores are just linear combinations with weights derived from the phenotype covariance matrix. Further, we demonstrated that if the phenotype correlation matrix is not available, we can use the correlation of test statistics for each phenotype across all genetic markers in its place with little loss of efficiency. These innovations mean that researchers can, using only PCSS, select the unique set of covariates they wish to adjust for and model a linear combination of phenotypes.</p>
<p>Importantly, these two approaches which require a priori specification of a phenotype of clinical interest, contrast to other recently developed methods which jointly and simultaneously analyze multiple phenotypes (<xref ref-type="bibr" rid="B6">Dutta et&#x20;al., 2019a</xref>,<xref ref-type="bibr" rid="B7">b</xref>; <xref ref-type="bibr" rid="B10">Guo and Wu, 2019</xref>; <xref ref-type="bibr" rid="B18">Li et&#x20;al., 2020</xref>; <xref ref-type="bibr" rid="B21">Ray and Boehnke, 2018</xref>) without an explicit characterization of the relationship between the phenotypes. These joint phenotype tests aim to simultaneously analyze multiple phenotypes while satisfying statistical objectives such as maximizing power under certain conditions. Furthermore, some of these approaches (<xref ref-type="bibr" rid="B21">Ray and Boehnke, 2018</xref>; <xref ref-type="bibr" rid="B10">Guo and Wu, 2019</xref>) do so using PCSS readily available from existing repositories.</p>
<p>Currently, our group&#x2019;s methods for using PCSS to analyze modified phenotypes with flexible covariate choices are limited to PCA and choosing a phenotype that is a linear combination of the phenotypes for which PCSS are available. Another meaningful way to combine phenotypes is through multiplication. That is, several phenotypes of interest can be viewed as multiplicative combinations of other phenotypes for which PCSS may be available. Examples of note include fatty acid conversion ratios (<xref ref-type="bibr" rid="B15">Kalsbeek et&#x20;al., 2018</xref>) and the body mass index (<xref ref-type="bibr" rid="B14">Justice et&#x20;al., 2017</xref>). Additionally, products of binary phenotypes can be interpreted as logical &#x201c;and&#x201d; and &#x201c;or&#x201d; statements (e.g., a phenotype <bold>
<italic>y</italic>
</bold> that is defined as &#x201c;<bold>
<italic>y</italic>
</bold>
<sub>1</sub> or <bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x201d;). Various medical conditions are defined through logical combinations of various phenotypes. For example, coding ischemic strokes based on the union and intersection of various stroke subtypes (<xref ref-type="bibr" rid="B25">von Berg et&#x20;al., 2020</xref>).</p>
<p>In this manuscript, we demonstrate how to analyze modified phenotypes which are multiplicative combinations of an arbitrarily large number of phenotypes for which PCSS are available. We also demonstrate how to flexibly adjust for covariates in these modified phenotype models. Importantly, we also show how the multiplication of phenotypes, when applied to binary phenotypes, allows for logical combination of phenotypes. After presenting a mathematical framework for the method, we validate the method using comprehensive simulations and demonstrate the method on real data from the Framingham Heart Study.</p>
</sec>
<sec sec-type="methods" id="s2">
<title>2 Methods</title>
<p>Consider the <italic>m</italic> phenotypes <bold>
<italic>y</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>y</italic>
</bold>
<sub>
<italic>m</italic>
</sub> where each is an <italic>n</italic>&#x20;&#xd7; 1 vector of measures across <italic>n</italic> subjects and the <italic>n</italic>&#x20;&#xd7; <italic>p</italic> design matrix <bold>
<italic>X</italic>
</bold>&#x20;&#x3d; (<bold>
<italic>x</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>x</italic>
</bold>
<sub>
<italic>p</italic>
</sub>) which consists of variables including genotypic information, covariates, and an intercept column. Moreover, let <bold>
<italic>w</italic>
</bold>
<sub>
<italic>m</italic>
</sub> &#x3d; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x22ef;<bold>
<italic>y</italic>
</bold>
<sub>
<italic>m</italic>
</sub> denote the pairwise Hadamard product of all <italic>m</italic> phenotypes for each subject. Our aim is to approximate the coefficients and standard errors of the covariate adjusted linear regression model for the product of <italic>m</italic> phenotypes: <inline-formula id="inf1">
<mml:math id="m1">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:math>
</inline-formula>&#x3f5; using only readily available&#x20;PCSS.</p>
<sec id="s2-1">
<title>2.1 Assumed Pre-Computed Summary Statistics and Information</title>
<p>We assume knowledge of the following PCSS: the means of every predictor (e.g., SNPs and covariates), the means of every phenotype, and the full variance-covariance matrix of all predictors and phenotypes (i.e.,&#x20;<inline-formula id="inf2">
<mml:math id="m2">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf3">
<mml:math id="m3">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf4">
<mml:math id="m4">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> for any <italic>i</italic>, <italic>j</italic>, <italic>k</italic>, <italic>l</italic> where 1 &#x2264; <italic>i</italic>, <italic>j</italic>&#x20;&#x2264; <italic>p</italic> and 1 &#x2264; <italic>k</italic>, <italic>l</italic>&#x20;&#x2264; <italic>m</italic>). These are all readily available in standard PCSS repositories. We also assume to know the marginal distribution that each predictor and phenotype follows (e.g., binomial, normal, etc.). <xref ref-type="fig" rid="F1">Figure&#x20;1</xref> displays the assumed information when modeling <italic>via</italic> both IPD and&#x20;PCSS.</p>
<fig id="F1" position="float">
<label>FIGURE 1</label>
<caption>
<p>Data assumed when modeling using individual participant data (IPD) and when using pre-computed summary statistics (PCSS) to model a product of <italic>m</italic> phenotypes (<bold>
<italic>y</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>y</italic>
</bold>
<sub>
<italic>m</italic>
</sub>) as a linear function of <italic>p</italic> covariates (<bold>
<italic>x</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>x</italic>
</bold>
<sub>
<italic>m</italic>
</sub>). While modeling using IPD requires <italic>n</italic>&#x20;&#xd7; (<italic>p</italic>&#x20;&#x2b; <italic>m</italic>) points of data, using PCSS only requires <italic>p</italic>
<sup>2</sup> &#x2b; <italic>pm</italic> &#x2b; <italic>m</italic>
<sup>2</sup> &#x2b; <italic>p</italic>&#x20;&#x2b; <italic>m</italic> values, which is far less when <italic>n</italic> is moderately large compared to <italic>p</italic> and <italic>m</italic>. All of these PCSS are readily available in existing PCSS repositories, or can be derived or approximated from other PCSS.</p>
</caption>
<graphic xlink:href="fgene-12-745901-g001.tif"/>
</fig>
<p>However, if some summary statistics are unknown, they may be able to be derived or approximated. For example, SNPs distributed in HWE can have their mean and variance approximated through a binomial distribution given the MAF. Furthermore, the covariance of a genetic variant and a non-genetic variable is calculated as the single-marker slope coefficient (for the model with the non-genetic variable as the response and the genetic variant as the predictor) divided by the variance of the genetic variant. Other published papers (<xref ref-type="bibr" rid="B16">Kim et&#x20;al., 2015</xref>; <xref ref-type="bibr" rid="B29">Zhu et&#x20;al., 2015</xref>) have shown that the correlation of two traits can be approximated by the correlation of <italic>Z</italic> statistics of SNPs not associated with either trait. This approximation method is described in detail in <xref ref-type="bibr" rid="B21">Ray and Boehnke (2018)</xref>. Two of our previous papers (<xref ref-type="bibr" rid="B9">Gasdaska et&#x20;al., 2019</xref>; <xref ref-type="bibr" rid="B26">Wolf et&#x20;al., 2020</xref>) have demonstrated the accuracy of these three methods through both simulation and real-data applications.</p>
</sec>
<sec id="s2-2">
<title>2.2 Linear Regression With Covariates Using Pre-Computed Summary Statistics</title>
<p>Given a response vector <bold>
<italic>w</italic>
</bold>
<sub>
<italic>m</italic>
</sub> and design matrix <bold>
<italic>X</italic>
</bold> &#x3d; (<bold>
<italic>x</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>x</italic>
</bold>
<sub>
<italic>p</italic>
</sub>) which includes <italic>p</italic> variables including SNPs&#x2019; minor allele counts, covariates, and a possible intercept column, the normal error regression model <inline-formula id="inf5">
<mml:math id="m5">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
<mml:mo>&#x2b;</mml:mo>
<mml:mi mathvariant="bold-italic">&#x025B;</mml:mi>
</mml:math>
</inline-formula> where &#x3f5;<inline-formula id="inf6">
<mml:math id="m6">
<mml:mo>&#x223c;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn mathvariant="bold-italic">0</mml:mn>
<mml:mo>,</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="bold-italic">I</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> has ordinary least squares estimate for <inline-formula id="inf7">
<mml:math id="m7">
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:math>
</inline-formula>: <inline-formula id="inf8">
<mml:math id="m8">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> with <inline-formula id="inf9">
<mml:math id="m9">
<mml:mi>Var</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula>. In a recent paper (<xref ref-type="bibr" rid="B26">Wolf et&#x20;al., 2020</xref>), we demonstrated how to calculate these values using only PCSS using the facts that:<disp-formula id="e1">
<mml:math id="m10">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>S</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:math>
<label>(1)</label>
</disp-formula>
<disp-formula id="e2">
<mml:math id="m11">
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
<label>(2)</label>
</disp-formula>
<disp-formula id="e3">
<mml:math id="m12">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
<label>(3)</label>
</disp-formula>and<disp-formula id="e4">
<mml:math id="m13">
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">X</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(4)</label>
</disp-formula>where <italic>S</italic>(<bold>
<italic>X</italic>
</bold>) is the <italic>p</italic>&#x20;&#xd7; <italic>p</italic> variance-covariance matrix of the columns of the design matrix <bold>
<italic>X</italic>
</bold>, <inline-formula id="inf10">
<mml:math id="m14">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:mo>&#x2026;</mml:mo>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> is the <italic>p</italic>&#x20;&#xd7; 1 vector of column means of <bold>
<italic>X</italic>
</bold>, <inline-formula id="inf11">
<mml:math id="m15">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the mean of <bold>
<italic>w</italic>
</bold>
<sub>
<italic>m</italic>
</sub>, and <inline-formula id="inf12">
<mml:math id="m16">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> is the sample covariance between <bold>
<italic>w</italic>
</bold>
<sub>
<italic>m</italic>
</sub> and&#x20;<bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>.</p>
<p>With these methods in mind and assumed access to standard PCSS, in order to approximate <inline-formula id="inf13">
<mml:math id="m17">
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
</mml:math>
</inline-formula>, and <inline-formula id="inf14">
<mml:math id="m18">
<mml:mi>Var</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> for this covariate adjusted multiple linear regression model, all that remains is to estimate <inline-formula id="inf15">
<mml:math id="m19">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf16">
<mml:math id="m20">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and <inline-formula id="inf17">
<mml:math id="m21">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>m</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> for each <italic>j</italic>. We will first demonstrate how to approximate these values using PCSS when <italic>m</italic>&#x20;&#x3d; 2 in <xref ref-type="sec" rid="s2-3-1">Section 2.3.1</xref> and later show how recursion can be used when <italic>m</italic>&#x20;&#x3e; 2 in <xref ref-type="sec" rid="s2-3-2">Section&#x20;2.3.2</xref>.</p>
</sec>
<sec id="s2-3">
<title>2.3 Covariance Estimation</title>
<sec id="s2-3-1">
<title>2.3.1 Covariance Estimation With the Product of 2 Phenotypes</title>
<p>Let <bold>
<italic>w</italic>
</bold>
<sub>2</sub> &#x3d; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>2</sub> be the pairwise Hadamard product of <bold>
<italic>y</italic>
</bold>
<sub>1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub>. Then, if <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> represents an &#x201c;intercept&#x201d; column of the design matrix with all elements unity [i.e.,&#x20;if <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> &#x3d; (1,&#x2026;,1)&#x2032;], we set <inline-formula id="inf18">
<mml:math id="m22">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:math>
</inline-formula>. Otherwise, we proceed as follows: We first approximate the conditional mean and variance of <bold>
<italic>y</italic>
</bold>
<sub>
<italic>k</italic>
</sub> given <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> &#x3d; <italic>x</italic> for <italic>k</italic>&#x20;&#x2208; {1, 2} through a linear regression model as<disp-formula id="e5">
<mml:math id="m23">
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mi>x</mml:mi>
</mml:math>
<label>(5)</label>
</disp-formula>and<disp-formula id="e6">
<mml:math id="m24">
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(6)</label>
</disp-formula>where <inline-formula id="inf19">
<mml:math id="m25">
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>/</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf20">
<mml:math id="m26">
<mml:msub>
<mml:mrow>
<mml:mi>a</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>b</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. We note that this conditional variance will be constant at any value of <italic>x</italic> following from the linear regression assumption of homoscedasticity.</p>
<p>Then, we calculate the sample partial correlation of <bold>
<italic>y</italic>
</bold>
<sub>1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub> controlling for <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>:<disp-formula id="e7">
<mml:math id="m27">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>.</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mfrac>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:msqrt>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:mrow>
</mml:mfrac>
<mml:mo>,</mml:mo>
</mml:math>
<label>(7)</label>
</disp-formula>setting <inline-formula id="inf21">
<mml:math id="m28">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>.</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:math>
</inline-formula> if either <inline-formula id="inf22">
<mml:math id="m29">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> or <inline-formula id="inf23">
<mml:math id="m30">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>. As the expectation of the conditional correlation equals the partial correlation under the assumption of a multivariate linear relationship between (<bold>
<italic>y</italic>
</bold>
<sub>1</sub>, <bold>
<italic>y</italic>
</bold>
<sub>2</sub>) and <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> (<xref ref-type="bibr" rid="B1">Baba et&#x20;al., 2004</xref>), we use the partial correlation as an estimate of the conditional correlation of <bold>
<italic>y</italic>
</bold>
<sub>1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub> at all possible values of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>. So, we approximate the covariance of <bold>
<italic>y</italic>
</bold>
<sub>1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub> conditional on <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>:<disp-formula id="e8">
<mml:math id="m31">
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>.</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:msqrt>
<mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
</mml:msqrt>
</mml:math>
<label>(8)</label>
</disp-formula>
</p>
<p>These terms let us approximate the conditional mean of <bold>
<italic>w</italic>
</bold>
<sub>2</sub> at a given value <italic>x</italic> of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>:<disp-formula id="e9">
<mml:math id="m32">
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(9)</label>
</disp-formula>
</p>
<p>Then, letting <italic>f</italic>
<sub>
<italic>j</italic>
</sub>(<italic>x</italic>) be an assumed probability distribution/mass function for <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> with support <inline-formula id="inf24">
<mml:math id="m33">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> [e.g., if <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> is a vector of minor allele counts with MAF <italic>p</italic>, letting <inline-formula id="inf25">
<mml:math id="m34">
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mo stretchy="false">(</mml:mo>
<mml:mtable>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
<mml:mo stretchy="false">)</mml:mo>
<mml:msup>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
<mml:msup>
<mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>p</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
</mml:msup>
</mml:math>
</inline-formula> and <inline-formula id="inf26">
<mml:math id="m35">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">{</mml:mo>
<mml:mrow>
<mml:mn>0,1,2</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">}</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula>) we approximate the sample covariance of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>2</sub>:<disp-formula id="e10">
<mml:math id="m36">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2248;</mml:mo>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>,</mml:mo>
</mml:math>
<label>(10)</label>
</disp-formula>swapping the sums for integrals across the support when appropriate.</p>
<p>We calculate the sample mean of <bold>
<italic>w</italic>
</bold>
<sub>2</sub> as<disp-formula id="e11">
<mml:math id="m37">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>/</mml:mo>
<mml:mi>n</mml:mi>
</mml:math>
<label>(11)</label>
</disp-formula>
</p>
<p>To approximate the variance, we first approximate the conditional variances of <bold>
<italic>w</italic>
</bold>
<sub>2</sub> at all levels of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>:<disp-formula id="e12">
<mml:math id="m38">
<mml:mtable class="aligned">
<mml:mtr>
<mml:mtd columnalign="right">
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
</mml:mtd>
<mml:mtd columnalign="left">
<mml:mspace width="0.3333em"/>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:mtd>
</mml:mtr>
<mml:mtr>
<mml:mtd columnalign="right"/>
<mml:mtd columnalign="left">
<mml:mspace width="0.3333em"/>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mtd>
</mml:mtr>
</mml:mtable>
</mml:math>
<label>(12)</label>
</disp-formula>And then approximate the sample variance as:<disp-formula id="e13">
<mml:math id="m39">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x2248;</mml:mo>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:munder>
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>x</mml:mi>
<mml:mo>&#x2208;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:munder>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:mfenced>
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
<mml:mi>n</mml:mi>
<mml:msub>
<mml:mrow>
<mml:mi>f</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mfenced open="(" close=")">
<mml:mrow>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mfenced open="" close=")">
</mml:mfenced>
</mml:mrow>
</mml:mfenced>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msup>
</mml:mrow>
</mml:mfenced>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(13)</label>
</disp-formula>once again swapping the sum for an integral across <inline-formula id="inf27">
<mml:math id="m40">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="script">S</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> when appropriate. This approach leads to a different variance estimate for each predictor <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>. We treat the median of these estimates across each <italic>j</italic> as the estimated variance.</p>
<p>Hence, taking the means, variances, and pairwise covariances of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>, <bold>
<italic>y</italic>
</bold>
<sub>1</sub>, and <bold>
<italic>y</italic>
</bold>
<sub>2</sub> and a distributional assumption about <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>, we approximate the covariance of one variable (<bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub>) with the product of the other two (<bold>
<italic>w</italic>
</bold>
<sub>2</sub> &#x3d; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>2</sub>) as well as the product&#x2019;s mean and variance.</p>
<p>Repeating this algorithm for each predictor <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and following the linear regression equations presented in <xref ref-type="sec" rid="s2-2">Section 2.2</xref> allows for calculation of covariate adjusted slope coefficients for the multiple regression model <inline-formula id="inf28">
<mml:math id="m41">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:math>
</inline-formula>&#x3f5; as well as the standard errors of these slope estimates.</p>
</sec>
<sec id="s2-3-2">
<title>2.3.2 Covariance Estimation With the Product of 3 or More Phenotypes</title>
<p>Regression models for larger products of phenotypes can also be approximated by applying the established method recursively: first estimating the covariance of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>2</sub>, then leveraging the covariance of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>2</sub> and <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and <bold>
<italic>y</italic>
</bold>
<sub>3</sub> to estimate the covariance of <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>3</sub>, and so forth. This recursion procedure is described in more detail in the appendix and software to carry it out is discussed in <xref ref-type="sec" rid="s2-8">Section&#x20;2.8</xref>.</p>
<p>Let <bold>
<italic>w</italic>
</bold>
<sub>
<italic>l</italic>
</sub> &#x3d; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x22ef;<bold>
<italic>y</italic>
</bold>
<sub>
<italic>l</italic>
</sub> &#x3d; <bold>
<italic>w</italic>
</bold>
<sub>
<italic>l</italic>&#x2212;1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>
<italic>l</italic>
</sub>. In order to estimate <inline-formula id="inf29">
<mml:math id="m42">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> through our established method, we use <inline-formula id="inf30">
<mml:math id="m43">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf31">
<mml:math id="m44">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf32">
<mml:math id="m45">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf33">
<mml:math id="m46">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf34">
<mml:math id="m47">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf35">
<mml:math id="m48">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf36">
<mml:math id="m49">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf37">
<mml:math id="m50">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, and <inline-formula id="inf38">
<mml:math id="m51">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> as inputs to the method described in <xref ref-type="sec" rid="s2-3-1">Section 2.3.1</xref>. That is, replacing <bold>
<italic>y</italic>
</bold>
<sub>1</sub> with <bold>
<italic>w</italic>
</bold>
<sub>
<italic>l</italic>&#x2212;1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub> with <bold>
<italic>y</italic>
</bold>
<sub>
<italic>l</italic>
</sub>. While <inline-formula id="inf39">
<mml:math id="m52">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf40">
<mml:math id="m53">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf41">
<mml:math id="m54">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, <inline-formula id="inf42">
<mml:math id="m55">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and <inline-formula id="inf43">
<mml:math id="m56">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> are assumed to be known, we must estimate <inline-formula id="inf44">
<mml:math id="m57">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf45">
<mml:math id="m58">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>.</p>
<p>Continuation of the recursive process starting at <bold>
<italic>w</italic>
</bold>
<sub>
<italic>l</italic>&#x2212;1</sub> and working down to <bold>
<italic>w</italic>
</bold>
<sub>2</sub> will yield an estimate for <inline-formula id="inf46">
<mml:math id="m59">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, or eventually the base case of <inline-formula id="inf47">
<mml:math id="m60">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>.</p>
<p>To approximate <inline-formula id="inf48">
<mml:math id="m61">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, we re-express the term as <inline-formula id="inf49">
<mml:math id="m62">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and approximate the covariance of <bold>
<italic>y</italic>
</bold>
<sub>
<italic>l</italic>
</sub> and <bold>
<italic>w</italic>
</bold>
<sub>
<italic>l</italic>&#x2212;2</sub>
<bold>
<italic>y</italic>
</bold>
<sub>
<italic>l</italic>&#x2212;1</sub> through the method described in <xref ref-type="sec" rid="s2-3-1">Section&#x20;2.3.1</xref>.</p>
<p>A diagram of the start of this recursion is displayed in <xref ref-type="sec" rid="s10">Supplementary Figure&#x20;S1</xref>.</p>
<p>This recursive estimation is impacted by the order in which the phenotypes are multiplied. So, any set of more than two phenotypes will render <italic>m</italic>!/2 possible ways to estimate the regression model through this method. Hence, we approximate the covariances and means using all permutations of <bold>
<italic>y</italic>
</bold>
<sub>1</sub>, &#x2026;, <bold>
<italic>y</italic>
</bold>
<sub>
<italic>m</italic>
</sub> unique up to the order of the first two terms as the order of our phenotypes, and take the median estimate of each term across all permutations as its predicted&#x20;value.</p>
</sec>
</sec>
<sec id="s2-4">
<title>2.4 Binary Phenotypes</title>
<p>Binary phenotypes present both new challenges to estimation and the opportunity to express logical combinations of phenotypes through products.</p>
<sec id="s2-4-1">
<title>2.4.1 Changes to Estimations</title>
<p>Our proposed method leverages several assumptions of the standard linear regression model in its approximations (namely those of homoscedasticity and linearity). As binary phenotypes notoriously violate both of these assumptions, we slightly adjust our approximation to better reflect this situation.</p>
<p>The covariance of two binary phenotypes is estimated using the same general framework as developed in <xref ref-type="sec" rid="s2-3-1">Section 2.3.1</xref>. The only changes are to mean the variance estimates. All estimated means (conditional and otherwise) are initially estimated as proposed in <xref ref-type="sec" rid="s2-3-1">Section 2.3.1</xref> and then restricted to the open interval (0, 1) (e.g., letting <italic>g</italic>(<italic>y</italic>
<sub>
<italic>k</italic>
</sub>&#x7c;<italic>x</italic>) &#x3d; <italic>a</italic>
<sub>
<italic>kj</italic>
</sub> &#x2b; <italic>b</italic>
<sub>
<italic>kj</italic>
</sub>
<italic>x</italic> if 0 &#x3c; <italic>a</italic>
<sub>
<italic>kj</italic>
</sub> &#x2b; <italic>b</italic>
<sub>
<italic>kj</italic>
</sub>
<italic>x</italic>&#x20;&#x3c; 1, &#x3f5; if <italic>a</italic>
<sub>
<italic>kj</italic>
</sub> &#x2b; <italic>b</italic>
<sub>
<italic>kj</italic>
</sub>
<italic>x</italic>&#x20;&#x2264; 0, and 1&#x20;&#x2212; &#x3f5; if 1 &#x2264; <italic>a</italic>
<sub>
<italic>kj</italic>
</sub> &#x2b; <italic>b</italic>
<sub>
<italic>kj</italic>
</sub>&#x2009;<italic>x</italic> for some small &#x3f5; &#x3e;&#x20;0).</p>
<p>Instead of estimating a phenotype&#x2019;s conditional variance from a linear model&#x2019;s residual variance, we estimate it as<disp-formula id="e14">
<mml:math id="m63">
<mml:mi>h</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:mi>g</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo stretchy="false">&#x7c;</mml:mo>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(14)</label>
</disp-formula>
</p>
<p>Further, we calculate the product&#x2019;s sample variance as<disp-formula id="e15">
<mml:math id="m64">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>/</mml:mo>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:mi>n</mml:mi>
<mml:mo>&#x2212;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
<label>(15)</label>
</disp-formula>
</p>
</sec>
<sec id="s2-4-2">
<title>2.4.2 Products as Logical Combinations</title>
<p>Binary phenotypes are of particular importance because their products can be interpreted as logical combinations.</p>
<p>We can represent the logical conjunction <bold>
<italic>y</italic>
</bold>
<sub>1</sub> &#x2227;<bold> <italic>y</italic>
</bold>
<sub>2</sub> (read as &#x201c;<bold>
<italic>y</italic>
</bold>
<sub>1</sub> and <bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x201d;) as the product <bold>
<italic>y</italic>
</bold>
<sub>1</sub>
<bold>
<italic>y</italic>
</bold>
<sub>2</sub>. Likewise, we express the logical disjunction <bold>
<italic>y</italic>
</bold>
<sub>1</sub> &#x2228;<bold> <italic>y</italic>
</bold>
<sub>2</sub> (&#x201c;<bold>
<italic>y</italic>
</bold>
<sub>1</sub> or <bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x201d;) as <bold>1</bold>
<sub>
<italic>n</italic>
</sub> &#x2212; ((<bold>1</bold>
<sub>
<italic>n</italic>
</sub> &#x2212; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>)(<bold>1</bold>
<sub>
<italic>n</italic>
</sub> &#x2212;&#x20;<bold>
<italic>y</italic>
</bold>
<sub>2</sub>)).</p>
<p>By framing both disjunctions and conjunctions in terms of phenotype multiplication, we can apply our established methods to approximate the covariances of these combinations with predictors and ultimately estimate linear models for these logical combinations.</p>
<p>While the case of the conjunction is a trivial application of the above methods of multiplying phenotypes, we will briefly describe how to model the disjunction. To do so, we consider the modified phenotypes <inline-formula id="inf50">
<mml:math id="m65">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> and <inline-formula id="inf51">
<mml:math id="m66">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mn mathvariant="bold">1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>n</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> (these represent the statements &#x201c;not <bold>
<italic>y</italic>
</bold>
<sub>1</sub>&#x201d; and &#x201c;not <bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x201d;). This gives us <bold>
<italic>y</italic>
</bold>
<sub>1</sub>&#x2009;&#x2228;&#x2009;<bold>
<italic>y</italic>
</bold>
<sub>2</sub> &#x3d; <bold>1</bold>
<sub>
<italic>n</italic>
</sub> &#x2212; <bold>
<italic>y</italic>
</bold>
<sub>1</sub>&#x2032;<bold>
<italic>y</italic>
</bold>
<sub>2</sub>&#x2032;. Then, <inline-formula id="inf52">
<mml:math id="m67">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf53">
<mml:math id="m68">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, and <inline-formula id="inf54">
<mml:math id="m69">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>l</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. If we set <inline-formula id="inf55">
<mml:math id="m70">
<mml:msubsup>
<mml:mrow>
<mml:mi mathvariant="bold-italic">w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, our method allow us to estimate <inline-formula id="inf56">
<mml:math id="m71">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula> for each <bold>
<italic>x</italic>
</bold>
<sub>
<italic>j</italic>
</sub> as well as <inline-formula id="inf57">
<mml:math id="m72">
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula> and <inline-formula id="inf58">
<mml:math id="m73">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>. Leveraging these estimates, <inline-formula id="inf59">
<mml:math id="m74">
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mo>&#x2212;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>, <inline-formula id="inf60">
<mml:math id="m75">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
<mml:mo>&#x2212;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, and <inline-formula id="inf61">
<mml:math id="m76">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msubsup>
<mml:mrow>
<mml:mi>w</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mo>&#x2032;</mml:mo>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, where <bold>
<italic>w</italic>
</bold>
<sub>2</sub> is equivalent to the disjunction <bold>
<italic>y</italic>
</bold>
<sub>1</sub>&#x2009;&#x2228;&#x2009;<bold>
<italic>y</italic>
</bold>
<sub>2</sub>. Using these terms as inputs for the framework presented in <xref ref-type="sec" rid="s2-2">Section 2.2</xref> allow for coefficient and standard error estimation for the linear model <inline-formula id="inf62">
<mml:math id="m77">
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2228;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi mathvariant="bold-italic">y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mi mathvariant="bold-italic">X</mml:mi>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi mathvariant="bold-italic">&#x3b2;</mml:mi>
</mml:mrow>
<mml:mo stretchy="false">&#x302;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mo>&#x2b;</mml:mo>
</mml:math>
</inline-formula>&#x3f5;.</p>
</sec>
</sec>
<sec id="s2-5">
<title>2.5 Simulation Studies</title>
<sec id="s2-5-1">
<title>2.5.1 Simulation 1: Type I Error Maintenance</title>
<p>To verify that our linear model with PCSS approach appropriately maintained the Type I error rate at a variety of <italic>&#x3b1;</italic> thresholds, we carried out a simulation under the null hypothesis that the predictor variant has no linear association with any of the phenotypes of interest. This null hypothesis represents a reasonable subset of the exact null hypothesis which is that the <italic>product</italic> of phenotypes has no linear relationship with the predictor. We carried out this simulation with varying sample size, MAF, phenotype means, phenotype correlations, and for continuous phenotypes, phenotype variances, for products of two binary phenotypes, two continuous phenotypes, and three continuous phenotypes. When simulating continuous phenotypes we assumed that <inline-formula id="inf63">
<mml:math id="m78">
<mml:msub>
<mml:mrow>
<mml:mi>Y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x223c;</mml:mo>
<mml:mi>N</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3bc;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
</mml:math>
</inline-formula> with Cor(<italic>Y</italic>
<sub>
<italic>ik</italic>
</sub>, <italic>Y</italic>
<sub>
<italic>il</italic>
</sub>) <italic>&#x3d; &#x3c1;</italic>
<sub>
<italic>kl</italic>
</sub> while when simulating binary phenotypes we let <italic>Y</italic>
<sub>
<italic>ik</italic>
</sub> &#x223c; Bernoulli<italic>(&#x3bc;</italic>
<sub>
<italic>k</italic>
</sub>
<italic>)</italic> (again with Cor(<italic>Y</italic>
<sub>
<italic>ik</italic>
</sub>
<italic>, Y</italic>
<sub>
<italic>il</italic>
</sub>
<italic>) &#x3d; &#x3c1;</italic>
<sub>
<italic>kl</italic>
</sub>). Simulation parameters (e.g., <italic>n</italic>, <italic>&#x3bc;</italic>
<sub>1</sub>, <italic>&#x3c1;</italic>
<sub>12</sub>) were randomly sampled from various distributions (full details are available in <xref ref-type="sec" rid="s10">Supplementary Table S1</xref>). We carried out 10<sup>8</sup> simulations for each collection of continuous phenotypes and 10<sup>7</sup> simulations for the case of binary phenotypes.</p>
</sec>
<sec id="s2-5-2">
<title>2.5.2 Simulation 2: Comparisons to IPD Models</title>
<p>To evaluate our method&#x2019;s ability to replicate the results of covariate adjusted linear models fit to IPD, we carried out three 2<sup>
<italic>k</italic>
</sup> factorial simulations&#x2014;one for the product of two binary phenotypes, one for the product of two positive continuous phenotypes, and one for the product of three positive continuous phenotypes. We carried out 1,000 simulations at each possible combination of parameters. In each simulation, we modeled the phenotype product as a function of a SNP and binary covariate. For the simulations with only two phenotypes, we also included a continuous covariate in our models.</p>
<p>In all simulations, we simulated <italic>n</italic> subjects&#x2019; SNP minor allele counts <bold>
<italic>x</italic>
</bold>
<sub>1</sub> at HWE with varying MAF. We simulated a binary covariate <bold>
<italic>x</italic>
</bold>
<sub>2</sub> &#x223c; Bernoulli(logit<sup>&#x2212;1</sup>&#x3b1;<sub>2</sub>
<bold>
<italic>x</italic>
</bold>
<sub>1</sub>). When generating sets of two phenotypes we also generated a continuous covariate <bold>
<italic>x</italic>
</bold>
<sub>3</sub> from a linear model with <bold>
<italic>x</italic>
</bold>
<sub>1</sub> such that <inline-formula id="inf64">
<mml:math id="m79">
<mml:msub>
<mml:mrow>
<mml:mover accent="true">
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mo>&#x304;</mml:mo>
</mml:mover>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>0</mml:mn>
</mml:math>
</inline-formula>, <inline-formula id="inf65">
<mml:math id="m80">
<mml:msubsup>
<mml:mrow>
<mml:mi>s</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:math>
</inline-formula>, and <inline-formula id="inf66">
<mml:math id="m81">
<mml:msub>
<mml:mrow>
<mml:mi>r</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>1</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>,</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:mrow>
</mml:msub>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b1;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>3</mml:mn>
</mml:mrow>
</mml:msub>
</mml:math>
</inline-formula>. This resulted in a SNP with two covariates (<italic>p</italic>&#x20;&#x3d; 3) in our two phenotype simulations, and a SNP with one covariate (<italic>p</italic>&#x20;&#x3d; 2) in our three phenotype simulation.</p>
<p>We generated individual phenotype measures through the model<disp-formula id="equ1">
<mml:math id="m82">
<mml:mi>u</mml:mi>
<mml:mrow>
<mml:mo stretchy="false">(</mml:mo>
<mml:mrow>
<mml:msub>
<mml:mrow>
<mml:mi>y</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:mrow>
<mml:mo stretchy="false">)</mml:mo>
</mml:mrow>
<mml:mo>&#x3d;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mn>0</mml:mn>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:munderover accentunder="false" accent="false">
<mml:mrow>
<mml:mo>&#x2211;</mml:mo>
</mml:mrow>
<mml:mrow>
<mml:mi>j</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn>1</mml:mn>
</mml:mrow>
<mml:mrow>
<mml:mi>p</mml:mi>
</mml:mrow>
</mml:munderover>
<mml:msub>
<mml:mrow>
<mml:mi>x</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3b2;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
<mml:mi>j</mml:mi>
</mml:mrow>
</mml:msub>
<mml:mo>&#x2b;</mml:mo>
<mml:msub>
<mml:mrow>
<mml:mi>&#x3f5;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>i</mml:mi>
<mml:mi>k</mml:mi>
</mml:mrow>
</mml:msub>
</mml:math>
</disp-formula>where <italic>u</italic>(<italic>y</italic>
<sub>
<italic>ik</italic>
</sub>) &#x3d; <italic>y</italic>
<sub>
<italic>ik</italic>
</sub> for continuous phenotypes, <italic>u</italic>(<italic>y</italic>
<sub>
<italic>ik</italic>
</sub>) &#x3d; logit(Pr(<italic>Y</italic>
<sub>
<italic>ik</italic>
</sub> &#x3d; 1)) for binary phenotypes, and &#x3f5;<sub>i</sub>
<sup>&#x2032;</sup> follows a multivariate normal distribution with <inline-formula id="inf68">
<mml:math id="m84">
<mml:mi mathvariant="bold-italic">&#x3bc;</mml:mi>
<mml:mo>&#x3d;</mml:mo>
<mml:mn mathvariant="bold-italic">0</mml:mn>
</mml:math>
</inline-formula> and <bold>&#x3a3;</bold>
<sub>(<italic>i</italic>,<italic>j</italic>)</sub> &#x3d; <italic>&#x3c3;</italic>
<sub>
<italic>i</italic>
</sub>
<italic>&#x3c3;</italic>
<sub>
<italic>j</italic>
</sub>
<italic>&#x3c1;</italic>
<sub>
<italic>ij</italic>
</sub>. Parameter values were selected such that, under optimal settings, empirical power was roughly 80&#x2013;90<italic>%</italic> at a significance threshold of 10<sup>&#x2212;8</sup>. Full details of simulation parameters are available in <xref ref-type="sec" rid="s10">Supplementary Table&#x20;S2</xref>.</p>
<p>In each simulation, we estimated coefficients, standard errors, <italic>t</italic> statistics, and two-sided <italic>p</italic>-values for the null hypothesis that there was no relationship between the product of phenotypes and the SNP (<bold>
<italic>x</italic>
</bold>
<sub>1</sub>) after adjusting for covariates both using IPD and using&#x20;PCSS.</p>
<p>Additionally, when simulating two binary phenotypes we fit covariate-adjusted logistic regression models for the logged odds that <italic>y</italic>
<sub>1<italic>i</italic>
</sub>
<italic>y</italic>
<sub>2<italic>i</italic>
</sub> &#x3d; 1 using IPD and returned the relevant two-sided <italic>p</italic>-value to compare the results of the linear model fit using PCSS to the correctly specified logistic&#x20;model.</p>
</sec>
</sec>
<sec id="s2-6">
<title>2.6 Real Data Application</title>
<p>Fatty acids are of broad importance for a wide range of cardiometabolic traits (<xref ref-type="bibr" rid="B12">Imamura et&#x20;al., 2020</xref>) with ratios of fatty acids are often used as a proxy for conversion efficiency. Previous genome wide association studies have explored the genetic architecture of fatty acids and their ratios (<xref ref-type="bibr" rid="B15">Kalsbeek et&#x20;al., 2018</xref>; <xref ref-type="bibr" rid="B17">Lemaitre et&#x20;al., 2011</xref>; <xref ref-type="bibr" rid="B24">Tintle et&#x20;al., 2015</xref>, <xref ref-type="bibr" rid="B23">2020</xref>).</p>
<p>We modeled 12 fatty acid ratios using both IPD and PCSS using data from the Framingham Heart Study&#x2019;s Generation-3 and Offspring cohorts downloaded from dbGaP (<xref ref-type="bibr" rid="B19">Mailman et&#x20;al., 2007</xref>). The specific ratios can be found in the first column of <xref ref-type="table" rid="T3">Table&#x20;3</xref>. <xref ref-type="sec" rid="s10">Supplementary Table S3</xref> lists all fatty acids used in at least one of the ratios alongside their abbreviations.</p>
<p>Quality control measures included setting Mendelian inconsistencies as missing and excluding SNPs with HWE <italic>p</italic>&#x20;&#x3c; 0.00001, MAF &#x3c;0.05, or missing values for over 10<italic>%</italic> of subjects. We excluded individuals missing over 10<italic>%</italic> of their genetic data after initial quality control and then took a subset of unrelated participants. After quality control we were left with 362,330 SNPs over 1,455 individuals (657 from the Offspring cohort and 888 from the Generation-3 cohort).</p>
<p>In addition to the standard PCSS described in <xref ref-type="sec" rid="s2-1">Section 2.1</xref>, we assumed access to pre-computed means and variances of the reciprocal of each fatty acid as well as the correlation between any fatty acid reciprocal and any other fatty acid, covariate, or SNP to model these ratios using&#x20;PCSS.</p>
<p>We analyzed each fatty acid ratio through the linear model: Ratio &#x223c; SNP &#x2b; age &#x2b; sex for each SNP in our sample using both IPD and PCSS and tested each SNP for statistical significance with the Bonferroni adjusted threshold <italic>&#x3b1;</italic>&#x20;&#x3d;&#x20;1.37 &#xd7; 10<sup>&#x2212;7</sup>.</p>
</sec>
<sec id="s2-7">
<title>2.7 Statistical Analysis</title>
<sec id="s2-7-1">
<title>2.7.1 Simulation 1</title>
<p>To analyze the results of our Type I Error simulations we calculated the empirical Type I Error rate when approximating linear models using PCSS at each specified significance threshold.</p>
</sec>
<sec id="s2-7-2">
<title>2.7.2 Simulation 2</title>
<p>For each of the three 2<sup>
<italic>k</italic>
</sup> factorial simulations, we assessed the PCSS model&#x2019;s accuracy relative to its IPD counterpart.</p>
<p>We calculated the bias and mean squared error when estimating the SNP&#x2019;s slope coefficient, standard error, and absolute value test statistic. We also modeled errors estimating the slope coefficient, standard error, and test statistic through multiple linear regression models with logical indicators for each of the <italic>k</italic> parameter settings as predictors, testing at the Bonferroni adjusted significance threshold of 0.05/<italic>k</italic> to determine which simulation parameters affected our method&#x2019;s accuracy.</p>
<p>We compared test decisions regarding the significance of the SNP when modeling the phenotype product after adjusting for covariates at significance thresholds 10<sup>&#x2212;1</sup>, 10<sup>&#x2212;2</sup>, &#x2026;, 10<sup>&#x2212;8</sup>. When analyzing binary phenotypes we also compared test decisions between the linear model fit using PCSS and the logistic regression model fit on IPD to demonstrate the robustness of linear models to model binary outcomes.</p>
</sec>
<sec id="s2-7-3">
<title>2.7.3 Real Data Application</title>
<p>We measured our overall bias and mean squared error in slope, standard error, and absolute value test statistic estimates for each outcome modeled. We recorded test decisions for both the IPD and PCSS models and whether the two results agreed or disagreed regarding the statistical significance of each SNP. When one approach found a SNP to be significant and the other did not, we noted if the non-significant result was &#x201c;borderline&#x201d; significant (<italic>&#x3b1;</italic>&#x20;&#x2264; <italic>p</italic>&#x20;&#x3c; 10<italic>&#x3b1;</italic>).</p>
</sec>
</sec>
<sec id="s2-8">
<title>2.8 Software</title>
<p>Software to perform these model approximations as well as those developed in <xref ref-type="bibr" rid="B26">Wolf et&#x20;al. (2020)</xref> is available on CRAN in the R package pcsstools<xref ref-type="fn" rid="fn5">
<sup>5</sup>
</xref>.</p>
</sec>
</sec>
<sec sec-type="results" id="s3">
<title>3 Results</title>
<sec id="s3-1">
<title>3.1 Simulation Studies</title>
<sec id="s3-1-1">
<title>3.1.1 Simulation 1</title>
<p>Empirical Type I error rates when using PCSS are displayed in <xref ref-type="table" rid="T1">Table&#x20;1</xref>. In all simulations, the approach&#x2019;s empirical Type I error rate was below the tested significance threshold.</p>
<table-wrap id="T1" position="float">
<label>TABLE 1</label>
<caption>
<p>Simulation studies of Type I error estimates when testing the linear association between a single SNP and a product of phenotypes using pre-computed summary statistics at significance thresholds: <italic>&#x3b1;</italic> &#x3d; 0.05, 0.001, 10<sup>&#x2212;5</sup>, and 10<sup>&#x2212;6</sup>. Each entry represents the proportion of <italic>p</italic>-values smaller than <italic>&#x3b1;</italic> when modeling the linear relation between a SNP and a product of phenotypes.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left"/>
<th colspan="4" align="center">Nominal <italic>&#x3b1;</italic>
</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">Phenotypes</td>
<td align="center">5.0E-02</td>
<td align="center">1.0E-03</td>
<td align="center">1.0E-05</td>
<td align="center">1.0E-06</td>
</tr>
<tr>
<td align="left">2 Continuous</td>
<td align="center">3.88E-02</td>
<td align="center">6.65E-04</td>
<td align="center">5.56E-06</td>
<td align="center">4.40E-07</td>
</tr>
<tr>
<td align="left">2 Binary</td>
<td align="center">2.39E-02</td>
<td align="center">2.06E-04</td>
<td align="center">8.00E-07</td>
<td align="center">1.00E-07</td>
</tr>
<tr>
<td align="left">3 Continuous</td>
<td align="center">2.70E-02</td>
<td align="center">3.81E-04</td>
<td align="center">2.91E-06</td>
<td align="center">3.40E-07</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3-1-2">
<title>3.1.2 Simulation 2</title>
<p>The PCSS method&#x2019;s errors when approximating slope coefficients, their standard errors, and test statistics are available in <xref ref-type="table" rid="T2">Table&#x20;2</xref>. When aggregated over all simulation settings we observe slight positive bias both when estimating the slope and the absolute value of the <italic>t</italic> test statistic for each collection of phenotypes. The magnitude of the mean test statistic error is comparable across all three simulations. <xref ref-type="fig" rid="F2">Figure&#x20;2</xref> displays our PCSS method&#x2019;s approximated slope coefficients compared to slope coefficients calculated using IPD for the SNP while modeling the phenotype product and adjusting for covariates. Similar graphical comparisons of standard error and test statistic estimates are available in <xref ref-type="sec" rid="s10">Supplementary Figures S2,&#x20;S3</xref>.</p>
<table-wrap id="T2" position="float">
<label>TABLE 2</label>
<caption>
<p>Simulation study approximating a linear model for a product of phenotypes using summary statistics. Summaries of errors when approximating slopes, slope standard errors, and the absolute value of the <italic>t</italic>-statistics for a SNP while adjusting for covariates when using pre-computed summary statistics (PCSS) compared to values obtained when calculating these statistics using individual participant data (IPD).</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Phenotypes</th>
<th colspan="3" align="center">
<italic>&#x3b2;</italic>
</th>
<th colspan="2" align="center">SE(<italic>&#x3b2;</italic>)</th>
<th colspan="2" align="center">&#x7c;<italic>t</italic>&#x7c;</th>
</tr>
<tr>
<th align="center">IPD Mean</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">2 Continuous</td>
<td align="center">6.09E-03</td>
<td align="center">3.72E-04</td>
<td align="center">1.43E-05</td>
<td align="center">2.39E-06</td>
<td align="center">6.58E-11</td>
<td align="center">2.33E-03</td>
<td align="center">4.13E-01</td>
</tr>
<tr>
<td align="left">2 Binary</td>
<td align="center">4.13E-01</td>
<td align="center">2.56E-02</td>
<td align="center">6.74E-03</td>
<td align="center">6.58E-05</td>
<td align="center">5.50E-07</td>
<td align="center">1.65E-01</td>
<td align="center">2.16E-01</td>
</tr>
<tr>
<td align="left">3 Continuous</td>
<td align="center">4.82E&#x2b;00</td>
<td align="center">3.33E-02</td>
<td align="center">3.79E-01</td>
<td align="center">-3.63E-02</td>
<td align="center">5.82E-04</td>
<td align="center">5.71E-02</td>
<td align="center">1.06E-01</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="F2" position="float">
<label>FIGURE 2</label>
<caption>
<p>Comparison of slope coefficients from a simulation study approximating a covariate adjusted linear model for a product of phenotypes using pre-computed summary statistics (PCSS) and individual participant data (IPD). <bold>(A)</bold> Modeling the product of two continuous phenotypes while adjusting for a binary and a continuous covariate. <bold>(B)</bold> Modeling the product of two binary phenotypes while adjusting for a binary and a continuous covariate. <bold>(C)</bold> Modeling the product of three continuous phenotypes while adjusting for a binary covariate.</p>
</caption>
<graphic xlink:href="fgene-12-745901-g002.tif"/>
</fig>
<p>When modeling estimation errors for two continuous phenotypes through a linear regression model with indicator variables for all of the simulation settings (<italic>k</italic>&#x20;&#x3d; 12, <italic>n</italic>&#x20;&#x3d; 2<sup>
<italic>k</italic>
</sup> &#xd7; 10<sup>3</sup>), our model for the slope error found all settings except the residual phenotype variances, <inline-formula id="inf69">
<mml:math id="m85">
<mml:msubsup>
<mml:mrow>
<mml:mi>&#x3c3;</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mi>k</mml:mi>
</mml:mrow>
<mml:mrow>
<mml:mn>2</mml:mn>
</mml:mrow>
</mml:msubsup>
</mml:math>
</inline-formula>, to be significantly associated with the PCSS model&#x2019;s slope estimate&#x2019;s error at the adjusted significance threshold 0.05/<italic>k</italic>. All settings had significant associations with our error when estimating the standard error of the slope coefficient, or the test statistic. In the case of two binary phenotypes (<italic>k</italic>&#x20;&#x3d; 14, <italic>n</italic>&#x20;&#x3d; 2<sup>
<italic>k</italic>
</sup> &#xd7; 10<sup>3</sup>), we found all settings to have significant associations with the error in slope, standard error, and test statistic estimates. For three continuous phenotypes (<italic>k</italic>&#x20;&#x3d; 13, <italic>n</italic>&#x20;&#x3d; 2<sup>
<italic>k</italic>
</sup> &#xd7; 10<sup>3</sup>), we also found all settings to have significant associations with the error when predicting the slope coefficient, its standard error, and its test statistic.</p>
<p>
<xref ref-type="fig" rid="F3">Figure&#x20;3</xref> shows comparisons of estimated and calculated <italic>p</italic>-values for a two-sided <italic>t</italic> test under the null hypothesis that the SNP had no linear association with the phenotype product after adjusting for covariates. (<xref ref-type="fig" rid="F3">Figure&#x20;3</xref> only includes <italic>p</italic>-values greater than 10<sup>&#x2212;15</sup> for the sake of visual clarity; <xref ref-type="sec" rid="s10">Supplementary Figure S4</xref> repeats this visualization without any restrictions.) <xref ref-type="fig" rid="F4">Figure&#x20;4</xref> shows various error rates rate between the IPD and PCSS models&#x2019; test decisions based on these <italic>p</italic>-values at differing significance thresholds. We see that all PCSS models overall disagreement rates to their IPD companions decrease as the significance threshold becomes more stringent. Likewise, when the IPD model rejected the null hypothesis, the PCSS model rarely failed to reject with error rates at most 13% which again decreased as the significance threshold decreased. When the IPD model failed to reject the null hypothesis, the PCSS approach&#x2019;s conditional error rate varied by the model&#x2019;s response. When modeling the product of two continuous or binary phenotypes, the error rate stayed relatively constant across all thresholds at around 3 and 15%, respectively. But, when modeling the product of three continuous phenotypes, the error rate increased as the significance threshold became more strict. Lastly, when compared to the test decisions of a covariate adjusted logistic regression model, our PCSS approximation of the related linear model tends to reach the same conclusions, with a moderate conservative tendency, especially at more strict significance thresholds.</p>
<fig id="F3" position="float">
<label>FIGURE 3</label>
<caption>
<p>Comparison of <italic>p</italic>-values from a simulation study approximating a covariate adjusted linear model for a product of phenotypes using pre-computed summary statistics (PCSS) and individual participant data (IPD). Two-sided <italic>p</italic>-values were computed for the null hypothesis that the SNP had no linear effect on the phenotype product while adjusting for covariates. All plots are restricted to the range (0, 15); unrestricted plots are available in the supplementary materials. <bold>(A)</bold> Modeling the product of two continuous phenotypes while adjusting for a binary and a continuous covariate. <bold>(B)</bold> Modeling the product of two binary phenotypes while adjusting for a binary and a continuous covariate. <bold>(C)</bold> Modeling the product of three continuous phenotypes while adjusting for a binary covariate.</p>
</caption>
<graphic xlink:href="fgene-12-745901-g003.tif"/>
</fig>
<fig id="F4" position="float">
<label>FIGURE 4</label>
<caption>
<p>Simulation studies&#x2019; test decision disagreement rates evaluating the significance of a SNP in a linear model for a product of phenotypes while adjusting for covariates using individualized participant data (IPD) and pre-computed summary statistics (PCSS) at various significance thresholds (<italic>&#x3b1;</italic>). Comparisons were also made between a logistic regression model fit using IPD on the product of two binary phenotypes and the PCSS model approximating the linear relationship. <bold>(A)</bold> Percentage of times the PCSS and IPD models&#x2019; test decisions disagreed across all simulations. <bold>(B)</bold> Percentage of times the PCSS model rejected the null hypothesis given that the IPD model failed to reject the null hypothesis. <bold>(C)</bold> Percentage of times the PCSS model failed to reject the null hypothesis given that the IPD model rejected the null hypothesis.</p>
</caption>
<graphic xlink:href="fgene-12-745901-g004.tif"/>
</fig>
</sec>
</sec>
<sec id="s3-2">
<title>3.2 Real Data Application</title>
<p>The bias and mean squared error of the PCSS model&#x2019;s approximation to the IPD model&#x2019;s slope, standard error, and absolute value test statistic for each fatty acid ratio are displayed in <xref ref-type="table" rid="T3">Table&#x20;3</xref>. Our mean slope error was &#x2212;2.93&#x20;&#xd7; 10<sup>&#x2212;3</sup> (Mean Squared Error 0.114) while the mean slope estimate when using IPD was &#x2212;1.3&#x20;&#xd7; 10<sup>&#x2212;3</sup>, demonstrating a slight bias towards zero. However, the standard error estimates under the PCSS model were on average 4.80&#x20;&#xd7; 10<sup>&#x2212;2</sup> lower than their respective estimates under the IPD model and absolute value tests statistics tended to be 1.79&#x20;&#xd7; 10<sup>&#x2212;2</sup> higher than their IPD counterparts indicating an overall minor anti-conservative&#x20;bias.</p>
<table-wrap id="T3" position="float">
<label>TABLE 3</label>
<caption>
<p>Summary of errors when approximating the linear model: FA Ratio &#x223c; snp &#x2b; age &#x2b; sex using pre-computed summary statistics (PCSS) compared to values obtained when calculating these statistics using individual participant data (IPD). Each fatty acid ratio was modeled across 362,330 SNPs from 1,455 subjects in the Framing Heart Study&#x2019;s Offspring and Generation-3 cohorts.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th rowspan="2" align="left">Ratio</th>
<th colspan="3" align="center">
<italic>&#x3b2;</italic>
</th>
<th colspan="2" align="center">SE(<italic>&#x3b2;</italic>)</th>
<th colspan="2" align="center">&#x7c;<italic>t</italic>&#x7c;</th>
</tr>
<tr>
<th align="center">IPD Mean</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
<th align="center">Bias</th>
<th align="center">MSE</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PA:POA</td>
<td align="center">5.58E-02</td>
<td align="center">1.49E-02</td>
<td align="center">4.18E-02</td>
<td align="center">&#x2212;4.80E-02</td>
<td align="center">2.08E-04</td>
<td align="center">4.38E-03</td>
<td align="center">1.45E-02</td>
</tr>
<tr>
<td align="left">PA:SA</td>
<td align="center">5.83E-05</td>
<td align="center">&#x2212;2.13E-05</td>
<td align="center">3.09E-07</td>
<td align="center">1.43E-04</td>
<td align="center">1.87E-09</td>
<td align="center">4.67E-03</td>
<td align="center">8.23E-03</td>
</tr>
<tr>
<td align="left">POA:OA</td>
<td align="center">&#x2212;2.08E-05</td>
<td align="center">&#x2212;2.10E-06</td>
<td align="center">8.08E-09</td>
<td align="center">&#x2212;2.69E-05</td>
<td align="center">4.69E-11</td>
<td align="center">5.92E-03</td>
<td align="center">1.70E-02</td>
</tr>
<tr>
<td align="left">SA:OA</td>
<td align="center">&#x2212;8.15E-05</td>
<td align="center">&#x2212;1.19E-05</td>
<td align="center">3.62E-07</td>
<td align="center">&#x2212;3.33E-05</td>
<td align="center">3.36E-10</td>
<td align="center">6.86E-03</td>
<td align="center">8.21E-03</td>
</tr>
<tr>
<td align="left">LA:GLA</td>
<td align="center">&#x2212;7.35E-02</td>
<td align="center">2.03E-02</td>
<td align="center">1.32E&#x2b;00</td>
<td align="center">&#x2212;5.10E-01</td>
<td align="center">1.95E-02</td>
<td align="center">3.11E-02</td>
<td align="center">3.17E-02</td>
</tr>
<tr>
<td align="left">LA:DGLA</td>
<td align="center">2.04E-03</td>
<td align="center">1.99E-04</td>
<td align="center">3.58E-04</td>
<td align="center">&#x2212;1.84E-03</td>
<td align="center">3.12E-07</td>
<td align="center">2.33E-02</td>
<td align="center">4.08E-02</td>
</tr>
<tr>
<td align="left">GLA:DGLA</td>
<td align="center">4.58E-05</td>
<td align="center">1.04E-05</td>
<td align="center">2.64E-07</td>
<td align="center">&#x2212;7.62E-06</td>
<td align="center">1.13E-10</td>
<td align="center">3.36E-02</td>
<td align="center">6.50E-02</td>
</tr>
<tr>
<td align="left">DGLA:AA</td>
<td align="center">&#x2212;1.99E-05</td>
<td align="center">&#x2212;1.54E-06</td>
<td align="center">2.84E-08</td>
<td align="center">3.79E-05</td>
<td align="center">1.09E-10</td>
<td align="center">9.37E-03</td>
<td align="center">1.33E-02</td>
</tr>
<tr>
<td align="left">AA:DTA</td>
<td align="center">1.76E-03</td>
<td align="center">&#x2212;4.20E-04</td>
<td align="center">7.21E-05</td>
<td align="center">&#x2212;3.76E-03</td>
<td align="center">9.42E-07</td>
<td align="center">9.73E-03</td>
<td align="center">2.38E-02</td>
</tr>
<tr>
<td align="left">EPA:DPA_N3</td>
<td align="center">2.29E-04</td>
<td align="center">&#x2212;1.38E-05</td>
<td align="center">1.31E-06</td>
<td align="center">&#x2212;4.04E-04</td>
<td align="center">1.02E-08</td>
<td align="center">2.34E-02</td>
<td align="center">4.72E-02</td>
</tr>
<tr>
<td align="left">DTA:DPA_N6</td>
<td align="center">&#x2212;1.49E-03</td>
<td align="center">1.66E-04</td>
<td align="center">7.46E-04</td>
<td align="center">&#x2212;1.16E-02</td>
<td align="center">8.24E-06</td>
<td align="center">4.71E-02</td>
<td align="center">1.43E-01</td>
</tr>
<tr>
<td align="left">DPA_N3:DHA</td>
<td align="center">&#x2212;4.12E-04</td>
<td align="center">6.03E-05</td>
<td align="center">3.45E-06</td>
<td align="center">&#x2212;1.52E-04</td>
<td align="center">2.53E-09</td>
<td align="center">1.56E-02</td>
<td align="center">3.75E-02</td>
</tr>
<tr>
<td align="left">Overall</td>
<td align="center">&#x2212;1.30E-03</td>
<td align="center">2.93E-03</td>
<td align="center">1.14E-01</td>
<td align="center">&#x2212;4.80E-02</td>
<td align="center">2.12E-02</td>
<td align="center">1.79E-02</td>
<td align="center">3.77E-02</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>
<xref ref-type="table" rid="T4">Table&#x20;4</xref> summarizes the number of SNPs found significant when modeling using both IPD and PCSS across all 12 &#xd7; 362330 models. Of the ten SNPs for which IPD and PCSS models disagreed, nine occurred when one approach found a SNP to have a significant effect while the other found it to have a borderline significant effect (<italic>&#x3b1;</italic> &#x2264; <italic>p</italic>&#x20;&#x3c; 10<italic>&#x3b1;</italic>).</p>
<table-wrap id="T4" position="float">
<label>TABLE 4</label>
<caption>
<p>Summary of test decisions for a real data application calculating the linear model Fatty Acid Ratio &#x223c; snp &#x2b; age &#x2b; sex using individual participant data (IPD) and pre-computed summary statistics (PCSS) across 362,330 SNPs from 1,455 subjects in the Framing Heart Study&#x2019;s Offspring and Generation-3 cohorts. Significance threshold of <italic>&#x3b1;</italic> &#x3d; 1.37 &#xd7; 10<sup>&#x2212;7</sup>.</p>
</caption>
<table>
<thead valign="top">
<tr>
<th align="left">Ratio</th>
<th align="center">IPD significant</th>
<th align="center">PCSS significant</th>
<th align="center">Both significant</th>
</tr>
</thead>
<tbody valign="top">
<tr>
<td align="left">PA:POA</td>
<td align="center">0</td>
<td align="center">0</td>
<td align="center">0</td>
</tr>
<tr>
<td align="left">PA:SA</td>
<td align="center">0</td>
<td align="center">0</td>
<td align="center">0</td>
</tr>
<tr>
<td align="left">POA:OA</td>
<td align="center">0</td>
<td align="center">0</td>
<td align="center">0</td>
</tr>
<tr>
<td align="left">SA:OA</td>
<td align="center">6</td>
<td align="center">9</td>
<td align="center">6</td>
</tr>
<tr>
<td align="left">LA:GLA</td>
<td align="center">5</td>
<td align="center">2</td>
<td align="center">2</td>
</tr>
<tr>
<td align="left">LA:DGLA</td>
<td align="center">9</td>
<td align="center">10</td>
<td align="center">9</td>
</tr>
<tr>
<td align="left">GLA:DGLA</td>
<td align="center">8</td>
<td align="center">8</td>
<td align="center">8</td>
</tr>
<tr>
<td align="left">DGLA:AA</td>
<td align="center">18</td>
<td align="center">19</td>
<td align="center">18</td>
</tr>
<tr>
<td align="left">AA:DTA</td>
<td align="center">0</td>
<td align="center">0</td>
<td align="center">0</td>
</tr>
<tr>
<td align="left">EPA:DPA_N3</td>
<td align="center">0</td>
<td align="center">1</td>
<td align="center">0</td>
</tr>
<tr>
<td align="left">DTA:DPA_N6</td>
<td align="center">5</td>
<td align="center">4</td>
<td align="center">4</td>
</tr>
<tr>
<td align="left">DPA_N3:DHA</td>
<td align="center">11</td>
<td align="center">11</td>
<td align="center">11</td>
</tr>
<tr>
<td align="left">Overall</td>
<td align="center">62</td>
<td align="center">64</td>
<td align="center">58</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec sec-type="discussion" id="s4">
<title>4 Discussion</title>
<p>We have developed a method that approximates the covariance of products of phenotypes with other variables using only bivariate and univariate pre-computed summary statistics (PCSS). We then demonstrated how this covariance estimation can be used to approximate linear models for products of phenotypes, how these can model logical &#x201c;and&#x201d; and &#x201c;or&#x201d; statements and how these models can include researchers choice of covariates. We demonstrated our approximation method&#x2019;s accuracy relative to models fit on individual participant data through multiple simulations and applications to real genetic&#x20;data.</p>
<p>The approximations showed good performance overall. In a wide variety of simulations, the Type I error was maintained, the bias in point estimates relative to models fit with IPD was minimal, and hypothesis tests almost always reached the same conclusion as would be obtained with IPD. Application of our method to real data from the Framingham Heart Study also showed good performance on similar metrics. In general, we have tried to formulate this PCSS method to only rely on commonly available or easily estimated PCSS. However, in our application we assumed that we had the PCSS for reciprocals of fatty acids. This may not always be the case in practice, but may suggest that these PCSS may be important to pre-compute to assist downstream analyses of ratios.</p>
<p>Despite these positive results, some limitations of our work are worth noting. First, we used linear regression for a binary response. Previous applications of PCSS have taken this approach (<xref ref-type="bibr" rid="B3">Canela-Xandri et&#x20;al., 2018</xref>), and it is generally robust; however, this approach is less precise than when the underlying relationship is truly linear. While some foundations for a logistic modelling approach were recently proposed by <xref ref-type="bibr" rid="B28">Wu et&#x20;al. (2021)</xref>, further work is needed to develop a comprehensive model for logistic regression using PCSS. We also note that our method makes assumptions about the fit of the linear model to the data. While these assumptions are the same as in the corresponding analysis of IPD data (e.g., true underlying linear relationship between <bold>
<italic>y</italic>
</bold> and <bold>
<italic>x</italic>
</bold>), these assumptions may be more acutely important in our PCSS method.</p>
<p>Second, while our simulation study was comprehensive and we demonstrated our method on real data we note that further testing on simulated and real data is encouraged to explore special cases not considered here. Situations that suggest further testing and methodological developments include modeling linear combinations of products and adjusting for clustered/family data. One scenario of particular importance that warrants further investigation is missing data resulting in PCSS being computed on different subsets of the full data. While our real data analysis showed good performance in a setting with missing genotype data, further investigations should be performed to address the robustness of this method (along with other methods based around PCSS) when dealing with missing phenotype information. Similarly, researchers may be limited in their ability to have additional <italic>a priori</italic> exclusion criteria applied to their analysis or obtain PCSS calculated using the same criteria for each phenotype. There are, however, two options. Either obtain summary statistics on the subgroup of interest or ignore the exclusion criteria by including the group in the &#x201c;controls&#x201d; in an analysis on a dichotomous outcome. More robust options are in development (e.g., multinomial regression, or methods that use summary statistics alone to allow researchers to apply exclusion criteria). Relatedly, while our method exhibited fair performance when modeling logical combinations of binary phenotypes with low case-control ratios (see <xref ref-type="sec" rid="s10">Supplementary Table S4</xref>), it would benefit from further and more thorough work to assess its robustness.</p>
<p>Third, when estimating the variance of a product of a set of continuous phenotypes we estimated this term as a function of each covariate and took the median all these estimates as the approximated value. While this approach works well in practice, it may be possible to utilize the joint distribution of the covariates to estimate the covariance of these estimations and derive a more optimal estimation. Relatedly, our simulations showed that when multiplying binary phenotypes that exhibit high negative correlation and when multiplying phenotypes that take on negative values care should be taken. Finally and relatedly is the issue of potential compounding of errors when the method is applied to products of <italic>m</italic> phenotypes (where <italic>m</italic> is large). Although there are meaningful combined phenotypes that consist of five or more phenotypes (upon which this method should be used cautiously), we note that many combined phenotypes [e.g., BMI, ratios of biomarkers, cardiovascular disease (defined as either coronary heart disease or stroke)] are combinations of four or fewer phenotypes. Additional simulation studies and methodological improvements are needed and caution should be exhibited when applying our method in these&#x20;cases.</p>
<p>Fourth and finally, this method does not support PCSS that describe score tests (where a null model with non-genetic covariates is first fit and then updated for each genetic variant instead of simultaneously estimating both the genetic and covariate effects) in its current form. While future research can likely expanded this method to work with this data, we note the considerable collection of PCSS repositories (e.g., PheWeb, GWAS Catalog, GeneAtlas) which do provide the PCSS needed to perform these proposed approximations.</p>
<p>The use of PCSS provides numerous advantages over IPD data including computational efficiency and reduced concerns about data privacy. However, substantially improved and flexible methods are needed in order to fully leverage PCSS in customized downstream analyses. Our method allows researchers further customization of analyzed phenotypes by opening the door to multiplicative combinations of phenotypes, including logical combinations of binary phenotypes. Approximations used are reasonable, with near perfect maintenance of the Type I error rate and power in most situations. Further work is needed to apply the method to additional datasets and to expand the method to larger classes of combined phenotypes.</p>
</sec>
</body>
<back>
<sec id="s5">
<title>Data Availability Statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found here: <ext-link ext-link-type="uri" xlink:href="https://www.ncbi.nlm.nih.gov/gap/">https://www.ncbi.nlm.nih.gov/gap/</ext-link> (accessions phs000007.v29.p10 and phs000342.v20.p13).</p>
</sec>
<sec id="s6">
<title>Author Contributions</title>
<p>JWo and JWe developed the proposed statistical method. JWo performed the simulation analyses. JWe performed the real data analysis. JWo wrote the first draft of the manuscript. JWo and NT wrote sections of the manuscript. All authors contributed to manuscript revision, read, and approved the submitted version.</p>
</sec>
<sec id="s7">
<title>Funding</title>
<p>The authors of this work were supported by NIH Grant 2R15HG006915-03 and Dordt University.</p>
</sec>
<sec sec-type="COI-statement" id="s8">
<title>Conflict of Interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x2019;s Note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<ack>
<p>The authors would like to thank Martha Barnard, Xueting Xia, Nathan Ryder, and Jason Vander Woude for their help with preliminary stages of this project. This manuscript originally appeared online as a preprint on bioRxiv (<xref ref-type="bibr" rid="B27">Wolf et&#x20;al., 2021</xref>).</p>
</ack>
<sec id="s10">
<title>Supplementary Material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/fgene.2021.745901/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/fgene.2021.745901/full&#x23;supplementary-material</ext-link>
</p>
<supplementary-material xlink:href="DataSheet1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="fn1">
<label>1</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://www.genome.gov/Funded-Programs-Projects/Electronic-Medical-Records-and-Genomics-Network-eMERGE">https://www.genome.gov/Funded-Programs-Projects/Electronic-Medical-Records-and-Genomics-Network-eMERGE</ext-link>
</p>
</fn>
<fn id="fn2">
<label>2</label>
<p>
<ext-link ext-link-type="uri" xlink:href="http://pheweb.sph.umich.edu/">http://pheweb.sph.umich.edu/</ext-link>
</p>
</fn>
<fn id="fn3">
<label>3</label>
<p>
<ext-link ext-link-type="uri" xlink:href="http://pheweb.sph.umich.edu/">http://pheweb.sph.umich.edu/</ext-link>
</p>
</fn>
<fn id="fn4">
<label>4</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://www.leelabsg.org/resources">https://www.leelabsg.org/resources</ext-link>
</p>
</fn>
<fn id="fn5">
<label>5</label>
<p>
<ext-link ext-link-type="uri" xlink:href="https://cran.r-project.org/package=pcsstools">https://cran.r-project.org/package&#x3d;pcsstools</ext-link>
</p>
</fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Baba</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Shibata</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Sibuya</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2004</year>). <article-title>Partial Correlation and Conditional Correlation as Measures of Conditional independence</article-title>. <source>Aust. New&#x20;Zealand J.&#x20;Stat.</source> <volume>46</volume>, <fpage>657</fpage>&#x2013;<lpage>664</lpage>. <pub-id pub-id-type="doi">10.1111/j.1467-842X.2004.00360.x</pub-id> </citation>
</ref>
<ref id="B2">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Bycroft</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Freeman</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Petkova</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Band</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Elliott</surname>
<given-names>L. T.</given-names>
</name>
<name>
<surname>Sharp</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>The UK Biobank Resource with Deep Phenotyping and Genomic Data</article-title>. <source>Nature</source> <volume>562</volume>, <fpage>203</fpage>&#x2013;<lpage>209</lpage>. <pub-id pub-id-type="doi">10.1038/s41586-018-0579-z</pub-id> </citation>
</ref>
<ref id="B3">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Canela-Xandri</surname>
<given-names>O.</given-names>
</name>
<name>
<surname>Rawlik</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Tenesa</surname>
<given-names>A.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>An Atlas of Genetic Associations in UK Biobank</article-title>. <source>Nat. Genet.</source> <volume>50</volume>, <fpage>1593</fpage>&#x2013;<lpage>1599</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-018-0248-z</pub-id> </citation>
</ref>
<ref id="B4">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Cox</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>UK Biobank Shares the Promise of Big Data</article-title>. <source>Nature</source> <volume>562</volume>, <fpage>194</fpage>&#x2013;<lpage>195</lpage>. <pub-id pub-id-type="doi">10.1038/d41586-018-06948-3</pub-id> </citation>
</ref>
<ref id="B5">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Diogo</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Tian</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Franklin</surname>
<given-names>C. S.</given-names>
</name>
<name>
<surname>Alanne-Kinnunen</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>March</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Spencer</surname>
<given-names>C. C. A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>Phenome-wide Association Studies across Large Population Cohorts Support Drug Target Validation</article-title>. <source>Nat. Commun.</source> <volume>9</volume>, <fpage>4285</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-018-06540-3</pub-id> </citation>
</ref>
<ref id="B6">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dutta</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Gagliano Taliun</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Weinstock</surname>
<given-names>J.&#x20;S.</given-names>
</name>
<name>
<surname>Zawistowski</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Sidore</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Fritsche</surname>
<given-names>L. G.</given-names>
</name>
<etal/>
</person-group> (<year>2019a</year>). <article-title>Meta-MultiSKAT: Multiple Phenotype Meta-Analysis for Region-Based Association Test</article-title>. <source>Genet. Epidemiol.</source> <volume>43</volume>, <fpage>800</fpage>&#x2013;<lpage>814</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.22248</pub-id> </citation>
</ref>
<ref id="B7">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Dutta</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Scott</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Boehnke</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lee</surname>
<given-names>S.</given-names>
</name>
</person-group> (<year>2019b</year>). <article-title>Multi-SKAT: General Framework to Test for Rare-Variant Association with Multiple Phenotypes</article-title>. <source>Genet. Epidemiol.</source> <volume>43</volume>, <fpage>4</fpage>&#x2013;<lpage>23</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.22156</pub-id> </citation>
</ref>
<ref id="B8">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gagliano Taliun</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>VandeHaar</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Boughton</surname>
<given-names>A. P.</given-names>
</name>
<name>
<surname>Welch</surname>
<given-names>R. P.</given-names>
</name>
<name>
<surname>Taliun</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Schmidt</surname>
<given-names>E. M.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Exploring and Visualizing Large-Scale Genetic Associations by Using PheWeb</article-title>. <source>Nat. Genet.</source> <volume>52</volume>, <fpage>550</fpage>&#x2013;<lpage>552</lpage>. <pub-id pub-id-type="doi">10.1038/s41588-020-0622-5</pub-id> </citation>
</ref>
<ref id="B9">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Gasdaska</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Friend</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Chen</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Westra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Zawistowski</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Lindsey</surname>
<given-names>W.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Leveraging Summary Statistics to Make Inferences about Complex Phenotypes in Large Biobanks</article-title>. <source>Pac. Symp. Biocomputing</source> <volume>24</volume>, <fpage>391</fpage>&#x2013;<lpage>402</lpage>. <pub-id pub-id-type="doi">10.1142/9789813279827_0036</pub-id> </citation>
</ref>
<ref id="B10">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Guo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Wu</surname>
<given-names>B.</given-names>
</name>
</person-group> (<year>2019</year>). <article-title>Integrate Multiple Traits to Detect Novel Trait&#x2013;Gene Association Using GWAS Summary Data with an Adaptive Test Approach</article-title>. <source>Bioinformatics</source> <volume>35</volume>, <fpage>2251</fpage>&#x2013;<lpage>2257</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/bty961</pub-id> </citation>
</ref>
<ref id="B11">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Heatherly</surname>
<given-names>R.</given-names>
</name>
</person-group> (<year>2016</year>). <article-title>Privacy and Security within Biobanking: The Role of Information Technology</article-title>. <source>J.&#x20;L. Med. Ethics</source> <volume>44</volume>, <fpage>156</fpage>&#x2013;<lpage>160</lpage>. <pub-id pub-id-type="doi">10.1177/1073110516644206</pub-id> </citation>
</ref>
<ref id="B12">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Imamura</surname>
<given-names>F.</given-names>
</name>
<name>
<surname>Fretts</surname>
<given-names>A. M.</given-names>
</name>
<name>
<surname>Marklund</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Korat</surname>
<given-names>A. V. A.</given-names>
</name>
<name>
<surname>Yang</surname>
<given-names>W.-S.</given-names>
</name>
<name>
<surname>Lankinen</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Fatty Acids in the De Novo Lipogenesis Pathway and Incidence of Type 2 Diabetes: A Pooled Analysis of Prospective Cohort Studies</article-title>. <source>PLOS Med.</source> <volume>17</volume>, <fpage>e1003102</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pmed.1003102</pub-id> </citation>
</ref>
<ref id="B13">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Jones</surname>
<given-names>E. M.</given-names>
</name>
<name>
<surname>Sheehan</surname>
<given-names>N. A.</given-names>
</name>
<name>
<surname>Masca</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Wallace</surname>
<given-names>S. E.</given-names>
</name>
<name>
<surname>Murtagh</surname>
<given-names>M. J.</given-names>
</name>
<name>
<surname>Burton</surname>
<given-names>P. R.</given-names>
</name>
</person-group> (<year>2012</year>). <article-title>DataSHIELD &#x2013; Shared Individual-Level Analysis without Sharing the Data: a Biostatistical Perspective</article-title>. <source>Norsk Epidemiologi</source> <volume>21</volume>, <fpage>1499</fpage>. <pub-id pub-id-type="doi">10.5324/nje.v21i2</pub-id> </citation>
</ref>
<ref id="B14">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Justice</surname>
<given-names>A. E.</given-names>
</name>
<name>
<surname>Winkler</surname>
<given-names>T. W.</given-names>
</name>
<name>
<surname>Feitosa</surname>
<given-names>M. F.</given-names>
</name>
<name>
<surname>Graff</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Fisher</surname>
<given-names>V. A.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>K.</given-names>
</name>
<etal/>
</person-group> (<year>2017</year>). <article-title>Genome-wide Meta-Analysis of 241,258 Adults Accounting for Smoking Behaviour Identifies Novel Loci for Obesity Traits</article-title>. <source>Nat. Commun.</source> <volume>8</volume>, <fpage>14977</fpage>. <pub-id pub-id-type="doi">10.1038/ncomms14977</pub-id> </citation>
</ref>
<ref id="B15">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kalsbeek</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Veenstra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Westra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Disselkoen</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Koch</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>McKenzie</surname>
<given-names>K. A.</given-names>
</name>
<etal/>
</person-group> (<year>2018</year>). <article-title>A Genome-wide Association Study of Red-Blood Cell Fatty Acids and Ratios Incorporating Dietary Covariates: Framingham Heart Study Offspring Cohort</article-title>. <source>PLoS ONE</source> <volume>13</volume>. <pub-id pub-id-type="doi">10.1371/journal.pone.0194882</pub-id> </citation>
</ref>
<ref id="B16">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Kim</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Bai</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Pan</surname>
<given-names>W.</given-names>
</name>
</person-group> (<year>2015</year>). <article-title>An Adaptive Association Test for Multiple Phenotypes with GWAS Summary Statistics</article-title>. <source>Genet. Epidemiol.</source> <volume>39</volume>, <fpage>651</fpage>&#x2013;<lpage>663</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.21931</pub-id> </citation>
</ref>
<ref id="B17">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Lemaitre</surname>
<given-names>R. N.</given-names>
</name>
<name>
<surname>Tanaka</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tang</surname>
<given-names>W.</given-names>
</name>
<name>
<surname>Manichaikul</surname>
<given-names>A.</given-names>
</name>
<name>
<surname>Foy</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Kabagambe</surname>
<given-names>E. K.</given-names>
</name>
<etal/>
</person-group> (<year>2011</year>). <article-title>Genetic Loci Associated with Plasma Phospholipid N-3 Fatty Acids: A Meta-Analysis of Genome-wide Association Studies from the CHARGE Consortium</article-title>. <source>PLOS Genet.</source> <volume>7</volume>, <fpage>e1002193</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1002193</pub-id> </citation>
</ref>
<ref id="B18">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Li</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Zhang</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Sha</surname>
<given-names>Q.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Joint Analysis of Multiple Phenotypes Using a Clustering Linear Combination Method Based on Hierarchical Clustering</article-title>. <source>Genet. Epidemiol.</source> <volume>44</volume>, <fpage>67</fpage>&#x2013;<lpage>78</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.22263</pub-id> </citation>
</ref>
<ref id="B19">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Mailman</surname>
<given-names>M. D.</given-names>
</name>
<name>
<surname>Feolo</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Jin</surname>
<given-names>Y.</given-names>
</name>
<name>
<surname>Kimura</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Tryka</surname>
<given-names>K.</given-names>
</name>
<name>
<surname>Bagoutdinov</surname>
<given-names>R.</given-names>
</name>
<etal/>
</person-group> (<year>2007</year>). <article-title>The NCBI dbGaP Database of Genotypes and Phenotypes</article-title>. <source>Nat. Genet.</source> <volume>39</volume>, <fpage>1181</fpage>&#x2013;<lpage>1186</lpage>. <pub-id pub-id-type="doi">10.1038/ng1007-1181</pub-id> </citation>
</ref>
<ref id="B20">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Pasaniuc</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Price</surname>
<given-names>A. L.</given-names>
</name>
</person-group> (<year>2017</year>). <article-title>Dissecting the Genetics of Complex Traits Using Summary Association Statistics</article-title>. <source>Nat. Rev. Genet.</source> <volume>18</volume>, <fpage>117</fpage>&#x2013;<lpage>127</lpage>. <pub-id pub-id-type="doi">10.1038/nrg.2016.142</pub-id> </citation>
</ref>
<ref id="B21">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Ray</surname>
<given-names>D.</given-names>
</name>
<name>
<surname>Boehnke</surname>
<given-names>M.</given-names>
</name>
</person-group> (<year>2018</year>). <article-title>Methods for Meta-Analysis of Multiple Traits Using GWAS Summary Statistics</article-title>. <source>Genet. Epidemiol.</source> <volume>42</volume>, <fpage>134</fpage>&#x2013;<lpage>145</lpage>. <pub-id pub-id-type="doi">10.1002/gepi.22105</pub-id> </citation>
</ref>
<ref id="B22">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Simell</surname>
<given-names>B. A.</given-names>
</name>
<name>
<surname>T&#xf6;rnwall</surname>
<given-names>O. M.</given-names>
</name>
<name>
<surname>H&#xe4;m&#xe4;l&#xe4;inen</surname>
<given-names>I.</given-names>
</name>
<name>
<surname>Wichmann</surname>
<given-names>H. E.</given-names>
</name>
<name>
<surname>Anton</surname>
<given-names>G.</given-names>
</name>
<name>
<surname>Brennan</surname>
<given-names>P.</given-names>
</name>
<etal/>
</person-group> (<year>2019</year>). <article-title>Transnational Access to Large Prospective Cohorts in Europe: Current Trends and Unmet Needs</article-title>. <source>New Biotechnol.</source> <volume>49</volume>, <fpage>98</fpage>&#x2013;<lpage>103</lpage>. <pub-id pub-id-type="doi">10.1016/j.nbt.2018.10.001</pub-id> </citation>
</ref>
<ref id="B23">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tintle</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Bassett</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Kuo-Liong</surname>
<given-names>C.</given-names>
</name>
<name>
<surname>Forouhi</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Kupers</surname>
<given-names>L.</given-names>
</name>
<name>
<surname>Lankinen</surname>
<given-names>M.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Circulating Omega-3 Fatty Acid Levels and Total and Cause-specific Mortality: Prospective Evidence from 14 Cohorts in the Fatty Acids and Outcomes Research Consortium</article-title>. <source>Circulation</source> <volume>141</volume>, <fpage>A43</fpage>. <pub-id pub-id-type="doi">10.1161/circ.141</pub-id> </citation>
</ref>
<ref id="B24">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Tintle</surname>
<given-names>N. L.</given-names>
</name>
<name>
<surname>Pottala</surname>
<given-names>J.&#x20;V.</given-names>
</name>
<name>
<surname>Lacey</surname>
<given-names>S.</given-names>
</name>
<name>
<surname>Ramachandran</surname>
<given-names>V.</given-names>
</name>
<name>
<surname>Westra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Rogers</surname>
<given-names>A.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>A Genome-wide Association Study of Saturated, Mono- and Polyunsaturated Red Blood Cell Fatty Acids in the Framingham Heart Offspring Study</article-title>. <source>Prostaglandins, Leukot. Essent. Fatty Acids</source> <volume>94</volume>, <fpage>65</fpage>&#x2013;<lpage>72</lpage>. <pub-id pub-id-type="doi">10.1016/j.plefa.2014.11.007</pub-id> </citation>
</ref>
<ref id="B25">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>von Berg</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>van der Laan</surname>
<given-names>S. W.</given-names>
</name>
<name>
<surname>McArdle</surname>
<given-names>P. F.</given-names>
</name>
<name>
<surname>Malik</surname>
<given-names>R.</given-names>
</name>
<name>
<surname>Kittner</surname>
<given-names>S. J.</given-names>
</name>
<name>
<surname>Mitchell</surname>
<given-names>B. D.</given-names>
</name>
<etal/>
</person-group> (<year>2020</year>). <article-title>Alternate Approach to Stroke Phenotyping Identifies a Genetic Risk Locus for Small Vessel Stroke</article-title>. <source>Eur. J.&#x20;Hum. Genet. EJHG</source> <volume>28</volume>, <fpage>963</fpage>&#x2013;<lpage>972</lpage>. <pub-id pub-id-type="doi">10.1038/s41431-020-0580-5</pub-id> </citation>
</ref>
<ref id="B26">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolf</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Barnard</surname>
<given-names>M.</given-names>
</name>
<name>
<surname>Xia</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Ryder</surname>
<given-names>N.</given-names>
</name>
<name>
<surname>Westra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tintle</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2020</year>). <article-title>Computationally Efficient, Exact, Covariate-Adjusted Genetic Principal Component Analysis by Leveraging Individual Marker Summary Statistics from Large Biobanks</article-title>. <source>Pac. Symp. Biocomputing</source> <volume>25</volume>, <fpage>719</fpage>&#x2013;<lpage>730</lpage>. <pub-id pub-id-type="doi">10.1142/9789811215636</pub-id> </citation>
</ref>
<ref id="B27">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wolf</surname>
<given-names>J.&#x20;M.</given-names>
</name>
<name>
<surname>Westra</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Tintle</surname>
<given-names>N.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Using Summary Statistics to Evaluate the Genetic Architecture of Multiplicative Combinations of Initially Analyzed Phenotypes with a Flexible Choice of Covariates</article-title>. <comment>bioRxiv</comment>. <pub-id pub-id-type="doi">10.1101/2021.03.08.433979</pub-id> </citation>
</ref>
<ref id="B28">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Wu</surname>
<given-names>P.</given-names>
</name>
<name>
<surname>Wang</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Lubitz</surname>
<given-names>S. A.</given-names>
</name>
<name>
<surname>Benjamin</surname>
<given-names>E. J.</given-names>
</name>
<name>
<surname>Meigs</surname>
<given-names>J.&#x20;B.</given-names>
</name>
<name>
<surname>Dupuis</surname>
<given-names>J.</given-names>
</name>
</person-group> (<year>2021</year>). <article-title>Approximate Conditional Phenotype Analysis Based on Genome Wide Association Summary Statistics</article-title>. <source>Scientific Rep.</source> <volume>11</volume>, <fpage>2518</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-021-82000-1</pub-id> </citation>
</ref>
<ref id="B29">
<citation citation-type="journal">
<person-group person-group-type="author">
<name>
<surname>Zhu</surname>
<given-names>X.</given-names>
</name>
<name>
<surname>Feng</surname>
<given-names>T.</given-names>
</name>
<name>
<surname>Tayo</surname>
<given-names>B.</given-names>
</name>
<name>
<surname>Liang</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Young</surname>
<given-names>J.</given-names>
</name>
<name>
<surname>Franceschini</surname>
<given-names>N.</given-names>
</name>
<etal/>
</person-group> (<year>2015</year>). <article-title>Meta-analysis of Correlated Traits <italic>via</italic> Summary Statistics from GWASs with an Application in Hypertension</article-title>. <source>Am. J.&#x20;Hum. Genet.</source> <volume>96</volume>, <fpage>21</fpage>&#x2013;<lpage>36</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2014.11.011</pub-id> </citation>
</ref>
</ref-list>
</back>
</article>