<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Dement.</journal-id>
<journal-title>Frontiers in Dementia</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Dement.</abbrev-journal-title>
<issn pub-type="epub">2813-3919</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frdem.2023.1120206</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Dementia</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>An alternative method of SNP inclusion to develop a generalized polygenic risk score analysis across Alzheimer&#x00027;s disease cohorts</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Brookes</surname> <given-names>Keeley J.</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/1648867/overview"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Guetta-Baranes</surname> <given-names>Tamar</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Thomas</surname> <given-names>Alan</given-names></name>
<xref ref-type="aff" rid="aff3"><sup>3</sup></xref>
</contrib>
<contrib contrib-type="author">
<name><surname>Morgan</surname> <given-names>Kevin</given-names></name>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/583104/overview"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Interdisciplinary Biomedical Research Centre, Biosciences, Clifton Campus, Nottingham Trent University</institution>, <addr-line>Nottingham</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff2"><sup>2</sup><institution>Human Genetics, Life Sciences, University Park, University of Nottingham</institution>, <addr-line>Nottingham</addr-line>, <country>United Kingdom</country></aff>
<aff id="aff3"><sup>3</sup><institution>Brains for Dementia Research Coordinating Centre, Institute of Neuroscience, Newcastle University</institution>, <addr-line>Newcastle upon Tyne</addr-line>, <country>United Kingdom</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Georgina Menzies, Cardiff University, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Adrienne Tin, University of Mississippi Medical Center, United States; Giuseppe Tosto, Columbia University, United States</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Keeley J. Brookes <email>keeley.brookes&#x00040;ntu.ac.uk</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>31</day>
<month>07</month>
<year>2023</year>
</pub-date>
<pub-date pub-type="collection">
<year>2023</year>
</pub-date>
<volume>2</volume>
<elocation-id>1120206</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>12</month>
<year>2022</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>06</month>
<year>2023</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2023 Brookes, Guetta-Baranes, Thomas and Morgan.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Brookes, Guetta-Baranes, Thomas and Morgan</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license> </permissions>
<abstract>
<sec>
<title>Introduction</title>
<p>Polygenic risk scores (PRSs) have great clinical potential for detecting late-onset diseases such as Alzheimer&#x00027;s disease (AD), allowing the identification of those most at risk years before the symptoms present. Although many studies use various and complicated machine learning algorithms to determine the best discriminatory values for PRSs, few studies look at the commonality of the Single Nucleotide Polymorphisms (SNPs) utilized in these models.</p></sec>
<sec>
<title>Methods</title>
<p>This investigation focussed on identifying SNPs that tag blocks of linkage disequilibrium across the genome, allowing for a generalized PRS model across cohorts and genotyping panels. PRS modeling was conducted on five AD development cohorts, with the best discriminatory models exploring for a commonality of linkage disequilibrium clumps. Clumps that contributed to the discrimination of cases from controls that occurred in multiple cohorts were used to create a generalized model of PRS, which was then tested in the five development cohorts and three further AD cohorts.</p></sec>
<sec>
<title>Results</title>
<p>The model developed provided a discriminability accuracy average of over 70% in multiple AD cohorts and included variants of several well-known AD risk genes.</p></sec>
<sec>
<title>Discussion</title>
<p>A key element of devising a polygenic risk score that can be used in the clinical setting is one that has consistency in the SNPs that are used to calculate the score; this study demonstrates that using a model based on commonality of association findings rather than meta-analyses may prove useful.</p></sec></abstract>
<kwd-group>
<kwd>polygenic risk score</kwd>
<kwd>dementia</kwd>
<kwd>Alzheimer&#x00027;s disease</kwd>
<kwd>cross-cohort</kwd>
<kwd>predictability</kwd>
</kwd-group>
<contract-sponsor id="cn001">Alzheimer&#x00027;s Research UK<named-content content-type="fundref-id">10.13039/501100002283</named-content></contract-sponsor>
<counts>
<fig-count count="2"/>
<table-count count="2"/>
<equation-count count="0"/>
<ref-count count="36"/>
<page-count count="10"/>
<word-count count="7883"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Genetics and Biomarkers of Dementia</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>Introduction</title>
<p>The investigation of genetic predisposition to complex disease via polygenic risk scores (PRS) has increased in recent years (Plomin and von Stumm, <xref ref-type="bibr" rid="B23">2022</xref>), and its popularity is guided by its long-term promise of clinical utility for early diagnosis and personalized therapeutic intervention, leading to disease prevention. Meanwhile, PRSs have the potential to further basic research by identifying causal variants, leading to novel drug development, the stratification of samples for clinical exploration, and therapeutic intervention trials involving targets who are most likely to respond.</p>
<p>Briefly, PRS analyses utilize large-scale genome-wide association study (GWAS) summary data (base data set) to score genotypes present in the target data set, i.e., an independent cohort of case and control samples, producing a single risk score for each individual for the disease under investigation. PRSs are calculated as a sum of the number of &#x0201C;effect&#x0201D; alleles present, weighted by their effect size (provided by the GWAS summary statistics), and are then tested for their correlation with the disease phenotype and the ability to discriminate the cases from controls using the area-under-the-curve (AUC) statistic from the receiver operating characteristic (ROC) curve. The selection of SNPs to be included in the PRS is normally determined by a significance threshold in the summary statistics, incorporating all SNPs with GWAS <italic>p</italic>-values below that cut-off (Purcell et al., <xref ref-type="bibr" rid="B26">2009</xref>; Euesden et al., <xref ref-type="bibr" rid="B10">2015</xref>).</p>
<p>However, this type of analysis has not been without its criticism, and the methodology is still being developed from its original and simplest form of taking significant GWAS SNPs and summing them by their weighted effect sizes (Purcell et al., <xref ref-type="bibr" rid="B26">2009</xref>) to the more complex machine learning algorithms being explored (Baker and Escott-Price, <xref ref-type="bibr" rid="B4">2020</xref>). A crucial caveat to the development of a clinically relevant PRS is the lack of translation of PRS models across datasets, with initial discriminability estimates failing to be maintained on additional independent cohorts (Janssens, <xref ref-type="bibr" rid="B16">2019</xref>; Baker and Escott-Price, <xref ref-type="bibr" rid="B4">2020</xref>). This is due to the lack of commonality of SNPs between cohorts not only because of differences in the SNP content on genotyping platforms but also potentially because of differences in allele frequencies between cases and controls in different cohorts, which governs the discriminability and is, therefore, subject to the same nuances as gene association studies (Nakaoka and Inoue, <xref ref-type="bibr" rid="B22">2009</xref>; Shi et al., <xref ref-type="bibr" rid="B31">2016</xref>). The lack of consistency in SNP content can be overcome by imputing data to fill in missing genotypes between cohorts or increase the coverage of SNPs between the target and base data sets. However, there are concerns about imputation accuracy and its effect on the subsequent analysis for SNPs with low minor allele frequencies, low levels of linkage disequilibrium, and those that may have gone through recent and dramatic allele frequency changes (Aleknonyte-Resch et al., <xref ref-type="bibr" rid="B2">2021</xref>; Ali et al., <xref ref-type="bibr" rid="B3">2022</xref>). The International Genomics of Alzheimer&#x00027;s Project (IGAP) Stage 1 (IGAP_S1) GWAS summary statistics (Lambert et al., <xref ref-type="bibr" rid="B19">2013</xref>) are often used as a base data set for PRS studies in Alzheimer&#x00027;s disease (AD), consisting of more than 7 million imputed genotypes. During the imputation and phasing of genetic data, some genotypes are lost in the process, for example, the rs9271192 (<italic>HLA-DRB5</italic>) SNP, which was one of the GWAS hits from the Lambert et al. (<xref ref-type="bibr" rid="B19">2013</xref>) study but was absent from the summary statistics, presumably having been lost in the imputation process. Furthermore, the focus in PRS studies is often on the ability to accurately identify cases from controls, with the extent of SNP commonality between the models developed across studies still to be fully explored to identify a consistent set of SNPs that could be used for clinical purposes.</p>
<p>Further commentary suggests that the traditional pruning/clumping of SNPs, selecting those that are the most significant (denoted as the &#x0201C;Index SNP&#x0201D;) within blocks of linkage disequilibrium (LD), followed by the &#x0201C;thresholding&#x0201D; for SNP inclusivity, falls short of explaining the heritability estimates and could discard information that increases prediction accuracy (Vilhj&#x000E1;lmsson et al., <xref ref-type="bibr" rid="B33">2015</xref>). Conversely, relaxing the thresholds to include SNPs outside the GWAS significance threshold has yielded the PRS models utilizing thousands of SNPs, yet whilst this has provided greater discriminability (Escott-Price et al., <xref ref-type="bibr" rid="B9">2015</xref>), the clinical utility of including such SNPs, often with negligible effect sizes, is questionable; if they do not impact the biology of the disease or suggest a druggable target, are they useful to include (Janssens, <xref ref-type="bibr" rid="B16">2019</xref>).</p>
<p>This investigation utilized the summary statistics from the Lambert et al. (<xref ref-type="bibr" rid="B19">2013</xref>) study and a cross-cohort methodology for SNP selection for PRS model generation in AD (see the Data Availability section for details of the cohorts used). The incorporation of SNPs into the PRS model is traditionally determined by their significance in the summary statistics base data set; however, the discrimination of cases and controls rests with the allele frequency difference of the selected SNPs in the target data set. Incorporating SNPs with large allele frequency differences between the cases and controls in the target data set will lead to more divergent PRSs between these two groups and, consequently, a more accurate prediction model. In contrast to the traditional thresholding method, in this study, SNPs were selected based on their consistent contribution to a highly discriminatory PRS model across several cohorts. It was hypothesized that, by doing this, SNPs that underlie the disease will be selected and improve the translatability of the PRS model across different data sets. The developed PRS models were then validated with three independent AD cohorts, showing an average of more than 70% accuracy in discriminability across all cohorts.</p></sec>
<sec sec-type="methods" id="s2">
<title>Methods</title>
<sec>
<title>Data sets</title>
<sec>
<title>Base data set</title>
<p>Summary statistics were obtained from the Lambert et al. (<xref ref-type="bibr" rid="B19">2013</xref>) study. These summary statistics were generated from imputation during the first stage of this study (IGAP_S1) by conducting a meta-analysis of four GWAS samples of European ancestry (<italic>n</italic> = 17,008 cases, <italic>n</italic> = 37,154 controls). This data set was selected because of the number of SNPs in the summary statistics and because of generating the number of samples that had a diagnosis of AD.</p></sec>
<sec>
<title>PRS development data sets</title>
<p>Five genotyping data sets were utilized for this study; see the Data Availability section for details on the cohorts. The data sets used were the Brains for Dementia Research (BDR) cohort, as described in a previous study by this group (Young et al., <xref ref-type="bibr" rid="B35">2021</xref>), and four cohorts previously utilized in large GWAS studies, which are freely available for download: ADC7, NIA, ROSMAP and TGEN data sets. ROSMAP data were obtained from the AMP-AD Knowledge Portal via the Synapse Data Access System (<ext-link ext-link-type="uri" xlink:href="https://www.synapse.org/">https://www.synapse.org/</ext-link>). The ADC7, NIA, and TGEN data sets were downloaded from NIAGADS (<ext-link ext-link-type="uri" xlink:href="https://www.niagads.org/">https://www.niagads.org/</ext-link>). The data sets that were selected had been genotyped on different platforms, had varying sample sizes, and had <italic>APOE</italic> isoform/genotype data (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>).</p>
<p>All data sets underwent quality control with PLINK v1.9 (Purcell et al., <xref ref-type="bibr" rid="B25">2007</xref>), and SNPs with a minor allele frequency of &#x0003C;1% were removed. In addition, genotype calls of &#x0003C;95% and those deviating significantly from Hardy&#x02013;Weinberg equilibrium (<italic>p</italic> &#x0003C; 0.0001) in the control samples were also removed. Furthermore, from the available information, non-Caucasian samples including Hispanic samples from the ROSMAP data set were also removed. Only samples with a definite or confirmed diagnosis of AD were included. When the genotyping of <italic>APOE</italic> isoform SNPs (rs429358 &#x00026; rs7412) was not included in the genotype data, the isoform information in the accompanying clinical data files for each data set was used to determine the likely SNP genotypes.</p></sec>
<sec>
<title>Validation data sets</title>
<p>Three validation data sets (MTC, WashU, and TARCC) were also obtained from NIAGADS (<ext-link ext-link-type="uri" xlink:href="https://www.niagads.org/">https://www.niagads.org/</ext-link>), and details on these cohorts can be found in the Data Availability section. Samples from these data sets were included in this study if diagnosed as AD or control and had <italic>APOE</italic> genotyping. Details of the sample sizes and genotyping platforms used to generate their data are available in <xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>.</p></sec></sec>
<sec>
<title>Clumping and LD clump SNP assignment</title>
<p>The IGAP_S1 summary statistics were clumped using the 1000Genomes European data set in PLINK v1.9 (Purcell et al., <xref ref-type="bibr" rid="B25">2007</xref>), with the parameters &#x02013;clump-p1 1, &#x02013;clump-p2 1, &#x02013;clump-kb 250, and &#x02013;clump-r2 0.8. These parameters were then used to define the LD clumps across the base data set SNPs. The IGAP_S1 clumped output file consists of rows of clumped data, with the most significant SNP of the clump denoted as the &#x0201C;Index&#x0201D; and all other SNPs that reside within that clump in subsequent columns. Traditionally, only these Index SNPs that are in common with the target dataset are used in PRS modeling.</p>
<p>In this investigation, the following alternative steps were taken to identify SNPs to be included in the PRS modeling (the R Script is available on request):</p>
<list list-type="order">
<list-item><p>The clump/LD blocks identified in the IGAP_S1 data were numbered, and each SNP from the IGAP_S1 data that resides in that clump was &#x0201C;tagged&#x0201D; with its corresponding clump/LD number.</p></list-item>
<list-item><p>Clump Tag SNPs in the IGAP_S1 summary statistics were matched by SNP ID with those in each development data set.</p></list-item>
<list-item><p>A single SNP representing each clump was selected for each development data set.</p></list-item>
<list-item><p>Clumps that had representative SNPs across all five of the development data sets were taken forward to create the &#x0201C;Common Clump SNP set&#x0201D; for analysis.</p></list-item>
<list-item><p>The beta effect size scores and reference alleles were obtained from the IGAP_S1 summary statistics for all SNPs identified for the polygenic risk score algorithm, producing five sets of SNPs to be used, one for each of the development cohorts.</p></list-item>
</list></sec>
<sec>
<title>Common clump SNP set quality control</title>
<p>SNPs representing each clump across the PRS development data sets were investigated for large deviations in beta effect sizes (obtained from the IGAP_S1 and used to generate the risk scores in the target data sets). The SNPs tagged in each clump across the data sets were examined for differences in beta effect direction and those that displayed a standard deviation (<italic>SD</italic>) of &#x000B1;1 across the beta effect size scores. Clumps with opposing directions of beta effects or <italic>SD</italic> of &#x0003E; 0.5 were removed; this resulted in the removal of 1,067 clumps based on the differential direction of the beta values only. No further clumps were removed based on the beta value <italic>SD</italic> of the five SNPs representing the clump across the five PRS development data sets; the maximum <italic>SD</italic> observed was 0.003. In addition, a total of 42,684 clump-tagged SNPs were available for analysis per cohort.</p></sec>
<sec>
<title>Threshold modeling</title>
<p>SNPs from the base data set at each IGAP_S1 significance threshold from 5 &#x000D7; 10<sup>&#x02212;8</sup> to 1, increasing at intervals of 10<sup>&#x02212;6</sup>, were used to generate the score file for the PRS modeling using the &#x02013;score parameter in PLINK v1.9 (Purcell et al., <xref ref-type="bibr" rid="B25">2007</xref>). Logistic regression was carried out in R v4.0.3 (R Core team, <xref ref-type="bibr" rid="B27">2021</xref>), followed by calculating the AUC (accuracy of the model to discrimination case from control) using the pROC package (Robin et al., <xref ref-type="bibr" rid="B29">2011</xref>). The results are presented in <xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>. R The scripts are available upon request.</p></sec>
<sec>
<title>Perfect discrimination modeling</title>
<p>SNPs were recruited into the PRS model on an individual basis based on the <italic>p-</italic>value of association (generated in PLINK &#x02013;assoc analysis) in the developmental data set and were used to generate the score file for PRS modeling using the &#x02013;score parameter in PLINK v1.9 (Purcell et al., <xref ref-type="bibr" rid="B25">2007</xref>). Logistic regression was carried out in R v4.0.3 (R Core team, <xref ref-type="bibr" rid="B27">2021</xref>), followed by calculating the AUC using the pROC package (Robin et al., <xref ref-type="bibr" rid="B29">2011</xref>). The results were used to determine the effect that the addition of each SNP had on the AUC (either increase or decrease) and labeled it with the effect direction. The SNPs were then re-ordered by <italic>p-</italic>value and the increase in AUC effect before undergoing single SNP addition PRS modeling to generate perfect discrimination curves, that is, AUC = 1, where all samples above a certain risk score were cases and all the samples that were below were controls.</p></sec>
<sec>
<title>Generation of absolutes</title>
<p>&#x0201C;Absolute&#x0201D; controls and cases were artificially generated to represent the absolute minimum score (control) and the absolute maximum score (case) that could be achieved for each SNP model. This involved creating genotypes for individuals that were homozygous for all the SNPs with effect allele beta scores with a protective (minus) direction, and an individual with the opposing genotypes. These &#x0201C;absolutes&#x0201D; were then subjected to risk scoring using the generalized PRS model to create the minimum and maximum scores that could possibly be achieved, which were then used to calculate centile bins (100 bins) at equal intervals across the entire range of possible scores. The application of risk scoring helped determine the visualization of the true spread of possible scores and where the scores obtained from the data sets resided on this full scale.</p></sec></sec>
<sec sec-type="results" id="s3">
<title>Results</title>
<p>Clumping was carried out on the IGAP_S1 summary statistics with an <italic>r</italic><sup>2</sup> of &#x02265; .8 in PLINK v1.9 (Purcell et al., <xref ref-type="bibr" rid="B25">2007</xref>). The coverage of the genome was explored using the BDR cohort using the traditional method of utilizing only the &#x0201C;Index&#x0201D; SNP and then any SNP in the LD clump. The greater coverage of the LD clumps of the IGAP_S1 data set was achieved by allowing any SNP that was representative of an LD block into the PRS analysis rather than just the Index SNP (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 1</xref>). Previous A previous study (Farrell and Brookes, <xref ref-type="bibr" rid="B11">2022</xref>) suggests that additional SNPs within the surrounding region of <italic>APOE</italic> isoform SNPs (rs429358 and rs7412) could possibly be independently contributing to the AD phenotype, and therefore, the entire <italic>APOE</italic> region was retained in this analysis.</p>
<p>In addition to the BDR data set, four additional AD genotyping data sets (ADC7, NIA, ROSMAP, and TGEN) obtained from the NIAGADS data repository were selected across a range of genotyping platforms to develop the PRS model (<xref ref-type="supplementary-material" rid="SM1">Supplementary Table 5</xref>). Each SNP from the cohort was &#x0201C;tagged&#x0201D; with a clump number, and clumps that were covered by SNPs in all five of the PRS developmental data sets were included in the analysis (44,291 clumps). Cross-cohort SNPs for each clump were explored for consistency in the reference allele, the direction of effect, and the beta coefficient values. When the direction of effect and the reference allele differed between the SNPs in the same LD clump, the clump was removed. The beta coefficient values were found to be very similar, with an average <italic>SD</italic> of 0.003 and a maximum observed <italic>SD</italic> of 0.1 observed across the cohorts for each LD clump. A total of 42,684 LD clumps uniquely tagged by cohort-relevant SNPs were available for PRS model generation as opposed to 33,679 SNPs in common across the cohorts, increasing the coverage of the genome and commonality between data sets.</p>
<p>Traditional PRS modeling was applied to these data sets for comparison, since this methodology mimics that of the PRS software tool PRSice (Euesden et al., <xref ref-type="bibr" rid="B10">2015</xref>), where SNPs are recruited into the model based on the <italic>p-</italic>value threshold in the base data set from 5 &#x000D7; 10<sup>&#x02212;5</sup> to 1 at an interval of 10<sup>&#x02212;6</sup>. These SNP models were subjected to logistic regression, and an AUC was generated from the R software package pROC. The best PRS models were identified for each PRS development cohort in terms of the most significant logistic regression <italic>p-</italic>value and AUC (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 2</xref>, <bold>Panel A</bold>, and <xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>). The results were highly variable between cohorts, with the best models utilizing between 18 and 41,216 SNPs and achieving an average discrimination accuracy of 0.76 (<italic>SD</italic> 0.138).</p>
<p>Perfect discrimination models were applied to the 42,684 LD clumps of each cohort. Briefly, SNPs tagged in each clump were ordered by cohort significance values and incorporated into the PRS model sequentially (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 2</xref>, <bold>Panel B</bold>). These SNPs were then reordered based on their cohort significance and positive direction of effect on the AUC (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 2</xref>, <bold>Panel C</bold>). All SNPs contributing to the point where perfect discrimination (AUC = 1) was first achieved were taken from each cohort. Two thresholds of commonality were set for testing: a lenient threshold requiring the LD clump to contribute to perfect discrimination in two of the five cohorts and a stringent threshold requiring the LD clump to contribute to the perfect discrimination in three of the five cohorts.</p>
<p>The lenient threshold identified a 2207-LD-clump PRS model, which, when applied to the development cohorts, yielded highly significant correlations and discriminatory accuracies of over 90% in four of the cohorts. The ROSMAP cohort demonstrated an accuracy of 53.82%, with the PRSs showing no significant correlation with disease outcome (<xref ref-type="table" rid="T1">Table 1</xref>). When they were applied to the lower discriminability estimates of the validation cohorts, although each of these cohorts was missing several of the LD clumps that contributed to the model, only a small number of alternative SNPs tagging the LD clumps were identified (<xref ref-type="table" rid="T1">Table 1</xref>).</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Results of the lenient PRS model, in the development cohorts and validation cohorts, indicating a high discriminability of the model in the development cohorts but a significant drop in the validation cohorts.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>Development</bold></th>
<th valign="top" align="center"><bold>SNP platform</bold></th>
<th valign="top" align="center"><bold>&#x00023; SNPs from PD contributing to PRS model</bold></th>
<th valign="top" align="center"><bold>Logistic regression <italic>P</italic>-value</bold></th>
<th valign="top" align="center"><bold>Area under the curve</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ADC7<break/> <italic>N =</italic> 1462</td>
<td valign="top" align="center">HumanOmni Express</td>
<td valign="top" align="center">1761</td>
<td valign="top" align="center">4.32 x 10<sup>&#x02212;85</sup></td>
<td valign="top" align="center">0.9004</td>
</tr> <tr>
<td valign="top" align="left">BDR<break/> <italic>N =</italic> 520</td>
<td valign="top" align="center">NeuroChip</td>
<td valign="top" align="center">612</td>
<td valign="top" align="center">1.12 x 10<sup>&#x02212;30</sup></td>
<td valign="top" align="center">0.9259</td>
</tr> <tr>
<td valign="top" align="left">NIA<break/> <italic>N =</italic> 1,386</td>
<td valign="top" align="center">Human 610-quad</td>
<td valign="top" align="center">932</td>
<td valign="top" align="center">1.89 x 10<sup>&#x02212;76</sup></td>
<td valign="top" align="center">0.9463</td>
</tr> <tr>
<td valign="top" align="left">ROSMAP<break/> <italic>N =</italic> 240</td>
<td valign="top" align="center">Affymetrix6.0/IlluminaOmni Quad</td>
<td valign="top" align="center">41</td>
<td valign="top" align="center">0.397</td>
<td valign="top" align="center">0.5382</td>
</tr> <tr>
<td valign="top" align="left">TGEN<break/> <italic>N =</italic> 1,510</td>
<td valign="top" align="center">Affymetrix 6.0</td>
<td valign="top" align="center">1,221</td>
<td valign="top" align="center">7.07 x 10<sup>&#x02212;84</sup></td>
<td valign="top" align="center">0.9357</td>
</tr> 
<tr style="background-color:#919497;color:#ffffff">
<td valign="top" align="left"><bold>Validation</bold></td>
<td valign="top" align="center"><bold>SNP Platform</bold></td>
<td valign="top" align="center"><bold>&#x00023; SNPs available</bold></td>
<td valign="top" align="center"><bold>Logistic regression</bold> <italic><bold>P</bold></italic><bold>-value</bold></td>
<td valign="top" align="center"><bold>Area under the curve</bold></td>
</tr> <tr>
<td valign="top" align="left">MTC<break/> <italic>N =</italic> 356</td>
<td valign="top" align="center">HumanOmni Express</td>
<td valign="top" align="center">ADC7 Panel<break/> 2204 SNPs (7 alt)</td>
<td valign="top" align="center">2.75 x 10<sup>&#x02212;8</sup></td>
<td valign="top" align="center">0.6948</td>
</tr> <tr>
<td valign="top" align="left">WashU<break/> <italic>N =</italic> 131</td>
<td valign="top" align="center">HumanOmni Express</td>
<td valign="top" align="center">ADC7 Panel<break/> 1862 SNPs (22 alt)</td>
<td valign="top" align="center">0.064</td>
<td valign="top" align="center">0.6288</td>
</tr>
<tr>
<td valign="top" align="left">TARCC<break/> <italic>N =</italic> 491</td>
<td valign="top" align="center">Affymetrix 6.0</td>
<td valign="top" align="center">TGEN Panel<break/> 2070 SNPs (74 alt)</td>
<td valign="top" align="center">8.17 x 10<sup>&#x02212;14</sup></td>
<td valign="top" align="center">0.7062</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>The lenient model consists of SNPs tagging the 2207 LD clumps identified to be contributing to the perfect discrimination (PD) model in at least 2 of the development cohorts. In the top panel, the results from the development cohorts indicate the number of LD clumps from that cohort that were found in at least one other cohort, note that there is no correlation between the number of clumps identified in each cohort and the accuracy in the model. In the bottom panel, the results from the validation cohorts indicate the SNP panels used in the PRS and the number of alternative (alt) SNPs that had to be found to substitute for SNPs missing in that cohort.</p>
</table-wrap-foot>
</table-wrap>
<p>Using the stringent threshold, 149 LD clumps were identified as contributing to at least three of the perfect discrimination models, and although significant correlations were achieved, the discriminability of this model was only &#x0007E;70% accurate in the development cohorts, with ROSMAP again showing less accuracy than the other data sets (<xref ref-type="table" rid="T2">Table 2</xref>). Even though some of the LD clumps were missing, the validation cohorts demonstrated similar accuracies in discrimination as the development cohorts.</p>
<table-wrap position="float" id="T2">
<label>Table 2</label>
<caption><p>Results of the stringent PRS model in the development cohorts and validation cohorts, indicating a more uniform discriminability of the model consisting of SNPs tagging the 149 LD clumps identified to be contributing to the perfect discrimination (PD) model in 3 out of the 5 development cohorts.</p></caption> 
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919497;color:#ffffff">
<th valign="top" align="left"><bold>Development</bold></th>
<th valign="top" align="center"><bold>SNP platform</bold></th>
<th valign="top" align="center"><bold>&#x00023; LD clumps contributing to PD model</bold></th>
<th valign="top" align="center"><bold>Logistic regression <italic>P</italic>-value</bold></th>
<th valign="top" align="center"><bold>Area under the curve</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left">ADC7<break/> <italic>N =</italic> 1462</td>
<td valign="top" align="center">HumanOmni express</td>
<td valign="top" align="center">137</td>
<td valign="top" align="center">1.19 x 10<sup>&#x02212;44</sup></td>
<td valign="top" align="center">0.7273</td>
</tr> <tr>
<td valign="top" align="left">BDR<break/> <italic>N =</italic> 520</td>
<td valign="top" align="center">NeuroChip</td>
<td valign="top" align="center">83</td>
<td valign="top" align="center">6.89 x 10<sup>&#x02212;22</sup></td>
<td valign="top" align="center">0.7911</td>
</tr> <tr>
<td valign="top" align="left">NIA<break/> <italic>N =</italic> 1386</td>
<td valign="top" align="center">Human610-quad</td>
<td valign="top" align="center">107</td>
<td valign="top" align="center">4.65 x 10<sup>&#x02212;72</sup></td>
<td valign="top" align="center">0.8259</td>
</tr> <tr>
<td valign="top" align="left">ROSMAP<break/> <italic>N =</italic> 240</td>
<td valign="top" align="center">Affymetrix6.0/IlluminaOmni Quad</td>
<td valign="top" align="center">8</td>
<td valign="top" align="center">0.449</td>
<td valign="top" align="center">0.5477</td>
</tr> <tr>
<td valign="top" align="left">TGEN<break/> <italic>N =</italic> 1510</td>
<td valign="top" align="center">Affymetrix 6.0</td>
<td valign="top" align="center">116</td>
<td valign="top" align="center">1.43 x 10<sup>&#x02212;54</sup></td>
<td valign="top" align="center">0.7734</td>
</tr> 
<tr style="background-color:#919497;color:#ffffff">
<td valign="top" align="left"><bold>Validation</bold></td>
<td valign="top" align="center"><bold>SNP platform</bold></td>
<td valign="top" align="center"><bold>&#x00023; SNPs available</bold></td>
<td valign="top" align="center"><bold>Logistic regression</bold> <italic><bold>P</bold></italic><bold>-value</bold></td>
<td valign="top" align="center"><bold>Area under the curve</bold></td>
</tr> <tr>
<td valign="top" align="left">MTC<break/> <italic>N =</italic> 356</td>
<td valign="top" align="center">HumanOmni Express</td>
<td valign="top" align="center">ADC7 Panel<break/> 149 SNPs (1 alt)</td>
<td valign="top" align="center">1.28 x 10<sup>&#x02212;11</sup></td>
<td valign="top" align="center">0.7294</td>
</tr> <tr>
<td valign="top" align="left">WashU<break/> <italic>N =</italic> 131</td>
<td valign="top" align="center">HumanOmni Express</td>
<td valign="top" align="center">ADC7 Panel<break/> 127 SNPs (2 alt)</td>
<td valign="top" align="center">0.035</td>
<td valign="top" align="center">0.6150</td>
</tr>
<tr>
<td valign="top" align="left">TARCC<break/> <italic>N =</italic> 491</td>
<td valign="top" align="center">Affymetrix 6.0</td>
<td valign="top" align="center">TGEN Panel<break/> 143 SNPs (6 alt)</td>
<td valign="top" align="center">8.32 x 10<sup>&#x02212;13</sup></td>
<td valign="top" align="center">0.7006</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<p>In Panel A, the results from the development cohorts, indicate the number of LD clumps from that cohort that were found in at least 2 other cohorts, note that there is no correlation between the number of clumps identified in each cohort and the accuracy in the model. In Panel B, the results from the validation cohorts indicate the SNP panels used in the PRS and the number of alternative (alt) SNPs that had to be found to substitute for SNPs missing in that cohort.</p>
</table-wrap-foot>
</table-wrap>
<p>The PRSs for &#x0201C;absolute&#x0201D; controls and cases were calculated and used to create a centile plot of scores within each data set for each model (<xref ref-type="fig" rid="F1">Figures 1</xref>, <xref ref-type="fig" rid="F2">2</xref>).</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>Bar chart for each cohort showing the numbers of cases (black) and controls (light gray) in each centile of the 2207 lenient model showing two-distinct normal distribution curves for the PRS corresponding to the controls and cases for each dataset.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frdem-02-1120206-g0001.tif"/>
</fig>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Bar chart for each cohort showing the numbers of cases (black) and controls (light gray) in each centile of the 149 stringent model. In contrast to the lenient model with more LD blocks utilized, the 149 suggests three distinct normal distributions of scores; one at the lower end of the scale made up predominantly of controls, a middle range distribution of both cases and controls, and a high range peak made up predominantly of cases.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frdem-02-1120206-g0002.tif"/>
</fig>
<p>The lenient model of 2,207 SNPs demonstrated a narrow range of scores that fall within the 46th to 54th centiles. Furthermore, there was an overlap of scores for cases and controls in the middle ranges; however, in all data sets, we observed two distinct peaks of scores, one for controls (48th&#x02212;49th centile) and one for cases (50th&#x02212;51st centile), with a 2- to 3-percentile difference between them. The ROSMAP data set has the smallest sample size and had insufficient cases in the higher score ranges, leading to a more predictive accuracy value. The stringent model of 149 SNPs created a far more dynamic range of scores from the 34th to the 75th centile; with this model, an interesting pattern of three distinct normal distributions was observed: one in the lower range from approximately the 38th to the 50th percentile, which consisted of a higher number of control samples; one in the higher range from the 64th to the 74th centile, consisting almost entirely of AD cases; and a middle third peak ranging from the 52nd to the 62nd centile, where there was an overlap of controls and cases but generally was biased toward AD case individuals. This model also indicated that the ROSMAP sample lacks AD cases with risk scores in the higher ranges, accounting for the lack of discriminability.</p>
<p>The stringent model had the best coverage of LD blocks in the replication samples, and therefore, the predictability of AD cases being present in the highest ranges was explored. This range of centiles suggests that more than 90% of individuals with scores 0.006 (63rd centile) or above will have AD, and this was supported in two out of the three replications cohorts. From the MTC cohort, 96.6% of individuals who scored above this threshold were AD cases. Similarly, in the TARCC cohort, 89.6% were AD cases. The WashU cohort did not support this fact, for the majority of the samples were controls for this score range; however, note that the number of individuals in this range was approximately a third of those in the other two replication cohorts.</p>
<p>A full table of the 2,207- and 149-model SNPs/LD blocks and genes located in the vicinity can be found in <xref ref-type="supplementary-material" rid="SM1">Supplementary material 3</xref>, <xref ref-type="supplementary-material" rid="SM1">4</xref>. <xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref> shows a summary table of the cohorts used and the PRS models created in this study, in addition to the average AUC obtained across the five data sets and when the ROSMAP data set was removed from the mean.</p></sec>
<sec sec-type="discussion" id="s4">
<title>Discussion</title>
<p>This investigation explored an alternative method for developing a PRS model to predict AD, aiming to improve the translatability of the model to other cohorts and identify areas of the genome that may house actionable targets for therapeutic intervention. Instead of modeling the data on single SNPs that are common between the base and target data sets, it used SNPs that tag linkage disequilibrium blocks that are common between different cohorts to improve the coverage of the genome and transferability to other cohorts. As an alternative to modeling the SNP inclusion based on significance value in the base data set, SNP inclusion was based on over-fitting the data to create perfect discriminatory models for five cohorts and then identifying the linkage disequilibrium blocks that repeatedly contribute to the AUC = 1 across the development cohorts. Lenient and stringent models were applied to both the cohorts they were developed on and three additional independent cohorts to test replicability. In all cases, the average AUC achieved in the validation cohorts was lower than that in the development cohorts; however, the stringent model demonstrated an average decrease in AUC of 0.05 compared to an average decrease in AUC of 0.17 in the lenient model.</p>
<p>One of the main caveats for PRS analysis is the generalization of the models developed to additional cohorts. Some authors suggest that even more GWAS data than those used in this study are required to accurately calculate the effect sizes of risk alleles; conversely, the current GWAS results could be inflated, subject to the &#x0201C;winner&#x00027;s curse&#x0201D; (Shi et al., <xref ref-type="bibr" rid="B31">2016</xref>) and the admixture of samples from the pooling of multiple cohorts (Janssens, <xref ref-type="bibr" rid="B16">2019</xref>).</p>
<p>Preliminary unpublished explorations of AUC results on a single target dataset using a GWAS SNP model and effect sizes from different GWAS summary statistics have demonstrated that differences in the effect sizes of reference alleles do not alter the accuracy of discriminating cases from controls. Furthermore, the pattern of the AUCs obtained during the threshold modeling process in this study (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 2</xref>) would also support the fact that, when effect sizes and SNP inclusion are similar, different AUCs can be obtained, suggesting that it is the allele frequency of the SNP in the target data set that determines the discriminatory accuracy, which is similar to the nuisances of gene association studies and significance. Therefore, an alternative hypothesis would be to identify SNPs that consistently display allele frequency differences that contribute to a more diverse risk score between cases and controls for a generalized AD risk model.</p>
<p>This alternative method of incorporating SNPs individually into the model based on their associated <italic>p-</italic>value within the target data set and directional effect on the AUC provided a model that perfectly separated the cases and controls. Importantly, when this was applied across multiple data sets, the common SNPs/LD clumps that appeared in at least two datasets (lenient model) provided a model that offered better and consistent discriminability across data sets (mean AUC = 0.8493, <italic>SD</italic> = 0.17) compared to when the traditional method of selecting SNPs on significance threshold models was used (mean AUC = 0.7621, <italic>SD</italic> = 0.14). Applying the stringent model resulted in an average AUC in the development cohort that was slightly less (mean AUC = 0.7331, <italic>SD</italic> = 0.11) than that of the traditional method; however, while the traditional method utilized vastly different SNPs in each model, the stringent model utilized the same SNPs in each cohort.</p>
<p>The ROSMAP data set was the exception, performing not as well as the other development cohorts in both models (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>). The ROSMAP data set sample size was the smallest of the development cohorts due to selecting only samples that had a confirmed diagnosis of AD and removing those classified as possible/probable cases of AD. Importantly, the <italic>APOE</italic> isoform SNPs were not significantly associated with the AD phenotype in this cohort (<italic>p</italic> &#x0003E; .05) and did not demonstrate discriminability in the PRS consisting of only these two SNPs (<xref ref-type="supplementary-material" rid="SM1">Supplementary material 5</xref>). As <italic>APOE</italic> is such a strong predictor for an AD diagnosis, the absence of this effect may go some way in explaining the lack of discriminability for this PRS compared to the other cohorts when the rs429358 and rs7412 SNPs are included in the model.</p>
<p>As shown in <xref ref-type="supplementary-material" rid="SM1">Supplementary material 3, 4</xref>, the genes located in the vicinity of the SNPs/LD blocks used in the models are those familiar in AD genetics, including not only established GWAS-hits genes, such as <italic>APOE, TOMM40, CLU, EPHA-AS1, PICALM, CD2AP, SLC24A4</italic>, and <italic>MSA4A6A</italic>, but also lesser known gene associations from previous studies and those that overlap with linkage peaks from early genetics investigations (Brookes and Morgan, <xref ref-type="bibr" rid="B6">2017</xref>), such as <italic>PACRG</italic> (Sirkis et al., <xref ref-type="bibr" rid="B32">2016</xref>); <italic>LRAT</italic> (Abraham et al., <xref ref-type="bibr" rid="B1">2008</xref>), <italic>UNC5C</italic> (Jiao et al., <xref ref-type="bibr" rid="B17">2014</xref>; Wetzel-Smith et al., <xref ref-type="bibr" rid="B34">2014</xref>); <italic>AKAP6</italic> (Seshadri et al., <xref ref-type="bibr" rid="B30">2010</xref>); <italic>DLG2</italic> (Lawingco et al., <xref ref-type="bibr" rid="B20">2020</xref>; Prokopenko et al., <xref ref-type="bibr" rid="B24">2022</xref>); and <italic>DAB1</italic> and <italic>ARID1B</italic> (Harold et al., <xref ref-type="bibr" rid="B12">2009</xref>). It is also promising that some novel genes (<italic>UMAD1, ABCAC1, SNX1, APP</italic>) found in a recent GWAS study with a very large sample size were also identified in the 2207 model (Bellenguez et al., <xref ref-type="bibr" rid="B5">2022</xref>).</p>
<p>A large proportion of the SNPs in the model produced here (81% and 74% for lenient and stringent models, respectively) had IGAP_S1 <italic>p-</italic>values of above .05 and, therefore, may have been omitted in traditional thresholding analysis. Furthermore, the beta coefficient values of these SNPs are not minuscule (&#x0003E;0.0019 and &#x0003C; -0.0012), suggesting that they might have observable biological effects and could be therapeutic targets.</p>
<p>Being that the IGAP_S1 summary statistics used in this study were already imputed, the further imputation of the genetic data of the cohorts in this study would not be beneficial due to increased variability that was observed in imputation in other data sets (Chen et al., <xref ref-type="bibr" rid="B7">2020</xref>). As an alternative measure, this investigation employed clumping the base data set and selected SNPs within the target data set to capture these LD blocks, taking the Index (most significant SNP in the clumped base data set) when possible or an alternative SNP in the same clump if the Index SNP was not present. This allowed a greater coverage of the genome both within our initial analyses of the BDR cohort and the cross-cohort analyses without imputation. Interestingly, even though the validation cohorts were selected for genotyping panels to match those in the developmental cohorts, not all SNPs/LD clumps were present for analysis, possibly due to those SNPs failing to genotype in that cohort or perhaps being removed at quality control.</p>
<p>The additional target data sets utilized are a key limitation of this study as they are not independent from the base data sets. The IGAP study was a milestone collaborative investigation that agglomerated multiple AD cohorts from around the world into a single meta-analysis to elucidate our currently well-established AD candidate genes (Lambert et al., <xref ref-type="bibr" rid="B19">2013</xref>); however, this leaves few data sets that are completely independent of this work. A previous study (Escott-Price et al., <xref ref-type="bibr" rid="B8">2017</xref>) also utilized a sample cohort that was part of the IGAP study. In their analysis, they recalculated the prediction accuracy based on the SNPs from the independent data set and found that it produced a similar AUC. Therefore, given the relatively small sample size of each of the cohorts utilized in this study, it is likely that the AUCs presented are not greatly biased.</p>
<p>One caveat for the LD clump approach is that, due to some alleles having effect sizes in the opposite direction, 1,067 clumps were removed from the analysis as their inclusion may have produced variation at the individual level, altering the risk &#x0201C;centile&#x0201D; in which the individual may reside. The scoring algorithm for PRS normally concerns the beta value of only the reference/effect allele; however, to allow more parity, perhaps the effect sizes of both alleles could be incorporated into this calculation, especially as the opposing allele may not have a neutral effect as assumed in the current scoring algorithm in PLINK.</p>
<p>Compared to other PRS investigations in AD, this investigation opted for a higher <italic>r</italic><sup>2</sup> value of 0.8 for clumping and included genetic variation surrounding the <italic>APOE</italic> gene. It is possible that the additional SNPs are tagging the effects of the <italic>APOE</italic> isoform SNPs, leading to inflated AUC values; however, what has been presented in this article is consistent with previous studies (Harrison et al., <xref ref-type="bibr" rid="B13">2020</xref>; Leonenko et al., <xref ref-type="bibr" rid="B21">2021</xref>). Regarding the <italic>APOE</italic> region, the LD between SNPs is generally low, with only moderate (<italic>r</italic><sup>2</sup> &#x0003C; 0.56) LD observed between SNPs in this region and the rs429358 and rs7412 isoform SNPs (Young et al., <xref ref-type="bibr" rid="B35">2021</xref>). Whilst the independent association with and contribution of several SNPs to AD within this area are supported by multiple studies (Huentelman et al., <xref ref-type="bibr" rid="B15">2010</xref>; Rao et al., <xref ref-type="bibr" rid="B28">2018</xref>; Zhou et al., <xref ref-type="bibr" rid="B36">2019</xref>), the extent to which the isoform SNPs are influencing these observations is still an area for investigation.</p>
<p>Furthermore, this study did not include environmental predictors for AD. The inclusion of &#x0201C;non-genetic&#x0201D; variables into PRS analyses is considered by some to be problematic, as they may not be entirely independent of genetic influences already being entered into the model. Quantitative genetics studies suggest that, due to genetic&#x02013;environment correlations, &#x0007E;25% of environmental measures are heritable (Kendler and Baker, <xref ref-type="bibr" rid="B18">2007</xref>). Female sex is viewed as a risk factor for AD, and this is genetically controlled. Therefore, to incorporate being female as a risk factor into the analysis may have increased the prediction accuracy. Similarly, longevity has also been found to have a genetic component (Herskind et al., <xref ref-type="bibr" rid="B14">1996</xref>), and thus, age at death was also omitted from the model.</p>
<p>In conclusion, this study aimed to provide an alternative to current PRS methods for developing a model that can be applied across multiple cohorts genotyped on various platforms using informative SNPs common across data sets. This is still a hypothetical model and requires further exploration, refinement, and testing on additional data. The authors invite other fellow researchers to test this model in their own cohorts, as additional data sets will help identify additional common SNPs and refine the model by removing false positives. When SNP or LD block consistency is achieved across studies, the biological consequences may then be identified and lead to the fulfillment of the clinical utility the PRS aims to bring.</p></sec>
<sec sec-type="data-availability" id="s5">
<title>Data availability statement</title>
<p>The Religious Orders Study and Memory and Aging Project, <underline><bold>ROSMAP</bold></underline>, Study genotype data were obtained from the AMP-AD Knowledge Portal via the Synapse Data Access System (SYN3219045, <ext-link ext-link-type="uri" xlink:href="https://www.synapse.org/">https://www.synapse.org/</ext-link>), with details about this cohort found in De Jager, P., Ma, Y., McCabe, C. et al. A multi-omic atlas of the human frontal cortex for aging and Alzheimer&#x00027;s disease research. Sci Data 5, 180142 (2018). <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/sdata.2018.142">https://doi.org/10.1038/sdata.2018.142</ext-link>. The genotypes for the Brains for Dementia Research <underline><bold>BDR</bold></underline> Cohort datasets are freely available via the Dementia Platform UK (<ext-link ext-link-type="uri" xlink:href="https://www.dementiasplatform.uk/">https://www.dementiasplatform.uk/</ext-link>), with details about this cohort described in Young J, Gallagher E, Koska K, Guetta-Baranes T, Morgan K, Thomas A, Brookes KJ. Genome-wide association findings from the brains for dementia research cohort. Neurobiol Aging. 2021 Nov; 107:159&#x02013;167. <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1016/j.neurobiolaging.2021.05.014">10.1016/j.neurobiolaging.2021.05.014</ext-link>. The Alzheimer&#x00027;s Disease Centres seventh set of ADC genotyped subjects used by the Alzheimer&#x00027;s Disease Genetics Consortium (ADGC) <underline><bold>ADC7</bold></underline> (NG00071), the National Institute on Aging Genetics Initiative for Late-Onset Alzheimer&#x00027;s Disease (NIA-LOAD), <underline><bold>NIA</bold></underline> (NG00020), <underline><bold>TGen II</bold></underline>, TGEN (NG00028), samples collected by the Knight ADRC at Washington University, <underline><bold>WashU</bold></underline> (NG00097), The Texas Alzheimer&#x00027;s Research Care Consortium, <underline><bold>TARCC</bold></underline> (NG00097) and University of Miami/ Texas Alzheimer&#x00027;s Research Care Consortium Wave 2/Case Western Reserve University, <underline><bold>MTC</bold></underline> (NG00096) datasets were downloaded from NIAGADS (<ext-link ext-link-type="uri" xlink:href="https://www.niagads.org/">https://www.niagads.org/</ext-link>). Details of the ADC7 and WashU cohort were first published in Kunkle, B.W., Grenier-Boley, B., Sims, R. et al. Genetic meta-analysis of diagnosed Alzheimer&#x00027;s disease identifies new risk loci and implicates A&#x003B2;, tau, immunity and lipid processing. Nat Genet 51, 414&#x02013;430 (2019). <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.1038/s41588-019-0358-2">https://doi.org/10.1038/s41588-019-0358-2</ext-link>.</p></sec>
<sec sec-type="ethics-statement" id="s6">
<title>Ethics statement</title>
<p>The Brains for Dementia Research cohort use for genetic analysis was reviewed and approved by the London-City and East NRES committee 08/H0704/128&#x0002B;5. The patients/participants provided their written informed consent to participate in this study.</p></sec>
<sec sec-type="author-contributions" id="s7">
<title>Author contributions</title>
<p>KB: study conception and design, data analysis and interpretation, manuscript preparation, and funding acquisition. TG-B: laboratory work. AT: provision of resources. KM: funding acquisition and manuscript review and editing. All authors contributed to the article and approved the submitted version.</p></sec>
</body>
<back>
<sec sec-type="funding-information" id="s8">
<title>Funding</title>
<p>The work presented here was supported by funding provided by an ARUK project grant, entitled &#x02018;Enabling high-throughput genomic approaches in Alzheimer&#x00027;s disease&#x00027; awarded to KM, and an ARUK extension grant entitled &#x02018;NeuroChip analysis of the entire Brains for Dementia Research (BDR) resource of 2,000 samples,&#x00027; awarded to KM and KB.</p>
</sec>
<ack><p>We would like to gratefully acknowledge all donors and their families for the samples provided for the BDR cohort and additional datasets who genetic data was also utilized here. Tissue samples from the BDR cohort were obtained from the Southwest Dementia Brain Bank, London Neurodegenerative Diseases Brain Bank, Manchester Brain Bank, Newcastle Brain Tissue Resource and Oxford Brain Bank, and we thank our colleagues of the BDR Network, in particular the neuropathologists at each center and BDR Brain Bank staff for the collection and classification of the samples. The BDR is jointly funded by Alzheimer&#x00027;s Research UK and the Alzheimer&#x00027;s Society in association with the Medical Research Council. Brains for Dementia Research has ethics approval from London-City and East NRES committee 08/H0704/128&#x0002B;5 and has deemed all approved requests for tissue to have been approved by the committee.</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
</sec>
<sec sec-type="disclaimer" id="s9">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s10">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frdem.2023.1120206/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frdem.2023.1120206/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.xlsx" id="SM1" mimetype="application/vnd.openxmlformats-officedocument.spreadsheetml.sheet" xmlns:xlink="http://www.w3.org/1999/xlink"/></sec>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Abraham</surname> <given-names>R.</given-names></name> <name><surname>Moskvina</surname> <given-names>V.</given-names></name> <name><surname>Sims</surname> <given-names>R.</given-names></name> <name><surname>Hollingworth</surname> <given-names>P.</given-names></name> <name><surname>Morgan</surname> <given-names>A.</given-names></name> <name><surname>Georgieva</surname> <given-names>L.</given-names></name> <etal/></person-group>. (<year>2008</year>). <article-title>A genome-wide association study for late-onset Alzheimer&#x00027;s disease using DNA pooling</article-title>. <source>BMC Med Genom</source>. <volume>1</volume>, <fpage>44</fpage>. <pub-id pub-id-type="doi">10.1186/1755-8794-1-44</pub-id><pub-id pub-id-type="pmid">18823527</pub-id></citation></ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Aleknonyte-Resch</surname> <given-names>M.</given-names></name> <name><surname>Szymczak</surname> <given-names>S.</given-names></name> <name><surname>Freitag-Wolf</surname> <given-names>S.</given-names></name> <name><surname>Dempfle</surname> <given-names>A.</given-names></name> <name><surname>Krawczak</surname> <given-names>M.</given-names></name></person-group> (<year>2021</year>). <article-title>Genotype imputation in case-only studies of gene-environment interaction: validity and power</article-title>. <source>Human Genet</source>. <volume>140</volume>, <fpage>1217</fpage>&#x02013;<lpage>1228</lpage>. <pub-id pub-id-type="doi">10.1007/s00439-021-02294-z</pub-id><pub-id pub-id-type="pmid">34041609</pub-id></citation></ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ali</surname> <given-names>A. T.</given-names></name> <name><surname>Liebert</surname> <given-names>A.</given-names></name> <name><surname>Lau</surname> <given-names>W.</given-names></name> <name><surname>Maniatis</surname> <given-names>N.</given-names></name> <name><surname>Swallow</surname> <given-names>D. M.</given-names></name></person-group> (<year>2022</year>). <article-title>The hazards of genotype imputation in chromosomal regions under selection: a case study using the Lactase gene region</article-title>. <source>Ann. Human Gene.</source> <volume>86</volume>, <fpage>24</fpage>&#x02013;<lpage>33</lpage>. <pub-id pub-id-type="doi">10.1111/ahg.12444</pub-id><pub-id pub-id-type="pmid">34523124</pub-id></citation></ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Baker</surname> <given-names>E.</given-names></name> <name><surname>Escott-Price</surname> <given-names>V.</given-names></name></person-group> (<year>2020</year>). <article-title>Polygenic risk scores in alzheimer&#x00027;s disease: current applications and future directions</article-title>. <source>Front. Dig. Health</source> <volume>2</volume>, <fpage>14</fpage>. <pub-id pub-id-type="doi">10.3389/fdgth.2020.00014</pub-id><pub-id pub-id-type="pmid">34713027</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bellenguez</surname> <given-names>C.</given-names></name> <name><surname>K&#x000FC;&#x000E7;&#x000FC;kali</surname> <given-names>F.</given-names></name> <name><surname>Jansen</surname> <given-names>I. E.</given-names></name> <name><surname>Kleineidam</surname> <given-names>L.</given-names></name> <name><surname>Moreno-Grau</surname> <given-names>S.</given-names></name> <name><surname>Amin</surname> <given-names>N.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>New insights into the genetic etiology of Alzheimer&#x00027;s disease and related dementias</article-title>. <source>Nat. Genetics</source>. <volume>4</volume>, <fpage>24</fpage>. <pub-id pub-id-type="doi">10.1038./s41588-022-01024-z</pub-id><pub-id pub-id-type="pmid">35379992</pub-id></citation></ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Brookes</surname> <given-names>K. J.</given-names></name> <name><surname>Morgan</surname> <given-names>K.</given-names></name></person-group> (<year>2017</year>). <article-title>Genetics of Alzheimer&#x00027;s disease</article-title>. <source>ELS</source> <volume>5</volume>, <fpage>228</fpage>. <pub-id pub-id-type="doi">10.1002./9780470015902.a0020228.pub2</pub-id></citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Chen</surname> <given-names>S-. F.</given-names></name> <name><surname>Dias</surname> <given-names>R.</given-names></name> <name><surname>Evans</surname> <given-names>D.</given-names></name> <name><surname>Salfati</surname> <given-names>E. L.</given-names></name> <name><surname>Liu</surname> <given-names>S.</given-names></name> <name><surname>Wineinger</surname> <given-names>N. E.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Genotype imputation and variability in polygenic risk score estimation</article-title>. <source>Genome Med.</source> <volume>12</volume>, <fpage>100</fpage>. <pub-id pub-id-type="doi">10.1186/s13073-020-00801-x</pub-id><pub-id pub-id-type="pmid">33225976</pub-id></citation></ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Escott-Price</surname> <given-names>V.</given-names></name> <name><surname>Myers</surname> <given-names>A. J.</given-names></name> <name><surname>Huentelman</surname> <given-names>M.</given-names></name> <name><surname>Hardy</surname> <given-names>J.</given-names></name></person-group> (<year>2017</year>). <article-title>Polygenic risk score analysis of pathologically confirmed Alzheimer&#x00027;s disease</article-title>. <source>Ann Neurol</source>. <volume>82</volume>, <fpage>311</fpage>&#x02013;<lpage>314</lpage>. <pub-id pub-id-type="doi">10.1002./ana.24999</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Escott-Price</surname> <given-names>V.</given-names></name> <name><surname>Sims</surname> <given-names>R.</given-names></name> <name><surname>Bannister</surname> <given-names>C.</given-names></name> <name><surname>Harold</surname> <given-names>D.</given-names></name> <name><surname>Vronskaya</surname> <given-names>M.</given-names></name> <name><surname>Majounie</surname> <given-names>E.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Common polygenic variation enhances risk prediction for Alzheimer&#x00027;s disease</article-title>. <source>Brain</source> <volume>138</volume>, <fpage>3673</fpage>&#x02013;<lpage>3684</lpage>. <pub-id pub-id-type="doi">10.1093/brain/awv268</pub-id><pub-id pub-id-type="pmid">26490334</pub-id></citation></ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Euesden</surname> <given-names>J.</given-names></name> <name><surname>Lewis</surname> <given-names>C. M.</given-names></name> <name><surname>O&#x00027;Reilly</surname> <given-names>P.F.</given-names></name></person-group> (<year>2015</year>). <article-title>PRSice: polygenic risk score software</article-title>. <source>Bioinformatics</source> <volume>31</volume>, <fpage>1466</fpage>&#x02013;<lpage>1468</lpage>. <pub-id pub-id-type="doi">10.1093/bioinformatics/btu848</pub-id><pub-id pub-id-type="pmid">31307061</pub-id></citation></ref>
<ref id="B11">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Farrell</surname> <given-names>C.</given-names></name> <name><surname>Brookes</surname> <given-names>K.J.</given-names></name></person-group> (<year>2022</year>). <article-title>Utilising polygenic risk score analysis for AD to determine the &#x02018;sphere of influence&#x00027; of the APOE isoform SNPs</article-title>. <source>J. Neurol. Neuromed.</source> <volume>6</volume>, <fpage>1</fpage>&#x02013;<lpage>7</lpage>. <pub-id pub-id-type="doi">10.29245/2572.942X/2022/2.1284</pub-id></citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harold</surname> <given-names>D.</given-names></name> <name><surname>Abraham</surname> <given-names>R.</given-names></name> <name><surname>Hollingworth</surname> <given-names>P.</given-names></name> <name><surname>Sims</surname> <given-names>R.</given-names></name> <name><surname>Gerrish</surname> <given-names>A.</given-names></name> <name><surname>Hamshere</surname> <given-names>M. L.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Genome-wide association study identifies variants at CLU and PICALM associated with Alzheimer&#x00027;s disease</article-title>. <source>Nat Genet</source>. <volume>41</volume>, <fpage>1088</fpage>&#x02013;<lpage>1093</lpage>. <pub-id pub-id-type="doi">10.1038/ng.440ng.440</pub-id> [pii]<pub-id pub-id-type="pmid">19734902</pub-id></citation></ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Harrison</surname> <given-names>J. R. J. R.</given-names></name> <name><surname>Mistry</surname> <given-names>S.</given-names></name> <name><surname>Muskett</surname> <given-names>N.</given-names></name> <name><surname>Escott-Price</surname> <given-names>V.</given-names></name> <name><surname>Brookes</surname> <given-names>K.</given-names></name></person-group> (<year>2020</year>). <article-title>From polygenic scores to precision medicine in Alzheimer&#x00027;s disease: a systematic review</article-title>. <source>J. Alzheimer&#x00027;s Dis.</source> 74 1271&#x02013;1283. <pub-id pub-id-type="doi">10.3233/JAD-191233</pub-id><pub-id pub-id-type="pmid">32250305</pub-id></citation></ref>
<ref id="B14">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Herskind</surname> <given-names>A. M.</given-names></name> <name><surname>McGue</surname> <given-names>M.</given-names></name> <name><surname>Holm</surname> <given-names>N. V.</given-names></name> <name><surname>S&#x000F8;rensen</surname> <given-names>T. I.</given-names></name> <name><surname>Harvald</surname> <given-names>B.</given-names></name> <name><surname>Vaupel</surname> <given-names>J.W.</given-names></name></person-group> (<year>1996</year>). <article-title>The heritability of human longevity: a population-based study of 2,872 Danish twin pairs born 1870&#x02013;1900</article-title>. <source>Human Gen</source>. 319&#x02013;323. <pub-id pub-id-type="doi">10.1007/BF02185763</pub-id><pub-id pub-id-type="pmid">8786073</pub-id></citation></ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Huentelman</surname> <given-names>M.</given-names></name> <name><surname>Corneveaux</surname> <given-names>J.</given-names></name> <name><surname>Myers</surname> <given-names>A.</given-names></name> <name><surname>Allen</surname> <given-names>A.</given-names></name> <name><surname>Pruzin</surname> <given-names>J.</given-names></name> <name><surname>Nalls</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>S4-03-02: genome-wide association study for Alzheimer&#x00027;s disease risk in a large cohort of clinically characterized and neuropathologically verified subjects</article-title>. <source>Alzheimer&#x00027;s and Dem</source>. <volume>6</volume>, <fpage>e13</fpage>&#x02013;<lpage>e13</lpage>. <pub-id pub-id-type="doi">10.1016/j.jalz.08041</pub-id></citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Janssens</surname> <given-names>A. C. J. W.</given-names></name></person-group> (<year>2019</year>). <article-title>Validity of polygenic risk scores: are we measuring what we think we are?</article-title> <source>Human Mol. Gen.</source> <volume>28</volume>, <fpage>205</fpage>. <pub-id pub-id-type="doi">10.1093./hmg/ddz205</pub-id><pub-id pub-id-type="pmid">31504522</pub-id></citation></ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Jiao</surname> <given-names>B.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Tang</surname> <given-names>B.</given-names></name> <name><surname>Hou</surname> <given-names>L.</given-names></name> <name><surname>Zhou</surname> <given-names>L.</given-names></name> <name><surname>Zhang</surname> <given-names>F.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>Investigation of TREM2, PLD3, and UNC5C variants in patients with Alzheimer&#x00027;s disease from mainland China</article-title>. <source>Neurobiol. Aging</source> <volume>35</volume>, <fpage>2422</fpage>. <pub-id pub-id-type="doi">10.1016/j.neurobiolaging.04025</pub-id><pub-id pub-id-type="pmid">24866402</pub-id></citation></ref>
<ref id="B18">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kendler</surname> <given-names>K. S.</given-names></name> <name><surname>Baker</surname> <given-names>J. H.</given-names></name></person-group> (<year>2007</year>). <article-title>Genetic influences on measures of the environment: a systematic review</article-title>. <source>Psychol. Med.</source> <volume>37</volume>, <fpage>615</fpage>&#x02013;<lpage>626</lpage>. <pub-id pub-id-type="doi">10.1017/S0033291706009524</pub-id><pub-id pub-id-type="pmid">17176502</pub-id></citation></ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lambert</surname> <given-names>J-. C.</given-names></name> <name><surname>Ibrahim-Verbaas</surname> <given-names>C. A.</given-names></name> <name><surname>Harold</surname> <given-names>D.</given-names></name> <name><surname>Naj</surname> <given-names>A. C.</given-names></name> <name><surname>Sims</surname> <given-names>R.</given-names></name> <name><surname>Bellenguez</surname> <given-names>C.</given-names></name> <etal/></person-group>. (<year>2013</year>). <article-title>Meta-analysis of 74,046 individuals identifies 11 new susceptibility loci for Alzheimer&#x00027;s disease</article-title>. <source>Nat. Gen.</source> <volume>45</volume>, <fpage>1452</fpage>&#x02013;<lpage>1458</lpage>. <pub-id pub-id-type="doi">10.1038/ng.2802</pub-id><pub-id pub-id-type="pmid">24162737</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lawingco</surname> <given-names>T.</given-names></name> <name><surname>Chaudhury</surname> <given-names>S.</given-names></name> <name><surname>Brookes</surname> <given-names>K. J.</given-names></name> <name><surname>Guetta-Baranes</surname> <given-names>T.</given-names></name> <name><surname>Guerreiro</surname> <given-names>R.</given-names></name> <name><surname>Bras</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2020</year>). <article-title>Genetic variants in glutamate-, A&#x003B2;-, and tau-related pathways determine polygenic risk for Alzheimer&#x00027;s disease</article-title>. <source>Neurobiol. Aging</source> <volume>4</volume>, <fpage>9</fpage>. <pub-id pub-id-type="doi">10.1016/j.neurobiolaging.11009</pub-id><pub-id pub-id-type="pmid">33303219</pub-id></citation></ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Leonenko</surname> <given-names>G.</given-names></name> <name><surname>Baker</surname> <given-names>E.</given-names></name> <name><surname>Stevenson-Hoare</surname> <given-names>J.</given-names></name> <name><surname>Sierksma</surname> <given-names>A.</given-names></name> <name><surname>Fiers</surname> <given-names>M.</given-names></name> <name><surname>Williams</surname> <given-names>J.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Identifying individuals with high risk of Alzheimer&#x00027;s disease using polygenic risk scores</article-title>. <source>Nature Commun.</source> <volume>12</volume>, <fpage>4506</fpage>. <pub-id pub-id-type="doi">10.1038/s41467-021-24082-z</pub-id><pub-id pub-id-type="pmid">34301930</pub-id></citation></ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Nakaoka</surname> <given-names>H.</given-names></name> <name><surname>Inoue</surname> <given-names>I.</given-names></name></person-group> (<year>2009</year>). <article-title>Meta-analysis of genetic association studies: methodologies, between-study heterogeneity and winner&#x00027;s curse</article-title>. <source>J. Human Gen.</source> <volume>54</volume>, <fpage>615</fpage>&#x02013;<lpage>623</lpage>. <pub-id pub-id-type="doi">10.1038/jhg.2009.95</pub-id><pub-id pub-id-type="pmid">19851339</pub-id></citation></ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Plomin</surname> <given-names>R.</given-names></name> <name><surname>von Stumm</surname> <given-names>S.</given-names></name></person-group> (<year>2022</year>). <article-title>Polygenic scores: prediction vs. explanation</article-title>. <source>Mol. Psychiatry</source> <volume>27</volume>, <fpage>49</fpage>&#x02013;<lpage>52</lpage>. <pub-id pub-id-type="doi">10.1038/s41380-021-01348-y</pub-id><pub-id pub-id-type="pmid">34686768</pub-id></citation></ref>
<ref id="B24">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Prokopenko</surname> <given-names>D.</given-names></name> <name><surname>Lee</surname> <given-names>S.</given-names></name> <name><surname>Hecker</surname> <given-names>J.</given-names></name> <name><surname>Mullin</surname> <given-names>K.</given-names></name> <name><surname>Morgan</surname> <given-names>S.</given-names></name> <name><surname>Katsumata</surname> <given-names>Y.</given-names></name> <etal/></person-group>. (<year>2022</year>). <article-title>Region-based analysis of rare genomic variants in whole-genome sequencing datasets reveal two novel Alzheimer&#x00027;s disease-associated genes: DTNB and DLG2</article-title>. <source>Mol. Psychiatry</source>. <pub-id pub-id-type="doi">10.1038./s41380-022-01475-0</pub-id><pub-id pub-id-type="pmid">35246634</pub-id></citation></ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Purcell</surname> <given-names>S.</given-names></name> <name><surname>Neale</surname> <given-names>B.</given-names></name> <name><surname>Todd-Brown</surname> <given-names>K.</given-names></name> <name><surname>Thomas</surname> <given-names>L.</given-names></name> <name><surname>Ferreira</surname> <given-names>M. A.</given-names></name> <name><surname>Bender</surname> <given-names>D.</given-names></name> <etal/></person-group>. (<year>2007</year>). <article-title>PLINK: a tool set for whole-genome association and population-based linkage analyses</article-title>. <source>Am. J. Hum. Genet.</source> <volume>81</volume>, <fpage>559</fpage>&#x02013;<lpage>575</lpage>. <pub-id pub-id-type="doi">10.1086/519795</pub-id><pub-id pub-id-type="pmid">17701901</pub-id></citation></ref>
<ref id="B26">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Purcell</surname> <given-names>S. M.</given-names></name> <name><surname>Wray</surname> <given-names>N. R.</given-names></name> <name><surname>Stone</surname> <given-names>J. L.</given-names></name> <name><surname>Visscher</surname> <given-names>P. M.</given-names></name> <name><surname>O&#x00027;Donovan</surname> <given-names>M. C.</given-names></name> <name><surname>Sullivan</surname> <given-names>P. F.</given-names></name> <etal/></person-group>. (<year>2009</year>). <article-title>Common polygenic variation contributes to risk of schizophrenia and bipolar disorder</article-title>. <source>Nature</source> <volume>460</volume>, <fpage>748</fpage>&#x02013;<lpage>752</lpage>. <pub-id pub-id-type="doi">10.1038/nature08185</pub-id><pub-id pub-id-type="pmid">19571811</pub-id></citation></ref>
<ref id="B27">
<citation citation-type="book"><person-group person-group-type="author"><collab>R Core team</collab></person-group> (<year>2021</year>). <source>R: A Language and Environment for Statistical Computing</source>. <publisher-loc>Vienna</publisher-loc>: <publisher-name>R Foundation for Statistical Computing</publisher-name>.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Rao</surname> <given-names>S.</given-names></name> <name><surname>Ghani</surname> <given-names>M.</given-names></name> <name><surname>Guo</surname> <given-names>Z.</given-names></name> <name><surname>Deming</surname> <given-names>Y.</given-names></name> <name><surname>Wang</surname> <given-names>K.</given-names></name> <name><surname>Sims</surname> <given-names>R.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>An APOE-independent cis-eSNP on chromosome 19q13.32 influences tau levels and late-onset Alzheimer&#x00027;s disease risk</article-title>. <source>Neurobiol. Aging</source> <volume>66</volume>, <fpage>178</fpage>. <pub-id pub-id-type="doi">10.1016/j.neurobiolaging.12027</pub-id><pub-id pub-id-type="pmid">29395286</pub-id></citation></ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Robin</surname> <given-names>X.</given-names></name> <name><surname>Turck</surname> <given-names>N.</given-names></name> <name><surname>Hainard</surname> <given-names>A.</given-names></name> <name><surname>Tiberti</surname> <given-names>N.</given-names></name> <name><surname>Lisacek</surname> <given-names>F.</given-names></name> <name><surname>Sanchez</surname> <given-names>J. C.</given-names></name> <etal/></person-group>. (<year>2011</year>). <article-title>pROC: an open-source package for R and S&#x0002B; to analyze and compare ROC curves</article-title>. <source>BMC Bioinform.</source> <volume>12</volume>, <fpage>77</fpage>. <pub-id pub-id-type="doi">10.1186/1471-2105-12-77</pub-id><pub-id pub-id-type="pmid">21414208</pub-id></citation></ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Seshadri</surname> <given-names>S.</given-names></name> <name><surname>Fitzpatrick</surname> <given-names>A. L.</given-names></name> <name><surname>Ikram</surname> <given-names>M. A.</given-names></name> <name><surname>DeStefano</surname> <given-names>A. L.</given-names></name> <name><surname>Gudnason</surname> <given-names>V.</given-names></name> <name><surname>Boada</surname> <given-names>M.</given-names></name> <etal/></person-group>. (<year>2010</year>). <article-title>Genome-wide analysis of genetic loci associated with Alzheimer disease</article-title>. <source>JAMA</source> <volume>303</volume>, <fpage>1832</fpage>&#x02013;<lpage>1840</lpage>. <pub-id pub-id-type="doi">10.1001/jama.2010.574</pub-id><pub-id pub-id-type="pmid">20460622</pub-id></citation></ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Shi</surname> <given-names>J.</given-names></name> <name><surname>Park</surname> <given-names>J. H.</given-names></name> <name><surname>Duan</surname> <given-names>J.</given-names></name> <name><surname>Berndt</surname> <given-names>S. T.</given-names></name> <name><surname>Moy</surname> <given-names>W.</given-names></name> <name><surname>Yu</surname> <given-names>K.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Winner&#x00027;s curse correction and variable thresholding improve performance of polygenic risk modeling based on genome-wide association study summary-level data</article-title>. <source>PLOS Gen.</source> <volume>12</volume>, <fpage>e1006493</fpage>. <pub-id pub-id-type="doi">10.1371/journal.pgen.1006493</pub-id><pub-id pub-id-type="pmid">28036406</pub-id></citation></ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sirkis</surname> <given-names>D. W.</given-names></name> <name><surname>Bonham</surname> <given-names>L. W.</given-names></name> <name><surname>Aparicio</surname> <given-names>R. E.</given-names></name> <name><surname>Geier</surname> <given-names>E. G.</given-names></name> <name><surname>Ramos</surname> <given-names>E. M.</given-names></name> <name><surname>Wang</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2016</year>). <article-title>Rare TREM2 variants associated with Alzheimer&#x00027;s disease display reduced cell surface expression</article-title>. <source>Acta Neuropathol. Commun</source>. 4, 98. <pub-id pub-id-type="doi">10.1186/s40478-016-0367-7</pub-id><pub-id pub-id-type="pmid">27589997</pub-id></citation></ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Vilhj&#x000E1;lmsson</surname> <given-names>B. J.</given-names></name> <name><surname>Yang</surname> <given-names>J.</given-names></name> <name><surname>Finucane</surname> <given-names>H. K.</given-names></name> <name><surname>Gusev</surname> <given-names>A.</given-names></name> <name><surname>Lindstr&#x000F6;m</surname> <given-names>S.</given-names></name> <name><surname>Ripke</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2015</year>). <article-title>Modeling linkage disequilibrium increases accuracy of polygenic risk scores</article-title>. <source>Am. J. Human Gen.</source> <volume>97</volume>, <fpage>576</fpage>&#x02013;<lpage>592</lpage>. <pub-id pub-id-type="doi">10.1016/j.ajhg.2015.09.001</pub-id><pub-id pub-id-type="pmid">26430803</pub-id></citation></ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wetzel-Smith</surname> <given-names>M. K.</given-names></name> <name><surname>Hunkapiller</surname> <given-names>J.</given-names></name> <name><surname>Bhangale</surname> <given-names>T. R.</given-names></name> <name><surname>Srinivasan</surname> <given-names>K.</given-names></name> <name><surname>Maloney</surname> <given-names>J. A.</given-names></name> <name><surname>Atwal</surname> <given-names>J. K.</given-names></name> <etal/></person-group>. (<year>2014</year>). <article-title>A rare mutation in UNC5C predisposes to late-onset Alzheimer&#x00027;s disease and increases neuronal cell death</article-title>. <source>Nat. Med</source>. <volume>20</volume>, <fpage>1452</fpage>&#x02013;<lpage>1457</lpage>. <pub-id pub-id-type="doi">10.1038/nm.3736</pub-id><pub-id pub-id-type="pmid">25419706</pub-id></citation></ref>
<ref id="B35">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Young</surname> <given-names>J.</given-names></name> <name><surname>Gallagher</surname> <given-names>E.</given-names></name> <name><surname>Koska</surname> <given-names>K.</given-names></name> <name><surname>Guetta-Baranes</surname> <given-names>T.</given-names></name> <name><surname>Morgan</surname> <given-names>K.</given-names></name> <name><surname>Thomas</surname> <given-names>A.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Genome-wide association findings from the brains for dementia research cohort</article-title>. <source>Neurobiol. Aging</source> <volume>107</volume>, <fpage>159</fpage>&#x02013;<lpage>167</lpage>. <pub-id pub-id-type="doi">10.1016/J.NEUROBIOLAGING.05014</pub-id><pub-id pub-id-type="pmid">34183186</pub-id></citation></ref>
<ref id="B36">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhou</surname> <given-names>X.</given-names></name> <name><surname>Chen</surname> <given-names>Y.</given-names></name> <name><surname>Mok</surname> <given-names>K. Y.</given-names></name> <name><surname>Kwok</surname> <given-names>T. C. Y.</given-names></name> <name><surname>Mok</surname> <given-names>V. C. T.</given-names></name> <name><surname>Guo</surname> <given-names>Q.</given-names></name> <etal/></person-group>. (<year>2019</year>). <article-title>Non-coding variability at the APOE locus contributes to the Alzheimer&#x00027;s risk</article-title>. <source>Nat. Commun.</source> <volume>10</volume>, <fpage>1</fpage>&#x02013;<lpage>16</lpage>. <pub-id pub-id-type="doi">10.1038/s41467-019-10945-z</pub-id><pub-id pub-id-type="pmid">31346172</pub-id></citation></ref>
</ref-list> 
</back>
</article> 