Great efforts are being made to develop advanced polygenic risk scores (PRS) to improve the prediction of complex traits and diseases. However, most existing PRS are primarily trained on European ancestry populations, limiting their transferability to non-European populations. In this article, we propose a novel method for generating multi-ancestry Polygenic Risk scOres based on enSemble of PEnalized Regression models (PROSPER). PROSPER integrates genome-wide association studies (GWAS) summary statistics from diverse populations to develop ancestry-specific PRS with improved predictive power for minority populations. The method uses a combination of L 1 (lasso) and L 2 (ridge) penalty functions, a parsimonious specification of the penalty parameters across populations, and an ensemble step to combine PRS generated across different penalty parameters. We evaluate the performance of PROSPER and other existing methods on large-scale simulated and real datasets, including those from 23andMe Inc., the Global Lipids Genetics Consortium, and All of Us. Results show that PROSPER can substantially improve multi-ancestry polygenic prediction compared to alternative methods across a wide variety of genetic architectures. In real data analyses, for example, PROSPER increased out-of-sample prediction R 2 for continuous traits by an average of 70% compared to a state-of-the-art Bayesian method (PRS-CSx) in the African ancestry population. Further, PROSPER is computationally highly scalable for the analysis of large SNP contents and many diverse populations.
Polygenic risk scores (PRSs) are now showing promising predictive performance on a wide variety of complex traits and diseases, but there exists a substantial performance gap across populations. We propose MUSSEL, a method for ancestry-specific polygenic prediction that borrows information in summary statistics from genome-wide association studies (GWASs) across multiple ancestry groups via Bayesian hierarchical modeling and ensemble learning. In our simulation studies and data analyses across four distinct studies, totaling 5.7 million participants with a substantial ancestral diversity, MUSSEL shows promising performance compared to alternatives. For example, MUSSEL has an average gain in prediction R2 across 11 continuous traits of 40.2% and 49.3% compared to PRS-CSx and CT-SLEB, respectively, in the African ancestry population. The best-performing method, however, varies by GWAS sample size, target ancestry, trait architecture, and linkage disequilibrium reference samples; thus, ultimately a combination of methods may be needed to generate the most robust PRSs across diverse populations.
Polygenic risk scores (PRS) increasingly predict complex traits, however, suboptimal performance in non-European populations raise concerns about clinical applications and health inequities. We developed CT-SLEB, a powerful and scalable method to calculate PRS using ancestry-specific GWAS summary statistics from multi-ancestry training samples, integrating clumping and thresholding, empirical Bayes and super learning. We evaluate CT-SLEB and nine-alternatives methods with large-scale simulated GWAS (∼19 million common variants) and datasets from 23andMe Inc., the Global Lipids Genetics Consortium, All of Us and UK Biobank involving 5.1 million individuals of diverse ancestry, with 1.18 million individuals from four non-European populations across thirteen complex traits. Results demonstrate that CT-SLEB significantly improves PRS performance in non-European populations compared to simple alternatives, with comparable or superior performance to a recent, computationally intensive method. Moreover, our simulation studies offer insights into sample size requirements and SNP density effects on multi-ancestry risk prediction.
Importance Twenty-three percent of 37.3M adults in the USA with diabetes are estimated to be undiagnosed, leading to potentially avoidable sequelae and morbidity. Objective To explore the utility of a polygenic risk score (PRS) at identifying individuals with undiagnosed diabetes and prediabetes. Design, Setting and Participants Individuals without doctor-diagnosed diabetes at study baseline in the UK Biobank (UKB) with HbA1c and BMI measurements. Participants were restricted to white individuals to use an ancestry-appropriate PRS. Undiagnosed diabetes and prediabetes were defined using HbA1c (≥6.5% and ≥5.7 - <6.5%, respectively). Exposures A diabetes PRS comprising 13,863 SNPs derived from the 23andMe Research Cohort, and measured BMI among UKB participants. Results Of 412,439 individuals self-reporting an absence of diagnosed diabetes and who had BMI and HbA1c measurements at baseline, 2,934 (0.7%) had undiagnosed diabetes, representing 11.9% of all (diagnosed and undiagnosed) diabetes. Nearly half (1,362, 46%) of undiagnosed diabetes cases were among individuals in the top 25% of the PRS distribution. Overweight individuals (BMI ≥25 - <30 kg/m2) who were in the top 12.5% of the PRS distribution had a similar frequency of undiagnosed diabetes (0.8-1.6% frequency) as individuals with obesity (BMI ≥30kg/m2) in the lowest 12.5% of the PRS distribution (0.7-1.7% frequency). Combining overweight and obesity with the PRS identified nearly all cases of undiagnosed diabetes: individuals with a BMI ≥25 kg/m2 (66% of the study population) or those in the top 54-69% of the PRS identified 98-99% of undiagnosed cases. Of the 199 undiagnosed diabetes cases occurring among individuals with a normal BMI (<25kg/m2), two-thirds were among individuals in the top 50% of the PRS. Prediabetes was common (14%), with measured BMI and PRS providing additive risk. Among those in the top 12.5% PRS with BMI ≥35kg/m2, 6.3% developed incident diabetes over 4 years follow-up, as compared to 0% among the bottom 12.5% PRS with BMI<25kg/m2. Conclusions A diabetes PRS is informative at identifying undiagnosed cases. PRS may have broader utility in detecting individuals with asymptomatic disease. Question Does a polygenic risk score (PRS) have utility in identifying individuals with undiagnosed type 2 diabetes (T2D)? Findings In this analysis of 412,439 individuals without doctor-diagnosed diabetes, a T2D PRS performed additively to body mass index (BMI) at identifying individuals with undiagnosed diabetes. Selecting individuals on the basis of overweight/obesity or a T2D PRS identified almost all cases of undiagnosed diabetes. The majority of undiagnosed diabetes cases among individuals with normal weight occurred among those at elevated polygenic risk. Meaning A T2D PRS identifies cases of undiagnosed diabetes among individuals with and without overweight or obesity. ### Competing Interest Statement All authors are employed by and hold stock or stock options in 23andMe, Inc. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: We used data from the 23andMe, Inc. Research Cohort16 to construct a T2D associated PRS. Individuals included were research participants of 23andMe, Inc., a direct-to-consumer genetics company, who were genotyped as part of the 23andMe Personal Genome Service. Participants provided informed consent and volunteered to participate in the research online, under a protocol approved by the external AAHRPP-accredited IRB, Ethical & Independent (E&I) Review Services. As of 2022, E&I Review Services is part of Salus IRB (). For the UK Biobank data, the UK Biobank has approval from the North West Multi-centre Research Ethics Committee (MREC) as a Research Tissue Bank (RTB) approval. This approval means that researchers do not require separate ethical clearance and can operate under the RTB approval (there are certain exceptions to this which are set out in the Access Procedures, such as re-contact applications). Find more information at This research has been conducted using the UK Biobank Resource under application number 95801. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes To be updated at the time of publication.
Background Human genetics provides opportunities for enhancing disease prediction through polygenic risk scores (PRS). Method We used a dataset from 23andMe (6.77M European, 1.30M Latine, and 0.45M African American individuals). Using cross-sectional data for PRS construction and a prospective cohort for evaluation, we estimated PRS-associated cumulative incidences after one year of follow-up for 12 clinical endpoints. Results The cumulative incidence of disease at one year was consistently higher among individuals in the top 10% of each PRS. Relative risks (RRs) comparing top to bottom 10% of PRS differed across diseases (e.g. European RR 2.12 for eczema vs 12.53 for T2D). Estimates were similar between Europeans and Latines however were more modest for African Americans (e.g. T2D RR 10.92 for Latines vs. 4.00 for African Americans). Clinical manifestation occurred earlier among those in top vs bottom 10% of polygenic risk: 16yrs for hypertension, and 9.5yrs for T2D. Among participants at elevated conventional risk of CHD or T2D, those in the top 10% PRS had a 10-20 fold higher RR of disease incidence vs those not at conventional risk. Among individuals at high polygenic risk of CHD or T2D, favorable lifestyle characteristics associated with 64-73% lower RR of developing disease over 1-year, with cumulative incidence equivalent to the population average. Conclusion In an ancestrally-diverse cohort, individuals in the top 10% PRS had higher 1-year disease incidence and earlier age of clinical manifestation. PRS provided risk stratification beyond conventional risk factors. Lifestyle characteristics markedly lowered disease incidence among those at elevated polygenic risk.
A substantial proportion of the adult United States population with type 2 diabetes (T2D) are undiagnosed, calling into question the comprehensiveness of current screening practices, which primarily rely on age, family history, and body mass index (BMI). We hypothesized that a polygenic score (PGS) may serve as a complementary tool to identify high-risk individuals. The T2D polygenic score maintained predictive utility after adjusting for family history and combining genetics with family history led to even more improved disease risk prediction. We observed that the PGS was meaningfully related to age of onset with implications for screening practices: there was a linear and statistically significant relationship between the PGS and T2D onset (−1.3 years per standard deviation of the PGS). Evaluation of U.S. Preventive Task Force and a simplified version of American Diabetes Association screening guidelines showed that addition of a screening criterion for those above the 90th percentile of the PGS provided a small increase the sensitivity of the screening algorithm. Among T2D-negative individuals, the T2D PGS was associated with prediabetes, where each standard deviation increase of the PGS was associated with a 23% increase in the odds of prediabetes diagnosis. Additionally, each standard deviation increase in the PGS corresponded to a 43% increase in the odds of incident T2D at one-year follow-up. Using complications and forms of clinical intervention (i.e., lifestyle modification, metformin treatment, or insulin treatment) as proxies for advanced illness we also found statistically significant associations between the T2D PGS and insulin treatment and diabetic neuropathy. Importantly, we were able to replicate many findings in a Hispanic/Latino cohort from our database, highlighting the value of the T2D PGS as a clinical tool for individuals with ancestry other than European. In this group, the T2D PGS provided additional disease risk information beyond that offered by traditional screening methodologies. The T2D PGS also had predictive value for the age of onset and for prediabetes among T2D-negative Hispanic/Latino participants. These findings strengthen the notion that a T2D PGS could play a role in the clinical setting across multiple ancestries, potentially improving T2D screening practices, risk stratification, and disease management.
Polygenic risk scores are becoming increasingly predictive of complex traits, but subpar performance in non-European populations raises concerns about their potential clinical applications. We develop a powerful and scalable method to calculate PRS using GWAS summary statistics from multi-ancestry training samples by integrating multiple techniques, including clumping and thresholding, empirical Bayes and super learning. We evaluate the performance of the proposed method and a variety of alternatives using large-scale simulated GWAS on ~19 million common variants and large 23andMe Inc. datasets, including up to 800K individuals from four non-European populations, across seven complex traits. Results show that the proposed method can substantially improve the performance of PRS in non-European populations relative to simple alternatives and has comparable or superior performance relative to a recent method that requires a higher order of computational time. Further, our simulation studies provide novel insights to sample size requirements and the effect of SNP density on multi-ancestry risk prediction.
BackgroundReliability of prostate cancer (PCa) genetic risk score (GRS), that is, the concordance between its estimated risk and observed risk, is required for genetic testing at the individual level. Reliability data are lacking for non-European racial/ethnic populations, which hinders its clinical use and exacerbates racial disparity.ObjectiveTo calibrate PCa ancestry-specific GRS in four racial/ethnic populations.Design, setting, and participantsPCa ancestry-specific GRSs, calculated from published risk-associated single-nucleotide polymorphisms in corresponding racial/ethnic populations, were evaluated in men who participated in 23andMe, Inc. genetic testing and consented for research, including 888 086 of European (EUR), 81 109 of Hispanic (HIS), 30 472 of African (AFR), and 13 985 of East Asian (EAS) ancestry, as classified by 23andMe's ancestry composition algorithm.Outcome measurements and statistical analysisThe concordance between the observed and estimated PCa risks at ten ancestry-specific GRS deciles was measured primarily by using the calibration slope (β), where 1 represents a perfect calibration. Platt scaling was used to correct the systematic bias of GRS.Results and limitationsA linear trend of an increased observed PCa prevalence in men with higher ancestry-specific GRS deciles was found in each racial population (all p-trend < 0.001). A calibration analysis revealed a systematic bias of GRS; β was considerably lower than 1 (0.73, 0.64, 0.66, and 0.75 in EUR, HIS, AFR, and EAS ancestries, respectively). This bias was reduced after the Platt scaling correction: β for scaled GRS in the testing dataset (40% of individuals) approximated 1 for all groups (0.95, 1.05, 1.02, and 1.01 in EUR, HIS, AFR, and EAS populations, respectively). The generalizability of the Platt correction needs to be validated in independent cohorts.ConclusionsA systematic bias of ancestry-specific GRS in the direction of an overestimated risk for men in the highest decile was found in EUR and non-EUR populations. GRS is well calibrated after correction and is appropriate for genetic testing at the individual level for personalized PCa screening.Patient summaryA corrected genetic risk score is more reliable (supported by the observed prostate cancer [PCa] risk) and appropriate for genetic testing for personalized PCa screening.
Current guidelines recommend BRCA1 and BRCA2 genetic testing for individuals with a personal or family history of certain cancers. Three BRCA1/2 founder variants — 185delAG (c.68_69delAG), 5382insC (c.5266dupC), and 6174delT (c.5946delT) — are common in the Ashkenazi Jewish population. We characterized a cohort of more than 2,800 research participants in the 23andMe database who carry one or more of the three Ashkenazi Jewish founder variants, evaluating two characteristics that are typically used to recommend individuals for BRCA testing: self-reported Jewish ancestry and family history of breast, ovarian, prostate, or pancreatic cancer. Of the 1,967 carriers who provided self-reported ancestry information, 21% did not self-report Jewish ancestry; of these individuals, more than half (62%) do have detectable Ashkenazi Jewish genetic ancestry. In addition, of the 343 carriers who provided both ancestry and family history information, 44% did not have a first-degree family history of a BRCA -related cancer and, in the absence of a personal history of cancer, would therefore be unlikely to qualify for clinical genetic testing. These findings may help inform the discussion around broader access to BRCA genetic testing.
With the rise in the prevalence of type 2 diabetes (T2D), as well as undiagnosed cases of T2D and prediabetes (25% and 90%, respectively), early detection is imperative to minimize individual and societal burden. T2D is highly heritable, and personal genetic information is increasingly available to the general public. Studies have suggested that T2D risk reduction strategies may be more effective for individuals with high T2D genetic risk, supporting the use of genetics as a screening tool to inform cost-effective interventions. We trained a polygenic risk score (PRS) for T2D based on >1,200 genotyped variants in >600,000 European consented research participants from a consumer genetic database who self-reported if they had been diagnosed with T2D. We tested the PRS' performance in separate sets of participants covering five different ancestries (African-American, East-Asian, European, Latino, and South-Asian; ~ 600,000 participants). The area under the receiver-operator curve of this PRS varied from 0.65 to 0.57, performing best in European and worst in African ancestries. The PRS was calibrated separately in each ancestry to account for differences in T2D prevalence. European participants with a PRS in the top 5% of the distribution have a T2D odds ratio of more than 3, and lifetime risk for this group exceeds 65%. Our PRS is strongly correlated with an independently derived T2D PRS from (Scott et al. 2017 GWAS, Spearman rho=0.44, p < 1x10^-200) but is more predictive in our dataset (AUC 0.65 vs. 0.59). Lastly, we defined an "increased likelihood" result based on the PRS threshold at which risk of T2D from genetics alone exceeds the risk of T2D due to being overweight. In our database, 19% of individuals met this criterion. We find that personalized PRS' have the potential to identify large numbers of individuals with increased T2D susceptibility equal to or greater than known risk factors and could prove useful in evaluating T2D risk at individual as well as population levels. Disclosure M.L. Multhaup: Employee; Self; 23andMe. R. Kita: Employee; Self; 23andMe. N. Eriksson: None. S. Aslibekyan: Employee; Self; 23andMe. J. Shelton: None. R.I. Tennen: Employee; Self; 23andMe. E. Kim: Employee; Self; 23andMe, inc. B. Koelsch: Employee; Self; 23andMe.
INTRODUCTION: The prevalence of celiac disease (CD) is widely variable throughout the United States (US), with a higher prevalence of disease in the Northeast. The reasons for this variability are unknown. In a prior study, we detected ethnic differences within the US, using a name-based algorithm, with prevalence in patients of Jewish ethnicity similar to the overall population, and lower in persons of East Asian ethnicity. CD etiology is dependent on human leukocyte antigen (HLA) haplotype. Typically, either HLA DQ2.5 or DQ8 is required (but not sufficient) for the development of CD, with DQ2.5 being the highest-risk haplotype. To date, no study has characterized regional or ethnic differences in the frequency of CD-compatible HLA haplotypes. Thus, we aimed to measure the frequencies of DQ2.5 and DQ8 across regions and ethnicities in the US. METHODS: We assessed the frequencies of HLA DQ2.5 (DQA1*05:DQB1*02) and DQ8 (DQA1*03:DQB1*03) in an unselected group of genotyped individuals who have used direct-to-consumer genetic testing between 2013 and 2017. Eligible participants were 23andMe customers who consented to participate in research. We assayed two SNPs to classify individuals as DQ2.5 homozygous, DQ2.5 heterozygous, DQ2.5/DQ8, DQ8 homozygous, DQ8 heterozygous, and 0 detected variants. We compared the frequency of each haplotype across four regions of the US. Additionally, we used genome-wide array data to cluster participants into 8 categories that correlate highly with self-reported race and ethnicity, and compared the frequencies of these haplotypes across these ethnic categories (Table 1). RESULTS: Of 1,290,668 individuals studied, at least one CD-compatible haplotype was present in 38.7% of individuals, and this frequency was similar across the four US regions. The frequencies of DQ2.5 homozygotes were also similar across the Northeast, Midwest, South, and West (1.25%, 1.43%, 1.38%, and 1.37%, respectively). In contrast, frequencies differ across ethnic groups: the highest DQ2.5 and DQ8 frequencies were observed in European (12.01%) and Ashkenazi Jewish (16.39%) participants, respectively. CONCLUSION: Previously reported regional variability in CD prevalence in the US may not be due to differences in HLA-based susceptibility; rather, other genetic or environmental factors likely play a role in disease pathogenesis. In addition, these differences carry great significance in view of the development of HLA haplotype-specific non-dietary therapies for CD.