Introduction:Parkinson's disease (PD) affects females and males differently, with differences in prevalence, clinical phenotypes, and therapeutic response, suggesting that biological sex may influence the underlying molecular mechanisms of PD. The extent to which genetic factors may contribute to these differences remains largely unknown. Objective:To investigate sex-specific autosomal genetic factors associated with PD risk. Methods:We performed a sex-stratified autosomal GWAS meta-analysis leveraging data from the Global Parkinson's Genetics Program, the International Parkinson's Disease Genomics Consortium, the UK Biobank, and the Fox Insight Genetics Study, including a total of 226,196 individuals from different European populations: 18,145 female PD cases, 95,558 female controls, 28,747 male PD cases, and 83,746 male controls. Results:We observed a high genetic correlation between the male and female PD meta-analyses (rg = 0.909, SE = 0.0403; p = 8.03E-113), and the heritability estimates were comparable between sexes (~10% in males and ~11% in females) and similar to estimates from prior sex-combined analyses. Our sex-stratified GWAS identified 57 genome-wide significant association signals, including five novel risk loci, three of which reached genome-wide significance in males only (RBM8A, ANKRD23, and CNTN4) and two in females only (RERE and ARL6IP6). Of the remaining previously identified GWAS loci regions, several showed differences in effect magnitude between sexes, with the GALC, RERE, ARL6IP6, and RBM8A loci demonstrating statistically significant sex-specific effects. Conclusions:Overall, PD genetic architecture appears broadly similar between females and males, but the identification of five novel loci and significant differences at select regions highlights the value of sex-stratified analyses for uncovering additional genetic contributors to PD risk beyond those detected in combined analyses.
Background and ObjectivesKnown pathogenic variants (PVs) in Parkinson disease (PD) contribute to disease development but have yet to be fully explored by arrays on a large scale. This study evaluated genotyping success of the NeuroBooster array (NBA) and determined the frequencies of PVs across ancestries.MethodsWe analyzed the presence and allele frequency of PVs in 28,710 PD cases, 9,614 other neurodegenerative disorder cases, and 15,821 controls across 11 ancestries within the Global Parkinson's Genetics Program (GP2) data set. Cluster plots were used to assess the quality of PVs genotyped on NBA.ResultsGenes previously predicted to have high or very high confidence of causing PD tend to have more PVs and are present across ancestry groups. Of 34 known PD gene PVs assessed, 25 were typed by NBA and classified as "good" (n = 12), "medium" (n = 4), or "bad" (n = 9) quality variants.DiscussionOur results confirm the likelihood that established PD genes are pathogenic and highlight the importance of ancestrally diverse research in PD. We also show the usefulness of the NBA as a reliable tool for the genotyping of rare variants of PD.
Elucidating the genetic contributions to Parkinson's disease aetiology across diverse ancestries is a critical priority for the development of targeted therapies in a global context. We conducted the largest sequencing characterization of potentially disease-causing, protein-altering and splicing mutations in 710 cases and 11 827 controls from genetically predicted African or African admixed ancestries. We explored copy number variants (CNVs) and runs of homozygosity in prioritized early onset and familial cases. Our study identified rare GBA1 coding variants to be the most frequent mutations among patients with Parkinson's disease, with a frequency of 4% in our case cohort. Of the 18 GBA1 variants identified, 10 were previously classified as pathogenic or likely pathogenic, four were novel and four were reported as of uncertain clinical significance. The most common known disease-associated GBA1 variants in the Ashkenazi Jewish and European populations, p.Asn409Ser, p.Leu483Pro, p.Thr408Met and p.Glu365Lys, were not identified among the screened Parkinson's disease cases of African and African admixed ancestry. Similarly, the European and Asian LRRK2 disease-causing mutational spectrum, including LRRK2 p.Gly2019Ser and p.Gly2385Arg genetic risk factors, did not appear to play a major role in Parkinson's disease aetiology among West African ancestry populations. However, we found three heterozygous novel missense LRRK2 variants of uncertain significance, with two (p.Glu268Ala and p.Arg1538Cys) displaying higher frequencies in the African ancestry population reference datasets. Structural variant analyses revealed the presence of PRKN CNVs with a frequency of 0.7% in African and African admixed cases, with 66% of CNVs detected being compound heterozygous or homozygous in early-onset cases, providing further insights into the genetic underpinnings in early-onset juvenile Parkinson's disease in these populations. Short tandem repeat analysis also identified ATXN3 CAG repeat expansions within the pathogenic range (CAGn > 45) in three patients with Parkinson's disease of African ancestry. Novel genetic variation among screened genes warrants further replication and functional prioritization to unravel their pathogenic potential. Here, we created the most comprehensive genetic catalogue of both known and novel coding and splicing variants potentially linked to Parkinson's disease aetiology in an underserved population and further conducted global and local ancestry analyses to further explore population-specific effects. Our study has the potential to guide the development of targeted therapies in the emerging era of precision medicine. By expanding genetics research to involve underrepresented populations, we hope that future Parkinson's disease treatments are not only effective but also inclusive, addressing the needs of diverse ancestral groups.
Motivation:Copy Number Variations (CNVs) play pivotal roles in complex disease etiology, often requiring large sample sizes to analyze disease associations. While genotyping arrays offer a cost-effective approach for CNV detection using Log R Ratio (LRR) and B Allele Frequency (BAF) signals, existing independent array-based callers suffer from high false positive rates and noise susceptibility, burdening manual validation. Results:We present CNV-Finder, a deep learning pipeline employing Long Short-Term Memory (LSTM) networks for large-scale CNV identification within user-defined genomic regions. Trained on expert-annotated samples from the Global Parkinson's Genetics Program across four neurodegenerative disease-associated genes (PRKN, LINGO2, MAPT, SNCA), CNV-Finder integrates human feedback to iteratively improve performance. In benchmarking across 105 936 samples spanning 11 ancestries and nearly 150 cohorts, the model achieved 91% and 89% visual confirmation rates for PRKN deletions and duplications at high-confidence thresholds. In two validation cohorts, CNV-Finder nominated 83% fewer candidates than a popular Hidden Markov Model-based caller while maintaining higher confirmation rates. Validation through MLPA, short-read, and long-read sequencing demonstrated robust performance, generalizing to diverse signatures including homozygous deletions and SNCA triplications absent from training. Our findings highlight human expertise's value in complex loci like 17q21.31. Availability and implementation:CNV-Finder is freely available at https://github.com/nvk23/CNV-Finder.
In the Global Parkinson's Genetics Program (GP2) we aim to advance precision medicine by integrating large-scale clinico-genetic data from diverse populations worldwide. We investigated potentially trial-eligible carriers of pathogenic and high-risk GBA1 and LRRK2 variants and conducted a global precision-medicine survey across GP2 sites. Among 65,509 individuals with Parkinson's disease, we identified 9,019 (13.8%) potentially trial-eligible genetic variant carriers, including 6,789 GBA1, 2,084 LRRK2, and 146 dual GBA1-LRRK2 carriers. Individuals were distributed across multiple global regions, many of which currently lack active gene-targeted trials, highlighting a global disparity between relevant variant carriers and the availability of disease modifying treatment trials. GP2's unified framework supports equitable recruitment for gene-targeted therapeutic studies and helps address critical gaps in Parkinson's disease genetics and future therapeutic development.
Introduction Genome-wide association studies (GWAS) have identified over 130 risk loci for Parkinson's disease (PD), yet the majority derive from studies performed in European ancestry populations. African (AFR) and African admixed (AAC) ancestry individuals remain underrepresented in PD genetics research, limiting our understanding of ancestry-specific genetic architecture and the generalizability of known risk factors. Methods We conducted GWAS in AFR and AAC populations by integrating individual-level genotype data from the Global Parkinson's Genetics Program (GP2) with summary statistics from 23andMe Research Institute and the Million Veterans Program. The combined dataset included 3,975 cases and 319,883 controls, representing a 64% increase in total sample size compared with prior analyses. We performed separate GWAS for AFR and AAC cohorts as well as a combined AFR/AAC meta-analysis. Results The intronic GBA1 variant rs3115534 was the most significant association across all analyses, reaching genome-wide significance in AAC individuals for the first time. In the AFR-only analysis, five loci achieved genome-wide significance: GBA1 (rs3115534), the SNCA signal previously reported in European ancestry GWAS (rs356182), a new protein-coding association at LRRK2 (rs72546327, p.T1410M), a non-coding RPL10P13 variant (rs12302417), and a novel signal on chromosome 16 (rs113244182). The combined AFR/AAC meta-analysis identified four genome-wide significant associations at GBA1 (rs3115534), SNCA (rs356182), SCARB2 (rs11547135), and LRRK2 (rs139283662, which is in LD with p.T1410M). Conclusions This study reports the largest GWAS of PD in AFR and AAC populations to date. Our findings confirm trans-ancestry risk loci (GBA1 and SCARB2) and identify an ancestry-enriched coding variant at LRRK2. This convergence of evidence around genes involved in glucocerebrosidase (GCase) trafficking and alpha-synuclein clearance supports current therapeutic strategies targeting this pathway and provides critical targets for developing precision medicine in African ancestry populations. Importantly, the identification of a novel association between a LRRK2 coding variant with disease in the AFR and AAC populations opens up a traditionally underrepresented population for ongoing LRRK2 targeted trials. Furthermore, the identification of novel ancestry-specific loci, including those that are directly relevant to current therapeutic deployment, underlines the importance of understanding the basis of disease in all populations. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This project was supported by the Global Parkinsons Genetics Program (GP2; https://gp2.org). GP2 is funded by the Aligning Science Across Parkinsons (ASAP) initiative and implemented by The Michael J. Fox Foundation for Parkinsons Research (MJFF). For a complete list of GP2 members see https://doi.org/10.5281/zenodo.7904831. This research was supported by the Aligning Science Across Parkinson's Initiative, the Intramural Research Program, National Institute on Aging, National Institutes of Health, Department of Health and Human Services, project ZO1 AG000949, and the Michael J. Fox Foundation for Parkinson's Research. This work utilized the computational resources of the NIH STRIDES Initiative (https://cloud.nih.gov) through the Other Transaction agreement - Azure: OT2OD032100, Google Cloud Platform: OT2OD027060, Amazon Web Services: OT2OD027852. We would like to thank the research participants, Paul Cannon, and employees of 23andMe Research Institute for making this work possible. We would also like to thank Dario Alessi and his team for their valuable contributions in functionally contextualizing this work. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data used in the preparation of this article were obtained from GP2. Specifically we used Tier 2 data from GP2 (release 11: DOI 10.5281/zenodo.17753486). GP2 data can be accessed through AMP PD (https://amp-pd.org). For the MVP dataset, PD summary statistics from the Million Veterans Program (MVP) were downloaded from dbGAP (accession number: phs002453.v1.p1; analysis accession: pha010400.1). Summary statistics from 23andMe were shared under a collaborative agreement submitted at https://research.23andme.com/collaborate/. All code generated for this article, and the identifiers for all software programs and packages used, are available on GitHub (https://github.com/GP2code/GP2-AFR-AAC-metaGWAS) and were given a persistent identifier via Zenodo (DOI: 10.5281/zenodo.7888140)
Background and Objectives:Known pathogenic variants (PVs) in Parkinson disease (PD) contribute to disease development but have yet to be fully explored by arrays on a large scale. This study evaluated genotyping success of the NeuroBooster array (NBA) and determined the frequencies of PVs across ancestries. Methods:We analyzed the presence and allele frequency of PVs in 28,710 PD cases, 9,614 other neurodegenerative disorder cases, and 15,821 controls across 11 ancestries within the Global Parkinson's Genetics Program (GP2) data set. Cluster plots were used to assess the quality of PVs genotyped on NBA. Results:Genes previously predicted to have high or very high confidence of causing PD tend to have more PVs and are present across ancestry groups. Of 34 known PD gene PVs assessed, 25 were typed by NBA and classified as "good" (n = 12), "medium" (n = 4), or "bad" (n = 9) quality variants. Discussion:Our results confirm the likelihood that established PD genes are pathogenic and highlight the importance of ancestrally diverse research in PD. We also show the usefulness of the NBA as a reliable tool for the genotyping of rare variants of PD.
Expanded short tandem repeats contribute to a broad spectrum of neurodegenerative diseases, yet their roles in Parkinson's disease (PD) and parkinsonism remain incompletely characterized, especially across diverse ancestries. We analyzed short-read whole-genome (WGS) and clinical exome sequencing (CES) data from 38,365 individuals (28,861 WGS; 9,504 CES), encompassing 23,242 patients with PD, 4,729 patients with atypical parkinsonism and 10,394 healthy controls from 11 genetic ancestries. To determine carrier frequencies and characterize repeat structures across diverse ancestries, we genotyped 12 established pathogenic loci where normal, intermediate, and pathogenic alleles can be reliably differentiated using short-read sequencing data. Additionally, we conducted threshold-based associations to determine the minimum threshold associated with increased PD risk in 15,995 individuals (8,591 PD, 7,404 controls) of European ancestry. Pathogenic repeat expansions were detected in 62 patients (56 PD and 6 atypical parkinsonism) and 5 controls across seven loci (AR, ATXN1, ATXN2, ATXN3, CACNA1A, HTT and THAP11), spanning seven ancestries. Among these, ATXN2 expansions were the most frequently observed in PD and were present in African, East Asian, European and Middle Eastern ancestries. Additionally, intermediate ATXN2 repeat expansions exhibited a strong, length-dependent association with PD risk in the European population, with individuals with ≥32 repeats having a more than four-fold increased risk (odds ratio 4.25, 95% confidence interval 1.80-12.05). Overall, >92% of expanded alleles harbor CAA interruptions within the CAG tract. Pathogenic expansions at other loci, such as ATXN3 and THAP11, showed more ancestry-specific distributions. Clinically, individuals with pathogenic ATXN2 and ATXN3 expansions most often presented with typical PD features but frequently showed earlier disease onset and a strong family history of PD. This large-scale, multi-ancestry study comprehensively maps the genetic landscape of pathogenic and intermediate repeat expansions in PD. Our findings confirm a length- and structure-dependent risk association for ATXN2 with PD in the European population and highlight the pleiotropic effects of repeat expansions across the parkinsonian spectrum.
BACKGROUND:The genetic architecture of Parkinson's disease varies considerably across ancestries, yet most previous genetic studies have focused on individuals of European ancestry. We aimed to characterise the distribution of established Parkinson's disease causal variants, as well as risk-associated variants with clinical implications (ie, variants in genes involved in pathways targeted by ongoing clinical trials), across ancestrally diverse populations. METHODS:We conducted a multi-ancestry, observational, cross-sectional genetic study using retrospective data from the Global Parkinson's Genetics Program (GP2) release 11 (released in December, 2025). The study investigated causal and risk variants, including copy number variants, in established Parkinson's disease and parkinsonism-associated genes, following the recommendations of the Movement Disorder Society (MDS) Task Force on the Nomenclature of Genetic Movement Disorders, including GBA1, LRRK2, SNCA, VPS35, RAB32, PINK1, PRKN, PARK7, ATP13A2, DCTN1, DNAJC6, FBXO7, JAM2, RAB39B, SLC20A2, SYNJ1, VPS13C, and WDR45. Individuals with Parkinson's disease were diagnosed based on established clinical criteria, including the Parkinson's UK Brain Bank or MDS diagnostic criteria (or both), and healthy control participants were defined as individuals without evidence of neurodegenerative disease and unrelated to participants with Parkinson's disease. We analysed genome and exome sequencing and array genotyping data of 99 783 individuals, including 58 559 individuals with Parkinson's disease and 41 224 controls, from 11 genetically inferred ancestries (African, African admixed, Ashkenazi Jewish, Latino and Indigenous people of the Americas, central Asian, complex admixture, east Asian, European, Finnish, Middle Eastern, and south Asian), defined using reference population-based ancestry inference methods. We calculated allele frequencies for all investigated variants in individuals with Parkinson's disease and controls, both overall and stratified by ancestry. FINDINGS:Approximately 29% of individuals (29 001 of 99 783; 15 443 [26·4%] of 58 559 individuals with Parkinson's disease and 13 558 [32·9%] of 41 224 controls) were from under-represented populations (ie, non-European and non-Ashkenazi Jewish). Our findings indicated both shared genetic contributors across ancestries as well as ancestry-specific differences in variant frequencies and the spectrum of variants within Parkinson's disease-associated genes. Overall, 1217 (2·1%) of 58 559 individuals with Parkinson's disease carried a causal variant, with substantial variations across ancestries ranging from ten (0·4%) of 2844 African individuals to 251 (10·7%) of 2343 individuals of Ashkenazi Jewish ancestry. Risk variants in GBA1 and LRRK2 were identified in 6893 (11·8%) of 58 559 individuals with Parkinson's disease and 3578 (8·7%) of 41 224 controls. GBA1 risk variants were most frequent overall and identified across all ancestries, but variant frequency and spectra differed substantially between ancestries, from 195 (4·1%) of 4773 in the east Asian ancestry group to 1505 (52·9%) of 2844 in the African ancestry group. Similarly, LRRK2 causal and risk variants showed ancestry-specific enrichment, with the highest frequencies of causal variants in the Ashkenazi Jewish (250 [10·7%] of 2343) and Middle Eastern (59 [4·4%] of 1347) ancestry groups, whereas risk variants were predominantly identified in the east Asian ancestry group (601 [12·6%] of 4773). Carriers of biallelic causal variants in PRKN, commonly including deletions and duplications, were also identified across all ancestries except Ashkenazi Jewish; the highest frequency was in the Middle Eastern ancestry group (17 [1·3%] of 1347), and frequencies in all other ancestries were less than 1%. INTERPRETATION:This large-scale, multi-ancestry genetic study offers crucial insights into the population-specific genetic architecture of Parkinson's disease. Whereas clinical trials targeting GBA1 and LRRK2 variant carriers are primarily performed in Europe and the USA, increased ancestral diversity in Parkinson's disease research will be crucial to improve diagnostic accuracy, enhance our understanding of disease mechanisms across populations, and ensure equitable application of and access to emerging genetically informed therapies. FUNDING:Aligning Science Across Parkinson's (ASAP) through the Global Parkinson's Genetics Program (GP2).
Latin America's diverse genetic landscape provides a unique opportunity to study Alzheimer's disease (AD) and frontotemporal dementia (FTD). The Multi-Partner Consortium to Expand Dementia Research in Latin America (ReDLat) recruited 2162 participants with AD, FTD, or healthy controls from six countries: Argentina, Brazil, Chile, Colombia, Mexico, and Peru. Participants underwent genomic sequencing, and population structure analyses were conducted using Principal Component Analysis and ADMIXTURE. The study revealed a predominant mix of American, African, and European ancestries, with an additional East Asian component in Brazil. Variant curation identified 17 pathogenic variants, a pathogenic C9orf72 expansion, and 44 variants of uncertain significance. Seventy families showed autosomal dominant inheritance, with 48 affected by AD and 22 by FTD. This represents the first large-scale genetic study of AD and FTD in Latin America, highlighting the need to consider diverse ancestries, social determinants of health, and cultural factors when assessing genetic risk for neurodegenerative diseases.
BACKGROUND:Large-scale sequencing initiatives have generated extensive genomic resources essential for variant interpretation, yet their effective use often requires bioinformatics expertise. To support identification of Parkinson's disease (PD) risk and disease-causing variants, we developed an open-access, summary-level genomic data browser. METHODS:We performed uniform joint variant calling to harmonize whole-genome sequencing (WGS) data from AMP-PD Release 4, GP2 Data Releases, and additional controls from the Alzheimer's Disease Sequencing Project. Clinical-exome sequencing (CES) data from GP2 Release 8 were also included. RESULTS:The integrated dataset included 31,665 WGS and 9,559 CES samples, spanning 11 ancestries and over 300 million variants. CONCLUSIONS:The GP2 Genome Browser is a lightweight, flexible platform providing intuitive gene- and variant-level summaries with ancestry-stratified allele frequencies and functional annotations. It is open source and freely accessible at https://gp2.broadinstitute.org, enabling broad access to PD genomic data and supporting global research efforts. © 2026 The Author(s). Movement Disorders published by Wiley Periodicals LLC on behalf of International Parkinson and Movement Disorder Society. © 2026 The Author(s). Movement Disorders published by Wiley Periodicals LLC on behalf of International Parkinson and Movement Disorder Society.
Although large-scale genetic association studies have proven opportunistic for the delineation of neurodegenerative disease processes, we still lack a full understanding of the pathological mechanisms of these diseases, resulting in few appropriate treatment options and diagnostic challenges. To mitigate these gaps, the Neurodegenerative Disease Knowledge Portal (NDKP) was created as an open-science initiative with the aim to aggregate, enable analysis, and display all available genomic datasets of neurodegenerative disease, while protecting the integrity and confidentiality of the underlying datasets. The portal contains 218 genomic datasets, including genotyping and sequencing studies, of individuals across ten different phenotypic groups, including neurological conditions such as Alzheimer's disease, amyotrophic lateral sclerosis, Lewy body dementia, and Parkinson's disease. In addition to securely hosting large genomic datasets, the NDKP provides accessible workflows and tools to effectively utilize the datasets and assist in the facilitation of customized genomic analyses. Here, we summarize the genomic datasets currently included within the portal, the bioinformatics processing of the datasets, and the variety of phenotypes captured. We also present example use-cases of the various user interfaces and integrated analytic tools to demonstrate their extensive utility in enabling the extraction of high-quality results at the source, for both genomics experts and those in other disciplines. Overall, the NDKP promotes open-science and collaboration, maximizing the potential for discovery from the large-scale datasets researchers and consortia are expending immense resources to produce and resulting in reproducible conclusions to improve diagnostic and therapeutic care for neurodegenerative disease patients.
Alzheimer's disease (AD) and Parkinson's disease (PD) are influenced by genetic and environmental factors. We conducted a biobank-scale study to (i) identify endocrine, nutritional, metabolic, and digestive disorders with potential causal or temporal associations with AD/PD risk before diagnosis; (ii) assess plasma biomarkers' specificity for AD/PD in the context of co-occurring gut related traits and disorders; and (iii) integrate multimodal datasets to enhance AD/PD prediction. Our findings show that several disorders were associated with increased AD/PD risk before diagnosis, with variation in the strength and timing of associations across conditions. Polygenic risk scores reveal lower genetic predisposition for AD/PD in individuals with co-occurring disorders. Moreover, the proteomic profile of AD/PD cases was influenced by comorbid gut-brain axis disorders. Last, our multimodal prediction models outperform single-modality paradigms in disease classification. This endeavor illuminates the interplay between factors involved in the gut-brain axis and the development of AD/PD, opening avenues for therapeutic targeting and early diagnosis.
Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5 years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55 kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.
[This corrects the article DOI: 10.1212/NXG.0000000000200246.].
Background:The genetic architecture of Parkinson's disease (PD) varies considerably across ancestries, yet most genetic studies have focused on individuals of European descent, limiting our insights into the genetic architecture of PD at a global scale. Methods:We conducted a large-scale, multi-ancestry investigation of causal and risk variants in PD-related genes. Using genetic datasets from the Global Parkinson's Genetics Program, we analyzed sequencing and genotyping data from 69,881 individuals, including 41,139 affected and 28,742 unaffected, from eleven different ancestries, including ~30% of individuals from non-European ancestries. Findings:Our findings revealed shared and ancestry-specific patterns in the prevalence and spectrum of PD-associated variants. Overall, ~2% of affected individuals carried a causative variant, with substantial variations across ancestries ranging from <0·5% in African, African-admixed, and Central Asian to >10% in Middle Eastern and Ashkenazi Jewish ancestries. Including disease-associated GBA1 and LRRK2 risk variants raised the yield to ~12.5%, largely driven by GBA1, except in East Asians, where LRRK2 risk variants dominated. GBA1 variants were most frequent globally, albeit with substantial differences in frequencies and variant spectra. While GBA1 variants were identified across all ancestries, frequencies ranged from 3·4% in Middle Eastern to 51·7% in African ancestry. Similarly, LRRK2 variants showed ancestry-specific enrichment, with G2019S most frequently seen in Middle Eastern and Ashkenazi Jewish, and risk variants predominating in East Asians. However, clinical trials targeting proteins encoded by these genes are primarily based in Europe and North America. Interpretation:This large-scale, multi-ancestry assessment offers crucial insights into the population-specific genetic architecture of PD. It underscores the critical need for increased diversity in PD genetic research to improve diagnostic accuracy, enhance our understanding of disease mechanisms across populations, and ensure the equitable development and application of emerging precision therapies.
A polygenic score (PGS) for Alzheimer's disease (AD) was derived recently from data on genome-wide significant loci in European ancestry populations. We applied this PGS to populations in 17 European countries and observed a consistent association with the AD risk, age at onset and cerebrospinal fluid levels of AD biomarkers, independently of apolipoprotein E locus (APOE). This PGS was also associated with the AD risk in many other populations of diverse ancestries. A cross-ancestry polygenic risk score improved the association with the AD risk in most of the multiancestry populations tested when the APOE region was included. Finally, we found that the PGS/polygenic risk score captured AD-specific information because the association weakened as the diagnosis was broadened. In conclusion, a simple PGS captures the AD-specific genetic information that is common to populations of different ancestries, although studies of more diverse populations are still needed to better characterize the genetics of AD.
Variants of uncertain significance (VUS) are a bottleneck for genetic discovery and complicate clinical decision-making in Alzheimer's disease and related neurological disorders (ADRD). We developed MoVUS: Model for Variants of Unknown Significance, a random-forest approach that integrates functional predictors to classify missense VUS. MoVUS leverages a balanced random forest model trained on dbNSFP v5.1a with high-confidence ClinVar and HGMD labels, using harmonized functional prediction rankscores. MoVUS produced confident, explainable calls, with ~98% accuracy (AUC ~0.998), prioritizing potentially pathogenic candidates and down-ranked likely benign variants on independent validation sets of ClinVar-only and HGMD-only variants. In our discovery analyses on ADRD-implicated variants in dbNSFP and from independent collaborator cohorts, we achieved high-confidence classifications on a majority of the unknown variants (average of 55% of discovery variants). We also had access to medical records and family trees for some variants, further validating our findings. Across held-out and external datasets, MoVUS reports high accuracy alongside confidence scores and helps prioritize actionable candidates, and reduces bias by considering multiple scores for each variant. To facilitate use, we developed a web app for users to browse across 100+ ADRD genes. MoVUS provides transparent, reproducible triage for research follow-up by pairing consensus predictors with SHAP-based visualizations and explanations.
GenoTools, a Python package, streamlines population genetics research by integrating ancestry estimation, quality control, and genome-wide association studies capabilities into efficient pipelines. By tracking samples, variants, and quality-specific measures throughout fully customizable pipelines, users can easily manage genetics data for large and small studies. GenoTools' "Ancestry" module renders highly accurate predictions, allowing for high-quality ancestry-specific studies, and enables custom ancestry model training and serialization specified to the user's genotyping or sequencing platform. As the genotype processing engine that powers several large initiatives, including the NIH's Center for Alzheimer's and Related Dementias and the Global Parkinson's Genetics Program, GenoTools was used to process and analyze the UK Biobank and major Alzheimer's disease and Parkinson's disease datasets with over 400,000 genotypes from arrays and 5,000 whole genome sequencing samples and has led to novel discoveries in diverse populations. It has provided replicable ancestry predictions, implemented rigorous quality control, and conducted genetic ancestry-specific genome-wide association studies to identify systematic errors or biases through a single command. GenoTools is a customizable tool that enables users to efficiently analyze and scale genotyping and sequencing (whole genome sequencing and exome) data with reproducible and scalable ancestry, quality control, and genome-wide association studies pipelines.