BACKGROUND:Circulating proteins represent robust drug targets with therapeutic potential. Many discoveries have focused on European-ancestry populations, disregarding minuscule yet substantial proteomic differences that may contribute to disease and alter drug generalizability in other ancestry groups. METHODS:Using 2-sample Mendelian randomization and colocalization, we analyzed the effects of 1562 circulating proteins on 145 cardiometabolic-centric outcomes to identify robust protein-phenotype associations in African-ancestry populations and reveal African-ancestry associations with heterogeneous effects. We further replicated these findings using the proteomic data available from the UK Biobank Pharma Proteomics Project and tested the effect of protein quantity in association with select phenotypes. Population branch statistics were also constructed to examine whether protein-genetic instruments under natural selection could lead to significant protein-outcome associations specific to the African ancestry. RESULTS:We identified 115 robust protein target-outcome associations in African-ancestry populations. Among these, 51 demonstrated heterogeneous effects between African- and European-ancestry populations. We further replicated 4 cross-platform African-ancestry associations in the UK Biobank Pharma Proteomics Project and also revealed 4 significant, direct associations between protein levels and phenotypes. Ultimately, based on our prioritization criteria, we found that CD36 (glycoprotein IIIb), APOC1 (apolipoprotein C1), GSTA1 (glutathione S-transferase alpha 1), and FOLH1 (folate hydrolase 1) were shown to influence lipids and heart diseases, and were uniquely represented in African-ancestry populations. In addition, using population branch statistics, we showed that 47.5% of the 115 significant protein-outcome associations were possibly driven by cis-acting protein quantitative trait loci under natural selection. CONCLUSIONS:Multiple lines of evidence were used to interrogate proteomic determinants of cardiometabolic diseases and traits in African-ancestry populations. We highlighted actionable circulating protein targets that could represent potential drug targets for cardiovascular diseases specific to populations with African ancestry.
While international efforts have characterized genetic variation in millions of individuals, the interplay of environmental, social, cultural, and genetic factors is poorly understood for most worldwide populations. The province of Quebec in Canada has been the site of numerous genetic studies, often focusing on individual Mendelian diseases in founder sub-populations. Here, we profiled and analyzed genome-wide genotyped variation in 29,337 Quebec residents from the large population-based cohort CARTaGENE (CaG), including rich phenotype and environmental data. We also sequenced the whole-genome of 2,173 CaG participants, including 163 and 132 individuals with grandparents born in Haiti and Morocco, respectively. We use this genetic information to gain insight into Quebec's demography and to help interpret the potential significance of variants identified in clinically important genes. We built an imputation panel by phasing the CaG whole-genome sequence data and showed, using genome-wide association studies (GWAS), how it improves the discovery of phenotype-genotype associations in this population. We provide allele frequency information and GWAS results through dedicated and publicly available websites. The genetic data, paired with phenotypic and environmental information, is also available for research use upon scientific and ethical review.
Summary:The X chromosome comprises approximately 5% of the human genome and encodes over 800 protein-coding genes, many of which exhibit sex-differentiated expression patterns due to escape from X chromosome inactivation (XCI) mechanisms. Despite its relevance to sex differences in complex traits, the X chromosome is routinely excluded from genome-wide association studies due to analytical challenges, and when analyzed, the impact of escape from XCI or sex is limitedly explored. No dedicated, publicly accessible browser for X chromosome-wide association study (XWAS) summary statistics currently exists, creating a barrier to systematic investigation of X-linked contributions to human traits. Here, we present geneXplore, an interactive web browser based on the PheWeb2 implementation, tailored for XWAS summary statistics across 1,944 phenotypes while distinguishing random XCI (rXCI), escape from XCI (eXCI), and sex-stratified analyses. Users can explore results via interactive plots (Manhattan and Miami, PheWAS and LocusZoom), searchable tables and access to cross-database lookup, with full summary statistics available for download. Availability and Implementation:geneXplore is freely available at https://genexplore.wustl.edu/ with no registration required and will be maintained for a minimum of two years following publication. Source code is available at https://github.com/Belloy-Lab/geneXplore_XWAS_Browser under an MIT license.
The genetic features of founder populations with recent bottlenecks, causing some deleterious variants to rise to higher frequencies, can enhance the power of rare variant association studies. French Canadians from Quebec represent a recent founder population with a particular disease heritage comprising more than 30 prevalent Mendelian conditions. Here, we characterize coding variation in this founder population using exome sequencing data from 2,820 French-Canadian participants - patients with inflammatory bowel diseases (IBD), parents and controls from the Quebec IBD cohort. We find that 18% of rare coding variants are 10-100 times more frequent than in non-Finnish Europeans (NFE). A total of 4,133 missense and loss-of-function variants were significantly enriched with a median 28-fold enrichment, revealing the potential for genotype-phenotype associations in this population. We describe significantly enriched pathogenic variants, including those known to account for the increased prevalence of rare diseases in FC compared to other European descent populations, such as Agenesis of corpus callosum and peripheral neuropathy (SLC12A6) and Leigh Syndrome French Canadian type (LRPPRC). Finally, we investigate whether rare protein-coding variants, enriched in French Canadians by the founder effect, contribute to the risk of IBD using trio and case/control cohorts. In addition to replicating associations in NOD2 and IL23R, we identified new candidate association signals, including enriched variants in SLC35E3, and ARSA. Our findings show that, even in well-characterized founder populations like the French Canadians, there remains untapped potential for genetic discovery, revealing both rare and complex disease risk factors through enriched coding variation.
Fraser syndrome (FS) is an autosomal recessive disorder, characterized by cryptophthalmos, syndactyly, and anomalies of the respiratory and urogenital tracts. Here we estimate the birth prevalence of FS in the French-Canadian founder population of Quebec, where no prevalence has been reported to date. We also describe the phenotype of probands with FS. Pathogenic allele frequency was estimated in the Quebec IBD cohort and the population-based cohort CARTaGENE (CaG), from exome sequencing ( n =2,323), short-read genome sequencing ( n =2,173) and genotyping array data ( n =29,330). Phenotypic data was collected for FS probands at CHU Ste-Justine (between 2013 and 2023). FRAS1 p.(Arg124Ter) was the most frequent pathogenic variant in the Quebec IBD and CaG cohorts, with frequencies 17-fold and 21-fold higher than Non-Finnish Europeans. Four French-Canadian probands were diagnosed with FS at CHU Ste-Justine, three of whom were homozygous for the variant. Birth prevalence was 0.23 per 100,000 births when predicted from the Quebec IBD cohort, 0.98 from CaG, and 1.34 from reported cases. FS shows an elevated birth prevalence in the French-Canadian population of Quebec. We propose FRAS1 p.(Arg124Ter) as a candidate founder pathogenic variant, informing clinical practices in this population.
CRISPR-dependent base editing (BE) enables the modeling and correction of genetic mutations at single-base resolution. Base editing screens, where point mutations are queried en masse, are powerful tools to systematically draw genotype–phenotype associations and characterise the function of genes and other genomic elements. However, the lack of user-friendly web-based tools for designing base editing screens can hinder broad technology adoption. Here, we introduce CRISPR-BEasy (https://crispr-beasy.cerc-genomic-medicine.ca), a free, automated web-based server that streamlines the creation of single guide (sg)RNA tiling libraries for base editing screens. Researchers can provide their genes or genomic features of interest, their base editors of choice, and target sequences to act as positive and negative controls. The server designs and annotates sgRNA libraries by integrating custom code with publicly available tools such as crisprVerse and Ensembl’s Variant Effect Predictor. CRISPR-BEasy provides downloadable results, including sgRNA on/off-target scores, predicted mutational outcomes per base editor, and intuitive interactive visualizations for data quality assessment. CRISPR-BEasy also provides a separate tool that assembles sgRNA libraries into oligonucleotides for cloning following the detailed protocol documented in the searchable web server manual. Together, CRISPR-BEasy ensures the seamless design of cloning-ready sgRNA libraries, seeking to democratise access to base editing screening technologies.
The lack of functionalities in web-based tools for interacting with stratified genome-wide association study (GWAS) summary-level results is currently hindering researchers from advancing knowledge of ancestry and sex on the genetics of complex human diseases and traits. Here we introduce PheWeb 2, a completely rewritten enhanced version of our original web-based tool, which offers intuitive and efficient interactive navigation and visual comparison across stratified GWAS results within a single framework. *Justin Bellavance & Hongyu Xiao contributed equally.
Gene genealogies represent the shared ancestry of a sample and are often encoded as ancestral recombination graphs (ARGs). It has recently become possible to infer these gene genealogies from sequencing or genotyping data and use them for many evolutionary and statistical genetics applications. Here, we use the ARG inference software ARG-needle and the pedigree imputation software ISGen to impute and trace the transmission of disease variants in founder populations where long shared haplotypes allow for accurate timing of relatedness. We applied these methods to the population of Quebec, where multiple founder events led to an uneven distribution of pathogenic variants across regions and where extensive population pedigrees are available via the BALSAC project. We validated this approach with nine founder mutations for the Saguenay-Lac-Saint-Jean region, demonstrating high accuracy for mutation age, imputation, and regional frequency estimation. We used imputed carrier status in a longitudinal cohort to highlight heterozygote effects for known recessive alleles. These heterozygote effects, together with regional frequency estimates, can inform the design of screening programs.
Intellectual disability (ID) is a neurodevelopmental disorder affecting up to 1-3% of people worldwide. Genetic factors, including rare de novo or rare homozygous mutations, explain many cases of autosomal dominant or recessive forms of ID. ID is clinically and genetically heterogeneous, with hundreds of genes associated with it. In this study, we performed high-depth whole-genome sequencing of twenty individuals from five consanguineous families from Pakistan, with nine individuals affected by mild or severe ID. We identified one splice and five missense rare variants (at allele frequencies below 0.001%) in a homozygous state in the affected individuals with supporting and moderate evidence of pathogenicity based on guidance from the American College of Medical Genetics and Genomics. These six variants mapped to different genes ( SRD5A3 , RDH11 , RTF2 , PCDHA2 , ADAMTS17 , and TRPC3 ), and only SRD5A3 had previously been known to cause ID. The p.Tyr169Cys mutation inside SRD5A3 was predicted to be deleterious and affect protein structure by multiple in silico tools. In addition, we found one missense mutation, p.Pro1505Ser, inside UNC13B with conflicting evidence of pathogenic and benign effects. Further functional studies are required to confirm the pathogenicity of these variants and understand their role in ID. Our findings provide additional needed information for interpreting rare variants in the genetic testing of ID. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement A.W. was supported by the Higher Education Commission of Pakistan. The sequencing experiments and analyses of sequencing data in this study were supported by the McGill Canada Excellence Research Chair Program in Genomic Medicine (V.M.). V.M. holds a Canada Excellence Research Chair. Calcul Quebec and the Digital Research Alliance of Canada provided the computational resources for this study. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The ORIC (Office of Research Innovation and Commercialization) and DSAR (Director of Advanced Study and Research) boards of the University of Azad Jammu and Kashmir, Muzaffarabad, Pakistan, approved the study ethics protocol. The DNA sequencing and genetic data analyses were also approved by the Research Ethics Office (IRB) of the Faculty of Medicine and Health Sciences at McGill University, Montreal, Canada. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes
Polygenic scores (PGS) enable the prediction of genetic predisposition for a wide range of traits and diseases by calculating the weighted sum of allele dosages for genetic variants associated with the trait or disease in question. Present approaches for calculating PGS from genotypes are often inefficient and labor-intensive, limiting transferability into clinical applications. Here, we present ‘Imputation Server PGS’, an extension of the Michigan Imputation Server designed to automate a standardized calculation of polygenic scores based on imputed genotypes. This extends the widely used Michigan Imputation Server with new functionality, bringing the simplicity and efficiency of modern imputation to the PGS field. The service currently supports over 4489 published polygenic scores from publicly available repositories and provides extensive quality control, including ancestry estimation to report population stratification. An interactive report empowers users to screen and compare thousands of scores in a fast and intuitive way. Imputation Server PGS provides a user-friendly web service, facilitating the application of polygenic scores to a wide range of genetic studies and is freely available at https://imputationserver.sph.umich.edu.
Type 2 diabetes (T2D) is a heterogeneous disease that develops through diverse pathophysiological processes1,2 and molecular mechanisms that are often specific to cell type3,4. Here, to characterize the genetic contribution to these processes across ancestry groups, we aggregate genome-wide association study data from 2,535,601 individuals (39.7% not of European ancestry), including 428,452 cases of T2D. We identify 1,289 independent association signals at genome-wide significance (P < 5 × 10-8) that map to 611 loci, of which 145 loci are, to our knowledge, previously unreported. We define eight non-overlapping clusters of T2D signals that are characterized by distinct profiles of cardiometabolic trait associations. These clusters are differentially enriched for cell-type-specific regions of open chromatin, including pancreatic islets, adipocytes, endothelial cells and enteroendocrine cells. We build cluster-specific partitioned polygenic scores5 in a further 279,552 individuals of diverse ancestry, including 30,288 cases of T2D, and test their association with T2D-related vascular outcomes. Cluster-specific partitioned polygenic scores are associated with coronary artery disease, peripheral artery disease and end-stage diabetic nephropathy across ancestry groups, highlighting the importance of obesity-related processes in the development of vascular outcomes. Our findings show the value of integrating multi-ancestry genome-wide association study data with single-cell epigenomics to disentangle the aetiological heterogeneity that drives the development and progression of T2D. This might offer a route to optimize global access to genetically informed diabetes care.
Whole genome sequencing (WGS) at high-depth (30X) allows the accurate discovery of variants in the coding and non-coding DNA regions and helps elucidate the genetic underpinnings of human health and diseases. Yet, due to the prohibitive cost of high-depth WGS, most large-scale genetic association studies use genotyping arrays or high-depth whole exome sequencing (WES). Here we propose a cost-effective method which we call “Whole Exome Genome Sequencing” (WEGS), that combines low-depth WGS and high-depth WES with up to 8 samples pooled and sequenced simultaneously (multiplexed). We experimentally assess the performance of WEGS with four different depth of coverage and sample multiplexing configurations. We show that the optimal WEGS configurations are 1.7–2.0 times cheaper than standard WES (no-plexing), 1.8–2.1 times cheaper than high-depth WGS, reach similar recall and precision rates in detecting coding variants as WES, and capture more population-specific variants in the rest of the genome that are difficult to recover when using genotype imputation methods. We apply WEGS to 862 patients with peripheral artery disease and show that it directly assesses more known disease-associated variants than a typical genotyping array and thousands of non-imputable variants per disease-associated locus.
The human leukocyte antigen (HLA) region on chromosome 6 is strongly associated with many immune-mediated and infection-related diseases. Due to its highly polymorphic nature and complex linkage disequilibrium patterns, traditional genetic association studies of single nucleotide polymorphisms do not perform well in this region. Instead, the field has adopted the assessment of the association of HLA alleles (i.e., entire HLA gene haplotypes) with disease. Often based on genotyping arrays, these association studies impute HLA alleles, decreasing accuracy and thus statistical power for rare alleles and in non-European ancestries. Here, we use whole-exome sequencing (WES) from 454,824 UK Biobank (UKB) participants to directly call HLA alleles using the HLA-HD algorithm. We show this method is more accurate than imputing HLA alleles and harness the improved statistical power to identify 360 associations for 11 auto-immune phenotypes (at least 129 likely novel), leading to better insights into the specific coding polymorphisms that underlie these diseases. We show that HLA alleles with synonymous variants, often overlooked in HLA studies, can significantly influence these phenotypes. Lastly, we show that HLA sequencing may improve polygenic risk scores accuracy across ancestries. These findings allow better characterization of the role of the HLA region in human disease.
Associations between human genetic variation and clinical phenotypes have become a foundation of biomedical research. Most repositories of these data seek to be disease-agnostic and therefore lack disease-focused views. The Type 2 Diabetes Knowledge Portal (T2DKP) is a public resource of genetic datasets and genomic annotations dedicated to type 2 diabetes (T2D) and related traits. Here, we seek to make the T2DKP more accessible to prospective users and more useful to existing users. First, we evaluate the T2DKP’s comprehensiveness by comparing its datasets with those of other repositories. Second, we describe how researchers unfamiliar with human genetic data can begin using and correctly interpreting them via the T2DKP. Third, we describe how existing users can extend their current workflows to use the full suite of tools offered by the T2DKP. We finally discuss the lessons offered by the T2DKP toward the goal of democratizing access to complex disease genetic results.
The substantial investments in human genetics and genomics made over the past three decades were anticipated to result in many innovative therapies. Here we investigate the extent to which these expectations have been met, excluding cancer treatments. In our search, we identified 40 germline genetic observations that led directly to new targets and subsequently to novel approved therapies for 36 rare and 4 common conditions. The median time between genetic target discovery and drug approval was 25 years. Most of the genetically driven therapies for rare diseases compensate for disease-causing loss-of-function mutations. The therapies approved for common conditions are all inhibitors designed to pharmacologically mimic the natural, disease-protective effects of rare loss-of-function variants. Large biobank-based genetic studies have the power to identify and validate a large number of new drug targets. Genetics can also assist in the clinical development phase of drugs—for example, by selecting individuals who are most likely to respond to investigational therapies. This approach to drug development requires investments into large, diverse cohorts of deeply phenotyped individuals with appropriate consent for genetically assisted trials. A robust framework that facilitates responsible, sustainable benefit sharing will be required to capture the full potential of human genetics and genomics and bring effective and safe innovative therapies to patients quickly.
Type 2 diabetes (T2D) is a heterogeneous disease that develops through diverse pathophysiological processes. To characterise the genetic contribution to these processes across ancestry groups, we aggregate genome-wide association study (GWAS) data from 2,535,601 individuals (39.7% non-European ancestry), including 428,452 T2D cases. We identify 1,289 independent association signals at genome-wide significance (P<5×10-8) that map to 611 loci, of which 145 loci are previously unreported. We define eight non-overlapping clusters of T2D signals characterised by distinct profiles of cardiometabolic trait associations. These clusters are differentially enriched for cell-type specific regions of open chromatin, including pancreatic islets, adipocytes, endothelial, and enteroendocrine cells. We build cluster-specific partitioned genetic risk scores (GRS) in an additional 137,559 individuals of diverse ancestry, including 10,159 T2D cases, and test their association with T2D-related vascular outcomes. Cluster-specific partitioned GRS are more strongly associated with coronary artery disease and end-stage diabetic nephropathy than an overall T2D GRS across ancestry groups, highlighting the importance of obesity-related processes in the development of vascular outcomes. Our findings demonstrate the value of integrating multi-ancestry GWAS with single-cell epigenomics to disentangle the aetiological heterogeneity driving the development and progression of T2D, which may offer a route to optimise global access to genetically-informed diabetes care.