Population-level nuclear magnetic resonance (NMR)-based lipoprotein profiling is a key tool for investigating dysregulated lipoprotein metabolism and its role in cardiovascular disorders. However, associations with size-resolved lipoprotein composition readouts are difficult to dissect, as these traits are often highly correlated. Derived variables can therefore be more relevant to biological interpretation. Here, we show that the total fatty acid (FA) content of the lipoproteins derived from their lipid headgroup concentrations using the formula FA = 3TG + 2PL + CE strongly correlates (Spearman rho = 0.98) with the independently measured total fatty acid content. This observation is not self-evident since these variables are determined using different portions of the NMR spectrum. Using NMR data acquired on the Nightingale platform for 274,303 UK Biobank (UKB) participants and genetic associations as a readout, we then show that this relationship also holds at the size-resolved lipoprotein level. We identified eight gene loci where the proposed FA variables display a significantly stronger genetic association signal than that of the corresponding lipid headgroup variables. Five of these loci (LIPC, LIPG, PLB1, LPL, and APOC3) have a direct function in FA metabolism. Including computed size-resolved FA variables may therefore improve the biological interpretation of future studies based on Nightingale's data.
Accurate quantification of circulating proteins is critical for assessing biological variation and integrating proteomics with other omics to understand biological processes and disease mechanisms. Protein measurements, however, can be substantially influenced by preanalytical variability arising from differences in sample collection, handling, and storage, whereas technical variation introduced by the assay and workflow is typically well controlled through established validation procedures. Identifying proteins that capture these systematic influences enables their incorporation into downstream analyzes, thereby improving statistical power. In this study, we applied highly multiplexed aptamer-based affinity proteomics to plasma samples from three independent cohorts─German, Arab-Asian and Qatari to evaluate how adjusting for all measured proteins influences protein quantitative trait loci (pQTLs) associations. Using the p-gain statistic as an indicator of improved association strength, we identified clusters of proteins whose covariation patterns suggested potential preanalytical effects. One cluster contained HSP90 (Heat Shock Protein 90), a marker linked to white blood cell lysis, while others were enriched for proteins involved in complement and coagulation cascades or platelet activation. Our work presents a data-driven framework for detecting latent sources of variation in large-scale proteomic data sets and lay the groundwork for future efforts to quantify the impact of hidden confounding factors.
Dysregulated blood lipids are a major predictor of cardiovascular events. A recent genome-wide association study (GWAS) with five clinically relevant lipid traits in 1.65 million individuals implicated over 770 genomic regions in regulating blood lipid metabolism. To translate these associations into clinical applications, a functional understanding of their roles in lipoprotein metabolism, transport and remodeling (LPmtr) is required. Here, we report the deep molecular fine-mapping of 554 of these lipid risk loci using 168 lipoprotein-related traits and all possible ratios between them in over 273,000 participants of the UK Biobank. We identified new ratio-based markers of pathways shared by multiple LPmtr genes, such as the linoleic acid fraction of the polyunsaturated fatty acid pool to reveal potential causal genes at poorly characterized lipid risk loci, the percentage of esterified cholesterol moieties in LDL particles as a proxy for soluble LDL receptor levels, and the HDL fraction of total lipoprotein particle number as a predictor of incident myocardial infarction. We demonstrate how lipoprotein fine-mapping can generate new hypotheses for drug target development while uncovering new mechanisms relevant to hyperlipidemia. Ratio-driven clustering further implicated miR-148 in TG secretion, linking ER-stress responses at postprandial state to VLDL metabolism via mTORC1, shown through series of integrated cellular assays and mouse studies. Moreover, consistent with its regulatory influence on lipid flux we identify miR-148a a previously unrecognized determinat of Lp(a) levels. Our study implements a novel approach of using metabolomic data to follow-up on genetic evidence from GWAS with clinical traits and generates new insights into the biology of lipoprotein particles, supporting the emerging view that assessing lipoprotein size and composition is essential for the understanding, prevention, and treatment of lipid-related disorders.
Most studies to date of protein quantitative trait loci (pQTLs) have relied on affinity proteomics platforms, which provide only limited information about the targeted protein isoforms and may be affected by genetic variation in their epitope binding. Here we show that mass spectrometry (MS)-based proteomics can complement these studies and provide insights into the role of specific protein isoform and epitope-altering variants. Using the Seer Proteograph nanoparticle enrichment MS platform, we identified and replicated new pQTLs in a genome-wide association study of proteins in blood plasma samples from two cohorts and evaluated previously reported pQTLs from affinity proteomics platforms. We found that >30% of the evaluated pQTLs were confirmed by MS proteomics to be consistent with the hypothesis that genetic variants induce changes in protein abundance, whereas another 30% could not be replicated and are possibly due to epitope effects, although alternative explanations for nonreplication need to be considered on a case-by-case basis.
The steroid hormone progesterone (P4) regulates multiple aspects of reproductive and metabolic physiology. Classical P4 signaling operates through nuclear receptors that regulate transcription. In addition, P4 signals through membrane P4 receptors (mPRs) in a rapid nongenomic modality. Despite the established physiological importance of P4 nongenomic signaling, the details of its signal transduction cascade remain elusive. Here, using Xenopus oocyte maturation as a well-established physiological readout of nongenomic P4 signaling, we identify the lipid hydrolase ABHD2 (α/β hydrolase domain-containing protein 2) as an essential mPRβ co-receptor to trigger meiosis. We show using functional assays coupled to unbiased and targeted cell-based lipidomics that ABHD2 possesses a phospholipase A2 (PLA2) activity that requires mPRβ. This PLA2 activity bifurcates P4 signaling by inducing clathrin-dependent endocytosis of mPRβ, resulting in the production of lipid messengers that are G-protein coupled receptor agonists. Therefore, P4 drives meiosis by inducing an ABHD2 PLA2 activity that requires both mPRβ and ABHD2 as obligate co-receptors.
Recent advances in high-throughput measurement technologies have enabled the analysis of molecular perturbations associated with disease phenotypes at the multi-omic level. Such perturbations can range in scale from fluctuations of individual molecules to entire biological pathways. Data-driven clustering algorithms have long been used to group interactions into interpretable functional modules; however, these modules are typically constrained to a fixed size or statistical cutoff. Furthermore, modules are often analyzed independently of their broader biological context. Consequently, such clustering approaches limit the ability to explore functional module associations with disease phenotypes across multiple scales. Here, we introduce AutoFocus, a data-driven method that hierarchically organizes biomolecules and tests for phenotype enrichment at every level within the hierarchy. As a result, the method allows disease-associated modules to emerge at any scale. We evaluated this approach using two datasets: First, we explored associations of biomolecules from the multi-omic QMDiab dataset (n = 388) with the well-characterized type 2 diabetes phenotype. Secondly, we utilized the ROS/MAP Alzheimer’s disease dataset (n = 500), consisting of high-throughput measurements of brain tissue to explore modules associated with multiple Alzheimer’s Disease-related phenotypes. Our method identifies modules that are multi-omic, span multiple pathways, and vary in size. We provide an interactive tool to explore this hierarchy at different levels and probe enriched modules, empowering users to examine the full hierarchy, delve into biomolecular drivers of disease phenotype within a module, and incorporate functional annotations.
Abstract Background Immunoglobulin (Ig) glycosylation modulates the immune response and plays a critical role in ageing and diseases. Studies have mainly focused on IgG glycosylation, and little is known about the genetics and epidemiology of IgA glycosylation. Methods We generated, using a novel liquid chromatography-mass spectrometry method, the first large-scale IgA glycomics dataset in serum from 2423 twins, encompassing 71 N- and O-glycan species. Results We showed that, despite the lack of a direct genetic template, glycosylation is highly heritable, and that glycopeptide structures are sex-specific, and undergo substantial changes with ageing. We observe extensive correlations between the IgA and IgG glycomes, and, exploiting the twin design, show that they are predominantly influenced by shared genetic factors. A genome-wide association study identified eight loci associated with both the IgA and IgG glycomes (ST6GAL1, ELL2, B4GALT1, ABCF2, TMEM121, SLC38A10, SMARCB1, and MGAT3) and two novel loci specifically modulating IgA O-glycosylation (C1GALT1 and ST3GAL1). Validation of our findings in an independent cohort of 320 individuals from Qatar showed that the underlying genetic architecture is conserved across ancestries. Conclusions Our study delineates the genetic landscape of IgA glycosylation and provides novel potential functional links with the aetiology of complex immune diseases, including genetic factors involved in IgA nephropathy risk.
In-depth multiomic phenotyping provides molecular insights into complex physiological processes and their pathologies. Here, we report on integrating 18 diverse deep molecular phenotyping (omics-) technologies applied to urine, blood, and saliva samples from 391 participants of the multiethnic diabetes Qatar Metabolomics Study of Diabetes (QMDiab). Using 6,304 quantitative molecular traits with 1,221,345 genetic variants, methylation at 470,837 DNA CpG sites, and gene expression of 57,000 transcripts, we determine (1) within-platform partial correlations, (2) between-platform mutual best correlations, and (3) genome-, epigenome-, transcriptome-, and phenome-wide associations. Combined into a molecular network of > 34,000 statistically significant trait-trait links in biofluids, our study portrays "The Molecular Human". We describe the variances explained by each omics in the phenotypes (age, sex, BMI, and diabetes state), platform complementarity, and the inherent correlation structures of multiomics data. Further, we construct multi-molecular network of diabetes subtypes. Finally, we generated an open-access web interface to "The Molecular Human" (http://comics.metabolomix.com), providing interactive data exploration and hypotheses generation possibilities.
Genome-wide association studies (GWAS) with proteomics are essential tools for drug discovery. To date, most studies have used affinity proteomics platforms, which have limited discovery to protein panels covered by the available affinity binders. Furthermore, it is not clear to which extent protein epitope changing variants interfere with the detection of protein quantitative trait loci (pQTLs). Mass spectrometry-based (MS) proteomics can overcome some of these limitations. Here we report a GWAS using the MS-based Seer Proteograph™ platform with blood samples from a discovery cohort of 1,260 American participants and a replication in 325 individuals from Asia, with diverse ethnic backgrounds. We analysed 1,980 proteins quantified in at least 80% of the samples, out of 5,753 proteins quantified across the discovery cohort. We identified 252 and replicated 90 pQTLs, where 30 of the replicated pQTLs have not been reported before. We further investigated 200 of the strongest associated cis-pQTLs previously identified using the SOMAscan and the Olink platforms and found that up to one third of the affinity proteomics pQTLs may be affected by epitope effects, while another third were confirmed by MS proteomics to be consistent with the hypothesis that genetic variants induce changes in protein expression. The present study demonstrates the complementarity of the different proteomics approaches and reports pQTLs not accessible to affinity proteomics, suggesting that many more pQTLs remain to be discovered using MS-based platforms.
Proteogenomics studies generate hypotheses on protein function and provide genetic evidence for drug target prioritization. Most previous work has been conducted using affinity-based proteomics approaches. These technologies face challenges, such as uncertainty regarding target identity, non-specific binding, and handling of variants that affect epitope affinity binding. Mass spectrometry-based proteomics can overcome some of these challenges. Here we report a pQTL study using the Proteograph™ Product Suite workflow ( Seer, Inc .) where we quantify over 18,000 unique peptides from nearly 3000 proteins in more than 320 blood samples from a multi-ethnic cohort in a bottom-up, peptide-centric, mass spectrometry-based proteomics approach. We identify 184 protein-altering variants in 137 genes that are significantly associated with their corresponding variant peptides, confirming target specificity of co-associated affinity binders, identifying putatively causal cis -encoded proteins and providing experimental evidence for their presence in blood, including proteins that may be inaccessible to affinity-based proteomics.
AbstractImmunoglobulin (Ig) glycosylation modulates the immune response, and plays a critical role in ageing and diseases. Studies have mainly focused on IgG glycosylation, and little is known about the genetics and epidemiology of IgA glycosylation. Here, we generated, using a novel LC-MS method, the first large-scale IgA glycomics dataset in serum from 2,423 twins, encompassing 71N-andO-glycan species. We showed that, despite the lack of a direct genetic template, glycosylation is highly heritable, and that glycopeptide structures are sex-specific, and undergo substantial changes with ageing. We observe extensive correlations between the IgA and IgG glycomes, and, exploiting the twin design, show that they are predominantly influenced by shared genetic factors. A genome-wide association study identified eight loci associated with both the IgA and IgG glycomes (ST6GAL1,ELL2,B4GALT1,ABCF2,TMEM121,SLC38A10,SMARCB1,MGAT3), and two novel loci specifically modulating IgAO-glycosylation (C1GALT1andST3GAL1). Validation of our findings in an independent cohort of 320 individuals from Qatar showed that the underlying genetic architecture is conserved across ethnicities. Our study delineates the genetic landscape of IgA glycosylation and provides novel potential functional links with the aetiology of complex immune diseases, including genetic factors involved in IgA nephropathy risk.
Cancer cells frequently undergo metabolic reprogramming as a mechanism of resistance against chemotherapeutic drugs. Metabolomic profiling provides a direct readout of metabolic changes and can thus be used to identify these tumor escape mechanisms. Here, we introduce piTracer, a computational tool that uses multi-scale molecular networks to identify potential combination therapies from pre- and post-treatment metabolomics data. We first demonstrate piTracer’s core ability to reconstruct cellular cascades by inspecting well-characterized molecular pathways and previously studied associations between genetic variants and metabolite levels. We then apply a new gene ranking algorithm on differential metabolomic profiles from human breast cancer cells after glutaminase inhibition. Four of the automatically identified gene targets were experimentally tested by simultaneous inhibition of the respective targets and glutaminase. Of these combination treatments, two were be confirmed to induce synthetic lethality in the cell line. In summary, piTracer integrates the molecular monitoring of escape mechanisms into comprehensive pathway networks to accelerate drug target identification. The tool is open source and can be accessed at https://github.com/krumsieklab/pitracer .
BACKGROUND:Human plasma contains a wide variety of circulating proteins. These proteins can be important clinical biomarkers in disease and also possible drug targets. Large scale genomics studies of circulating proteins can identify genetic variants that lead to relative protein abundance.METHODS:We conducted a meta-analysis on genome-wide association studies of autosomal chromosomes in 22,997 individuals of primarily European ancestry across 12 cohorts to identify protein quantitative trait loci (pQTL) for 92 cardiometabolic associated plasma proteins.RESULTS:We identified 503 (337 cis and 166 trans) conditionally independent pQTLs, including several novel variants not reported in the literature. We conducted a sex-stratified analysis and found that 118 (23.5%) of pQTLs demonstrated heterogeneity between sexes. The direction of effect was preserved but there were differences in effect size and significance. Additionally, we annotate trans-pQTLs with nearest genes and report plausible biological relationships. Using Mendelian randomization, we identified causal associations for 18 proteins across 19 phenotypes, of which 10 have additional genetic colocalization evidence. We highlight proteins associated with a constellation of cardiometabolic traits including angiopoietin-related protein 7 (ANGPTL7) and Semaphorin 3F (SEMA3F).CONCLUSION:Through large-scale analysis of protein quantitative trait loci, we provide a comprehensive overview of common variants associated with plasma proteins. We highlight possible biological relationships which may serve as a basis for further investigation into possible causal roles in cardiometabolic diseases.
Exercise-induced hemolysis occurs as the result of intense physical exercise and is caused by metabolic and mechanical factors including repeated muscle contractions leading to capillary vessels compression, vasoconstriction of internal organs and foot strike among others. We hypothesized that exercise-induced hemolysis occurred in endurance racehorses and its severity was associated with the intensity of exercise. To provide further insight into the hemolysis of endurance horses, the aim of the study was to deployed a strategy for small molecules (metabolites) profiling, beyond standard molecular methods. The study included 47 Arabian endurance horses competing for either 80, 100, or 120 km distances. Blood plasma was collected before and after the competition and analyzed macroscopically, by ELISA and non-targeted metabolomics with liquid chromatography-mass spectrometry. A significant increase in all hemolysis parameters was observed after the race, and an association was found between the measured parameters, average speed, and distance completed. Levels of hemolysis markers were highest in horses eliminated for metabolic reasons in comparison to finishers and horses eliminated for lameness (gait abnormality), which may suggest a connection between the intensity of exercise, metabolic challenges, and hemolysis. Utilization of omics methods alongside conventional methods revealed a broader insight into the exercise-induced hemolysis process by displaying, apart from commonly measured hemoglobin and haptoglobin, levels of hemoglobin degradation metabolites. Obtained results emphasized the importance of respecting horse limitations in regard to speed and distance which, if underestimated, may lead to severe damages.
ABSTRACT Genome-wide association studies (GWAS) with proteomics generate hypotheses on protein function and offer genetic evidence for drug target prioritization. Although most protein quantitative loci (pQTLs) have so far been identified by high-throughput affinity proteomics platforms, these methods also have some limitations, such as uncertainty about target identity, non-specific binding of aptamers, and inability to handle epitope-modifying variants that affect affinity binding. Mass spectrometry (MS) proteomics has the potential to overcome these challenges and broaden the scope of pQTL studies. Here, we employ the recently developed MS-based Proteograph™ workflow ( Seer, Inc .) to quantify over 18,000 unique peptides from almost 3,000 proteins in more than 320 blood samples from a multi-ethnic cohort. We implement a bottom-up MS-proteomics approach for the detection and quantification of blood-circulating proteins in the presence of protein altering variants (PAVs). We identify 184 PAVs located in 137 genes that are significantly associated with their corresponding variant peptides in MS data (MS-PAVs). Half of these MS-PAVs (94) overlap with cis -pQTLs previously identified by affinity proteomics pQTL studies, thus confirming the target specificity of the affinity binders. An additional 54 MS-PAVs overlap with trans -pQTLs (and not cis -pQTLs) in affinity proteomics studies, thus identifying the putatively causal cis -encoded protein and providing experimental evidence for its presence in blood. The remaining 36 MS-PAVs have not been previously reported and include proteins that may be inaccessible to affinity proteomics, such as a variant in the incretin pro-peptide (GIP) that associates with type 2 diabetes and cardiovascular disease. Overall, our study introduces a novel approach for analyzing MS-based proteomics data within the GWAS context, provides new insights relevant to genetics-based drug discovery, and highlights the potential of MS-proteomics technologies when applied at population scale. Highlights This is the first pQTL study that uses the Proteograph ™ ( Seer Inc .) mass spectrometry-based proteomics workflow. We introduce a novel bottom-up proteomics approach that accounts for protein altering variants in the detection of pQTLs. We confirm the target and potential epitope effects of affinity binders for cis- pQTLs from affinity proteomics studies. We establish putatively causal proteins for known affinity proteomics trans -pQTLs and confirm their presence in blood. We identify novel protein altering variants in proteins of clinical relevance that may not be accessible to affinity proteomics. Graphical abstract
BACKGROUND:Bardet-Biedl syndrome (BBS) is an autosomal recessive, genetically heterogeneous, pleiotropic disorder caused by variants in genes involved in the function of the primary cilium. We have harnessed genomics to identify BBS and ophthalmic technologies to describe novel features of BBS.CASE PRESENTATION:A patient with an unclear diagnosis of syndromic type 2 diabetes mellitus, another affected sibling and unaffected siblings and parents were sequenced using DNA extracted from saliva samples. Corneal confocal microscopy (CCM) and retinal spectral domain optical coherence tomography (SD-OCT) were used to identify novel ophthalmic features in these patients. The two affected individuals had a homozygous variant in C8orf37 (p.Trp185*). SD-OCT and CCM demonstrated a marked and patchy reduction in the retinal nerve fiber layer thickness and loss of corneal nerve fibers, respectively.CONCLUSION:This report highlights the use of ophthalmic imaging to identify novel retinal and corneal abnormalities that extend the phenotype of BBS in a patient with syndromic type 2 diabetes.
EDITORIAL article Front. Endocrinol., 22 June 2022Sec. Systems Endocrinology https://doi.org/10.3389/fendo.2022.948991
Type 2 diabetes (T2D) has a heterogeneous etiology influencing its progression, treatment, and complications. A data driven cluster analysis in European individuals with T2D previously identified four subtypes: severe insulin deficient (SIDD), severe insulin resistant (SIRD), mild obesity-related (MOD), and mild age-related (MARD) diabetes. Here, the clustering approach was applied to individuals with T2D from the Qatar Biobank and validated in an independent set. Cluster-specific signatures of circulating metabolites and proteins were established, revealing subtype-specific molecular mechanisms, including activation of the complement system with features of autoimmune diabetes and reduced 1,5-anhydroglucitol in SIDD, impaired insulin signaling in SIRD, and elevated leptin and fatty acid binding protein levels in MOD. The MARD cluster was the healthiest with metabolomic and proteomic profiles most similar to the controls. We have translated the T2D subtypes to an Arab population and identified distinct molecular signatures to further our understanding of the etiology of these subtypes.
Genome-wide association studies (GWAS) with non-targeted metabolomics have identified many genetic loci of biomedical interest. However, metabolites with a high degree of missingness, such as drug metabolites and xenobiotics, are often excluded from such studies due to a lack of statistical power and higher uncertainty in their quantification. Here we propose ratios between related drug metabolites as GWAS phenotypes that can drastically increase power to detect genetic associations between pairs of biochemically related molecules. As a proof-of-concept we conducted a GWAS with 520 individuals from the Qatar Biobank for who at least five of the nine available acetaminophen metabolites have been detected. We identified compelling evidence for genetic variance in acetaminophen glucuronidation and methylation by UGT2A15 and COMT, respectively. Based on the metabolite ratio association profiles of these two loci we hypothesized the chemical structure of one of their products or substrates as being 3-methoxyacetaminophen, which we then confirmed experimentally. Taken together, our study suggests a novel approach to analyze metabolites with a high degree of missingness in a GWAS setting with ratios, and it also demonstrates how pharmacological pathways can be mapped out using non-targeted metabolomics measurements in large population-based studies.
In-depth multiomics phenotyping can provide a molecular understanding of complex physiological processes and their pathologies. Here, we report on the application of 18 diverse deep molecular phenotyping (omics-) technologies to urine, blood, and saliva samples from 391 participants of the multiethnic diabetes study QMDiab. We integrated quantitative readouts of 6,304 molecular traits with 1,221,345 genetic variants, methylation at 470,837 DNA CpG sites, and gene expression of 57,000 transcripts using between-platform mutual best correlations, within-platform partial correlations, and genome-, epigenome-, transcriptome-, and phenome-wide associations. The achieved molecular network covers over 34,000 statistically significant trait-trait links and illustrates “The Molecular Human”. We describe the variances explained by each omics layer in the phenotypes age, sex, BMI, and diabetes state, platform complementarity, and the inherent correlation structures of multiomics. Finally, we discuss biological aspects of the networks relevant to the molecular basis of complex disorders. We developed a web-based interface to “The Molecular Human”, which is freely accessible at http://comics.metabolomix.com and allows dynamic interaction with the data.